CompendiumLabs
/

mistral-nemo-instruct-2407-gguf

Inference Endpoints

Model card Files Files and versions Community

iamlemec commited on Jul 20

Commit

c332381

•

1 Parent(s): 15a9dba

Update README.md

Files changed (1) hide show

README.md +9 -3

README.md CHANGED Viewed

@@ -1,3 +1,9 @@
----
-license: apache-2.0
----

+---
+license: apache-2.0
+---
+This is a quantized GGUF of [mistralai/Mistral-Nemo-Instruct-2407](https://huggingface.co/mistralai/Mistral-Nemo-Instruct-2407).
+Right now, to run it in `llama.cpp` you'll need to use [PR #8577](https://github.com/ggerganov/llama.cpp/pull/8604) or equivalently the fork [iamlemec/llama.cpp](https://github.com/iamlemec/llama.cpp/tree/mistral-nemo).
+Currently, we just have a [`Q5_K`](https://huggingface.co/CompendiumLabs/mistral-nemo-instruct-2407-gguf/blob/main/mistral-nemo-instruct-q5_k.gguf) quantization which comes in at 8.73 GB. If you're interested other quantizations, just ping me [@iamlemec](https://twitter.com/iamlemec) on Twitter.