Triangle104/Chronos-Gold-12B-1.0-Q4_K_S-GGUF

This model was converted to GGUF format from elinas/Chronos-Gold-12B-1.0 using llama.cpp via the ggml.ai's GGUF-my-repo space. Refer to the original model card for more details on the model.

Model details:

Chronos Gold 12B 1.0 is a very unique model that applies to domain areas such as general chatbot functionatliy, roleplay, and storywriting. The model has been observed to write up to 2250 tokens in a single sequence. The model was trained at a sequence length of 16384 (16k) and will still retain the apparent 128k context length from Mistral-Nemo, though it deteriorates over time like regular Nemo does based on the RULER Test

As a result, is recommended to keep your sequence length max at 16384, or you will experience performance degredation.

The base model is mistralai/Mistral-Nemo-Base-2407 which was heavily modified to produce a more coherent model, comparable to much larger models.

Chronos Gold 12B-1.0 re-creates the uniqueness of the original Chronos with significiantly enhanced prompt adherence (following), coherence, a modern dataset, as well as supporting a majority of "character card" formats in applications like SillyTavern.

It went through an iterative and objective merge process as my previous models and was further finetuned on a dataset curated for it.

The specifics of the model will not be disclosed at the time due to dataset ownership. Instruct Template

This model uses ChatML - below is an example. It is a preset in many frontends.

Use with llama.cpp

Install llama.cpp through brew (works on Mac and Linux)

brew install llama.cpp

Invoke the llama.cpp server or the CLI.

CLI:

llama-cli --hf-repo Triangle104/Chronos-Gold-12B-1.0-Q4_K_S-GGUF --hf-file chronos-gold-12b-1.0-q4_k_s.gguf -p "The meaning to life and the universe is"

Server:

llama-server --hf-repo Triangle104/Chronos-Gold-12B-1.0-Q4_K_S-GGUF --hf-file chronos-gold-12b-1.0-q4_k_s.gguf -c 2048

Note: You can also use this checkpoint directly through the usage steps listed in the Llama.cpp repo as well.

Step 1: Clone llama.cpp from GitHub.

git clone https://github.com/ggerganov/llama.cpp

Step 2: Move into the llama.cpp folder and build it with LLAMA_CURL=1 flag along with other hardware-specific flags (for ex: LLAMA_CUDA=1 for Nvidia GPUs on Linux).

cd llama.cpp && LLAMA_CURL=1 make

Step 3: Run inference through the main binary.

./llama-cli --hf-repo Triangle104/Chronos-Gold-12B-1.0-Q4_K_S-GGUF --hf-file chronos-gold-12b-1.0-q4_k_s.gguf -p "The meaning to life and the universe is"

./llama-server --hf-repo Triangle104/Chronos-Gold-12B-1.0-Q4_K_S-GGUF --hf-file chronos-gold-12b-1.0-q4_k_s.gguf -c 2048

Downloads last month: 6

GGUF

Model size

12B params

Architecture

llama

Hardware compatibility

4-bit

Inference Providers NEW

This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Triangle104/Chronos-Gold-12B-1.0-Q4_K_S-GGUF

Base model

mistralai/Mistral-Nemo-Base-2407

Finetuned

elinas/Chronos-Gold-12B-1.0

Quantized

(10)

this model

Collections including Triangle104/Chronos-Gold-12B-1.0-Q4_K_S-GGUF

Mistral

Collection

MistralAI-based models • 927 items • Updated Aug 1 • 1

RP

Collection

Roleplaying Models • 1903 items • Updated Aug 1 • 10

Evaluation results

strict accuracy on IFEval (0-Shot)
Open LLM Leaderboard

31.660
normalized accuracy on BBH (3-Shot)
Open LLM Leaderboard

35.910
exact match on MATH Lvl 5 (4-Shot)
Open LLM Leaderboard

4.380
acc_norm on GPQA (0-shot)
Open LLM Leaderboard

9.060
acc_norm on MuSR (0-shot)
Open LLM Leaderboard

19.420
accuracy on MMLU-PRO (5-shot)
test set Open LLM Leaderboard

27.980

View on Papers With Code