Upload folder using huggingface_hub

Browse files

Files changed (12) hide show

README.md +35 -175
added_tokens.json +40 -0
config.json +48 -0
mergekit_config.yml +13 -0
merges.txt +0 -0
model-00001-of-00002.safetensors +3 -0
model-00002-of-00002.safetensors +3 -0
model.safetensors.index.json +1 -0
special_tokens_map.json +23 -0
tokenizer.json +0 -0
tokenizer_config.json +323 -0
vocab.json +0 -0

README.md CHANGED Viewed

@@ -1,187 +1,47 @@
-# mergekit
-`mergekit` is a toolkit for merging pre-trained language models. `mergekit` uses an out-of-core approach to perform unreasonably elaborate merges in resource-constrained situations. Merges can be run entirely on CPU or accelerated with as little as 8 GB of VRAM. Many merging algorithms are supported, with more coming as they catch my attention.
-Features:
-- Supports Llama, Mistral, GPT-NeoX, StableLM, and more
-- Many [merge methods](#merge-methods)
-- GPU or CPU execution
-- Lazy loading of tensors for low memory use
-- Interpolated gradients for parameter values (inspired by Gryphe's [BlockMerge_Gradient](https://github.com/Gryphe/BlockMerge_Gradient) script)
-- Piecewise assembly of language models from layers ("Frankenmerging")
-🔊 Call to Evolve - to solve evolutionary merge methods as a community - please see https://github.com/arcee-ai/mergekit/issues/207.
-## Installation
-```sh
-git clone https://github.com/cg123/mergekit.git
-cd mergekit
-pip install -e .  # install the package and make scripts available
-```
-If the above fails with the error of:
-```
-ERROR: File "setup.py" or "setup.cfg" not found. Directory cannot be installed in editable mode:
-(A "pyproject.toml" file was found, but editable mode currently requires a setuptools-based build.)
-```
-You may need to upgrade pip to > 21.3 with the command `python3 -m pip install --upgrade pip`
-## Usage
-The script `mergekit-yaml` is the main entry point for `mergekit`. It takes a YAML configuration file and an output path, like so:
-```sh
-mergekit-yaml path/to/your/config.yml ./output-model-directory [--cuda] [--lazy-unpickle] [--allow-crimes] [... other options]
-```
-This will run the merge and write your merged model to `./output-model-directory`.
-For more information on the arguments accepted by `mergekit-yaml` run the command `mergekit-yaml --help`.
-### Uploading to Huggingface
-When you have a merged model you're happy with, you may want to share it on the Hugging Face Hub. `mergekit` generates a `README.md` for your merge with some basic information for a model card. You can edit it to include more details about your merge, like giving it a good name or explaining what it's good at; rewrite it entirely; or use the generated `README.md` as-is. It is also possible to edit your `README.md` online once it has been uploaded to the Hub.
-Once you're happy with your model card and merged model, you can upload it to the Hugging Face Hub using the [huggingface_hub](https://huggingface.co/docs/huggingface_hub/index) Python library.
-```
-# log in to huggingface with an access token (must have write permission)
-huggingface-cli login
-# upload your model
-huggingface-cli upload your_hf_username/my-cool-model ./output-model-directory .
-```
-The [documentation](https://huggingface.co/docs/huggingface_hub/guides/cli#huggingface-cli-upload) for `huggingface_hub` goes into more detail about other options for uploading.
-## Merge Configuration
-Merge configurations are YAML documents specifying the operations to perform in order to produce your merged model.
-Below are the primary elements of a configuration file:
-- `merge_method`: Specifies the method to use for merging models. See [Merge Methods](#merge-methods) for a list.
-- `slices`: Defines slices of layers from different models to be used. This field is mutually exclusive with `models`.
-- `models`: Defines entire models to be used for merging. This field is mutually exclusive with `slices`.
-- `base_model`: Specifies the base model used in some merging methods.
-- `parameters`: Holds various parameters such as weights and densities, which can also be specified at different levels of the configuration.
-- `dtype`: Specifies the data type used for the merging operation.
-- `tokenizer_source`: Determines how to construct a tokenizer for the merged model.
-### Parameter Specification
-Parameters are flexible and can be set with varying precedence. They can be specified conditionally using tensor name filters, which allows finer control such as differentiating between attention heads and fully connected layers.
-Parameters can be specified as:
-- **Scalars**: Single floating-point values.
-- **Gradients**: List of floating-point values, specifying an interpolated gradient.
-The parameters can be set at different levels, with decreasing precedence as follows:
-1. `slices.*.sources.parameters` - applying to a specific input slice
-2. `slices.*.parameters` - applying to a specific output slice
-3. `models.*.parameters` or `input_model_parameters` - applying to any tensors coming from specific input models
-4. `parameters` - catchall
-### Tokenizer Source
-The `tokenizer_source` field of a configuration file determines what tokenizer is used by the merged model. This also effects how embeddings and language model heads are merged.
-This functionality is still experimental and may break. Please file an issue if you encounter any issues with it.
-Valid values:
-- `base`: use the tokenizer from the base model
-- `union`: construct a tokenizer with all tokens from all models
-- `model:<model_path>`: use the tokenizer from a specific model
-If set, mergekit will find a mapping between each model's vocabulary and the output tokenizer. This allows models with different vocabularies or added tokens to be meaningfully merged.
-`tokenizer_source` is compatible with all merge methods, but when used `lm_head`/`embed_tokens` will be merged linearly. For two-model merges, the `embed_slerp` parameter can be set to `true` to use SLERP instead.
-If the `tokenizer_source` field is not set, mergekit will fall back to its legacy default behavior. The tokenizer for the base model (or first model in the merge, if no base model is specified) will be copied to the output directory. The parameter matrices for `lm_head`/`embed_tokens` will be truncated to the smallest size present in the merge. In _most_ cases this corresponds to using the tokenizer for the base model.
-### Examples
-Several examples of merge configurations are available in [`examples/`](examples/).
-## Merge Methods
-A quick overview of the currently supported merge methods:
-| Method                                                                                       | `merge_method` value | Multi-Model | Uses base model |
-| -------------------------------------------------------------------------------------------- | -------------------- | ----------- | --------------- |
-| Linear ([Model Soups](https://arxiv.org/abs/2203.05482))                                     | `linear`             | ✅          | ❌              |
-| SLERP                                                                                        | `slerp`              | ❌          | ✅              |
-| [Task Arithmetic](https://arxiv.org/abs/2212.04089)                                          | `task_arithmetic`    | ✅          | ✅              |
-| [TIES](https://arxiv.org/abs/2306.01708)                                                     | `ties`               | ✅          | ✅              |
-| [DARE](https://arxiv.org/abs/2311.03099) [TIES](https://arxiv.org/abs/2306.01708)            | `dare_ties`          | ✅          | ✅              |
-| [DARE](https://arxiv.org/abs/2311.03099) [Task Arithmetic](https://arxiv.org/abs/2212.04089) | `dare_linear`        | ✅          | ✅              |
-| Passthrough                                                                                  | `passthrough`        | ❌          | ❌              |
-| [Model Stock](https://arxiv.org/abs/2403.19522)                                              | `model_stock`        | ✅          | ✅              |
-### Linear
-The classic merge method - a simple weighted average.
-Parameters:
-- `weight` - relative (or absolute if `normalize=False`) weighting of a given tensor
-- `normalize` - if true, the weights of all models contributing to a tensor will be normalized. Default behavior.
-### SLERP
-Spherically interpolate the parameters of two models. One must be set as `base_model`.
-Parameters:
-- `t` - interpolation factor. At `t=0` will return `base_model`, at `t=1` will return the other one.
-### [Task Arithmetic](https://arxiv.org/abs/2212.04089)
-Computes "task vectors" for each model by subtracting a base model. Merges the task vectors linearly and adds back the base. Works great for models that were fine tuned from a common ancestor. Also a super useful mental framework for several of the more involved merge methods.
-Parameters: same as [Linear](#linear)
-### [TIES](https://arxiv.org/abs/2306.01708)
-Builds on the task arithmetic framework. Resolves interference between models by sparsifying the task vectors and applying a sign consensus algorithm. Allows you to merge a larger number of models and retain more of their strengths.
-Parameters: same as [Linear](#linear), plus:
-- `density` - fraction of weights in differences from the base model to retain
-### [DARE](https://arxiv.org/abs/2311.03099)
-In the same vein as TIES, sparsifies task vectors to reduce interference. Differs in that DARE uses random pruning with a novel rescaling to better match performance of the original models. DARE can be used either with the sign consensus algorithm of TIES (`dare_ties`) or without (`dare_linear`).
-Parameters: same as [TIES](#ties) for `dare_ties`, or [Linear](#linear) for `dare_linear`
-### Passthrough
-`passthrough` is a no-op that simply passes input tensors through unmodified. It is meant to be used for layer-stacking type merges where you have only one input model. Useful for frankenmerging.
-### [Model Stock](https://arxiv.org/abs/2403.19522)
-Uses some neat geometric properties of fine tuned models to compute good weights for linear interpolation. Requires at least three models, including a base model.
-Parameters:
-- `filter_wise`: if true, weight calculation will be per-row rather than per-tensor. Not recommended.
-# Citation
-We now have a [paper](https://arxiv.org/abs/2403.13257) you can cite for the MergeKit library:
-```bibtex
-@article{goddard2024arcee,
-  title={Arcee's MergeKit: A Toolkit for Merging Large Language Models},
-  author={Goddard, Charles and Siriwardhana, Shamane and Ehghaghi, Malikeh and Meyers, Luke and Karpukhin, Vlad and Benedict, Brian and McQuade, Mark and Solawetz, Jacob},
-  journal={arXiv preprint arXiv:2403.13257},
-  year={2024}
-}
 ```

+---
+base_model:
+- Shaleen123/phi-2-maths
+- Shaleen123/phi-2-code
+- Shaleen123/phi-2-4bits
+library_name: transformers
+tags:
+- mergekit
+- merge
+---
+# merge
+This is a merge of pre-trained language models created using [mergekit](https://github.com/cg123/mergekit).
+## Merge Details
+### Merge Method
+This model was merged using the [linear](https://arxiv.org/abs/2203.05482) merge method.
+### Models Merged
+The following models were included in the merge:
+* [Shaleen123/phi-2-maths](https://huggingface.co/Shaleen123/phi-2-maths)
+* [Shaleen123/phi-2-code](https://huggingface.co/Shaleen123/phi-2-code)
+* [Shaleen123/phi-2-4bits](https://huggingface.co/Shaleen123/phi-2-4bits)
+### Configuration
+The following YAML configuration was used to produce this model:
+```yaml
+models:
+  - model: Shaleen123/phi-2-code
+    parameters:
+      weight: 0.5
+  - model: Shaleen123/phi-2-maths
+    parameters:
+      weight: 0.3
+  - model: Shaleen123/phi-2-4bits
+    parameters:
+      weight: 1.0
+merge_method: linear
+dtype: float16
 ```

added_tokens.json ADDED Viewed

	@@ -0,0 +1,40 @@

+{
+  "\t\t": 50294,
+  "\t\t\t": 50293,
+  "\t\t\t\t": 50292,
+  "\t\t\t\t\t": 50291,
+  "\t\t\t\t\t\t": 50290,
+  "\t\t\t\t\t\t\t": 50289,
+  "\t\t\t\t\t\t\t\t": 50288,
+  "\t\t\t\t\t\t\t\t\t": 50287,
+  "  ": 50286,
+  "   ": 50285,
+  "    ": 50284,
+  "     ": 50283,
+  "      ": 50282,
+  "       ": 50281,
+  "        ": 50280,
+  "         ": 50279,
+  "          ": 50278,
+  "           ": 50277,
+  "            ": 50276,
+  "             ": 50275,
+  "              ": 50274,
+  "               ": 50273,
+  "                ": 50272,
+  "                 ": 50271,
+  "                  ": 50270,
+  "                   ": 50269,
+  "                    ": 50268,
+  "                     ": 50267,
+  "                      ": 50266,
+  "                       ": 50265,
+  "                        ": 50264,
+  "                         ": 50263,
+  "                          ": 50262,
+  "                           ": 50261,
+  "                            ": 50260,
+  "                             ": 50259,
+  "                              ": 50258,
+  "                               ": 50257
+}

config.json ADDED Viewed

	@@ -0,0 +1,48 @@

+{
+  "_name_or_path": "Shaleen123/phi-2-maths",
+  "architectures": [
+    "PhiForCausalLM"
+  ],
+  "attention_dropout": 0.0,
+  "auto_map": {
+    "AutoConfig": "microsoft/phi-2--configuration_phi.PhiConfig",
+    "AutoModelForCausalLM": "microsoft/phi-2--modeling_phi.PhiForCausalLM"
+  },
+  "bos_token_id": 50256,
+  "embd_pdrop": 0.0,
+  "eos_token_id": 50256,
+  "hidden_act": "gelu_new",
+  "hidden_size": 2560,
+  "initializer_range": 0.02,
+  "intermediate_size": 10240,
+  "layer_norm_eps": 1e-05,
+  "max_position_embeddings": 2048,
+  "model_type": "phi",
+  "num_attention_heads": 32,
+  "num_hidden_layers": 32,
+  "num_key_value_heads": 32,
+  "partial_rotary_factor": 0.4,
+  "qk_layernorm": false,
+  "quantization_config": {
+    "_load_in_4bit": true,
+    "_load_in_8bit": false,
+    "bnb_4bit_compute_dtype": "float32",
+    "bnb_4bit_quant_type": "fp4",
+    "bnb_4bit_use_double_quant": false,
+    "llm_int8_enable_fp32_cpu_offload": false,
+    "llm_int8_has_fp16_weight": false,
+    "llm_int8_skip_modules": null,
+    "llm_int8_threshold": 6.0,
+    "load_in_4bit": true,
+    "load_in_8bit": false,
+    "quant_method": "bitsandbytes"
+  },
+  "resid_pdrop": 0.1,
+  "rope_scaling": null,
+  "rope_theta": 10000.0,
+  "tie_word_embeddings": false,
+  "torch_dtype": "float16",
+  "transformers_version": "4.38.2",
+  "use_cache": true,
+  "vocab_size": 51200
+}

mergekit_config.yml ADDED Viewed

	@@ -0,0 +1,13 @@

+models:
+  - model: Shaleen123/phi-2-code
+    parameters:
+      weight: 0.5
+  - model: Shaleen123/phi-2-maths
+    parameters:
+      weight: 0.3
+  - model: Shaleen123/phi-2-4bits
+    parameters:
+      weight: 1.0
+merge_method: linear
+dtype: float16

merges.txt ADDED Viewed

The diff for this file is too large to render. See raw diff

model-00001-of-00002.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:e6a858e3c69e1ba3c09c6d49a7dee2460ccd4c13b599704e608c88df6fa46c64
+size 1993680248

model-00002-of-00002.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:6a42859d88e8973ad35321d724fc93498b3b74ce153db25b5e975d0c5eeb1395
+size 1049154408

model.safetensors.index.json ADDED Viewed

	@@ -0,0 +1 @@

+ {"metadata": {"mergekit_version": "0.0.4.2", "total_size": 3042785280}, "weight_map": {"model.final_layernorm.weight": "model-00001-of-00002.safetensors", "model.final_layernorm.bias": "model-00001-of-00002.safetensors", "lm_head.weight": "model-00001-of-00002.safetensors", "lm_head.bias": "model-00001-of-00002.safetensors", "model.layers.31.mlp.fc2.weight": "model-00001-of-00002.safetensors", "model.layers.31.mlp.fc2.bias": "model-00001-of-00002.safetensors", "model.layers.31.mlp.fc1.weight": "model-00001-of-00002.safetensors", "model.layers.31.mlp.fc1.bias": "model-00001-of-00002.safetensors", "model.layers.31.self_attn.v_proj.weight": "model-00001-of-00002.safetensors", "model.layers.31.self_attn.v_proj.bias": "model-00001-of-00002.safetensors", "model.layers.31.self_attn.k_proj.weight": "model-00001-of-00002.safetensors", "model.layers.31.self_attn.k_proj.bias": "model-00001-of-00002.safetensors", "model.layers.31.self_attn.q_proj.weight": "model-00001-of-00002.safetensors", "model.layers.31.self_attn.q_proj.bias": "model-00001-of-00002.safetensors", "model.layers.31.self_attn.dense.weight": "model-00001-of-00002.safetensors", "model.layers.31.self_attn.dense.bias": "model-00001-of-00002.safetensors", "model.layers.31.input_layernorm.weight": "model-00001-of-00002.safetensors", "model.layers.31.input_layernorm.bias": "model-00001-of-00002.safetensors", "model.layers.30.mlp.fc2.weight": "model-00001-of-00002.safetensors", "model.layers.30.mlp.fc2.bias": "model-00001-of-00002.safetensors", "model.layers.30.mlp.fc1.weight": "model-00001-of-00002.safetensors", "model.layers.30.mlp.fc1.bias": "model-00001-of-00002.safetensors", "model.layers.30.self_attn.v_proj.weight": "model-00001-of-00002.safetensors", "model.layers.30.self_attn.v_proj.bias": "model-00001-of-00002.safetensors", "model.layers.30.self_attn.k_proj.weight": "model-00001-of-00002.safetensors", "model.layers.30.self_attn.k_proj.bias": "model-00001-of-00002.safetensors", "model.layers.30.self_attn.q_proj.weight": "model-00001-of-00002.safetensors", "model.layers.30.self_attn.q_proj.bias": "model-00001-of-00002.safetensors", "model.layers.30.self_attn.dense.weight": "model-00001-of-00002.safetensors", "model.layers.30.self_attn.dense.bias": "model-00001-of-00002.safetensors", "model.layers.30.input_layernorm.weight": "model-00001-of-00002.safetensors", "model.layers.30.input_layernorm.bias": "model-00001-of-00002.safetensors", "model.layers.29.mlp.fc2.weight": "model-00001-of-00002.safetensors", "model.layers.29.mlp.fc2.bias": "model-00001-of-00002.safetensors", "model.layers.29.mlp.fc1.weight": "model-00001-of-00002.safetensors", "model.layers.29.mlp.fc1.bias": "model-00001-of-00002.safetensors", "model.layers.29.self_attn.v_proj.weight": "model-00001-of-00002.safetensors", "model.layers.29.self_attn.v_proj.bias": "model-00001-of-00002.safetensors", "model.layers.29.self_attn.k_proj.weight": "model-00001-of-00002.safetensors", "model.layers.29.self_attn.k_proj.bias": "model-00001-of-00002.safetensors", "model.layers.29.self_attn.q_proj.weight": "model-00001-of-00002.safetensors", "model.layers.29.self_attn.q_proj.bias": "model-00001-of-00002.safetensors", "model.layers.29.self_attn.dense.weight": "model-00001-of-00002.safetensors", "model.layers.29.self_attn.dense.bias": "model-00001-of-00002.safetensors", "model.layers.29.input_layernorm.weight": "model-00001-of-00002.safetensors", "model.layers.29.input_layernorm.bias": "model-00001-of-00002.safetensors", "model.layers.28.mlp.fc2.weight": "model-00001-of-00002.safetensors", "model.layers.28.mlp.fc2.bias": "model-00001-of-00002.safetensors", "model.layers.28.mlp.fc1.weight": "model-00001-of-00002.safetensors", "model.layers.28.mlp.fc1.bias": "model-00001-of-00002.safetensors", "model.layers.28.self_attn.v_proj.weight": "model-00001-of-00002.safetensors", "model.layers.28.self_attn.v_proj.bias": "model-00001-of-00002.safetensors", "model.layers.28.self_attn.k_proj.weight": "model-00001-of-00002.safetensors", "model.layers.28.self_attn.k_proj.bias": "model-00001-of-00002.safetensors", "model.layers.28.self_attn.q_proj.weight": "model-00001-of-00002.safetensors", "model.layers.28.self_attn.q_proj.bias": "model-00001-of-00002.safetensors", "model.layers.28.self_attn.dense.weight": "model-00001-of-00002.safetensors", "model.layers.28.self_attn.dense.bias": "model-00001-of-00002.safetensors", "model.layers.28.input_layernorm.weight": "model-00001-of-00002.safetensors", "model.layers.28.input_layernorm.bias": "model-00001-of-00002.safetensors", "model.layers.27.mlp.fc2.weight": "model-00001-of-00002.safetensors", "model.layers.27.mlp.fc2.bias": "model-00001-of-00002.safetensors", "model.layers.27.mlp.fc1.weight": "model-00001-of-00002.safetensors", "model.layers.27.mlp.fc1.bias": "model-00001-of-00002.safetensors", "model.layers.27.self_attn.v_proj.weight": "model-00001-of-00002.safetensors", "model.layers.27.self_attn.v_proj.bias": "model-00001-of-00002.safetensors", "model.layers.27.self_attn.k_proj.weight": "model-00001-of-00002.safetensors", "model.layers.27.self_attn.k_proj.bias": "model-00001-of-00002.safetensors", "model.layers.27.self_attn.q_proj.weight": "model-00001-of-00002.safetensors", "model.layers.27.self_attn.q_proj.bias": "model-00001-of-00002.safetensors", "model.layers.27.self_attn.dense.weight": "model-00001-of-00002.safetensors", "model.layers.27.self_attn.dense.bias": "model-00001-of-00002.safetensors", "model.layers.27.input_layernorm.weight": "model-00001-of-00002.safetensors", "model.layers.27.input_layernorm.bias": "model-00001-of-00002.safetensors", "model.layers.26.mlp.fc2.weight": "model-00001-of-00002.safetensors", "model.layers.26.mlp.fc2.bias": "model-00001-of-00002.safetensors", "model.layers.26.mlp.fc1.weight": "model-00001-of-00002.safetensors", "model.layers.26.mlp.fc1.bias": "model-00001-of-00002.safetensors", "model.layers.26.self_attn.v_proj.weight": "model-00001-of-00002.safetensors", "model.layers.26.self_attn.v_proj.bias": "model-00001-of-00002.safetensors", "model.layers.26.self_attn.k_proj.weight": "model-00001-of-00002.safetensors", "model.layers.26.self_attn.k_proj.bias": "model-00001-of-00002.safetensors", "model.layers.26.self_attn.q_proj.weight": "model-00001-of-00002.safetensors", "model.layers.26.self_attn.q_proj.bias": "model-00001-of-00002.safetensors", "model.layers.26.self_attn.dense.weight": "model-00001-of-00002.safetensors", "model.layers.26.self_attn.dense.bias": "model-00001-of-00002.safetensors", "model.layers.26.input_layernorm.weight": "model-00001-of-00002.safetensors", "model.layers.26.input_layernorm.bias": "model-00001-of-00002.safetensors", "model.layers.25.mlp.fc2.weight": "model-00001-of-00002.safetensors", "model.layers.25.mlp.fc2.bias": "model-00001-of-00002.safetensors", "model.layers.25.mlp.fc1.weight": "model-00001-of-00002.safetensors", "model.layers.25.mlp.fc1.bias": "model-00001-of-00002.safetensors", "model.layers.25.self_attn.v_proj.weight": "model-00001-of-00002.safetensors", "model.layers.25.self_attn.v_proj.bias": "model-00001-of-00002.safetensors", "model.layers.25.self_attn.k_proj.weight": "model-00001-of-00002.safetensors", "model.layers.25.self_attn.k_proj.bias": "model-00001-of-00002.safetensors", "model.layers.25.self_attn.q_proj.weight": "model-00001-of-00002.safetensors", "model.layers.25.self_attn.q_proj.bias": "model-00001-of-00002.safetensors", "model.layers.25.self_attn.dense.weight": "model-00001-of-00002.safetensors", "model.layers.25.self_attn.dense.bias": "model-00001-of-00002.safetensors", "model.layers.25.input_layernorm.weight": "model-00001-of-00002.safetensors", "model.layers.25.input_layernorm.bias": "model-00001-of-00002.safetensors", "model.layers.24.mlp.fc2.weight": "model-00001-of-00002.safetensors", "model.layers.24.mlp.fc2.bias": "model-00001-of-00002.safetensors", "model.layers.24.mlp.fc1.weight": "model-00001-of-00002.safetensors", "model.layers.24.mlp.fc1.bias": "model-00001-of-00002.safetensors", "model.layers.24.self_attn.v_proj.weight": "model-00001-of-00002.safetensors", "model.layers.24.self_attn.v_proj.bias": "model-00001-of-00002.safetensors", "model.layers.24.self_attn.k_proj.weight": "model-00001-of-00002.safetensors", "model.layers.24.self_attn.k_proj.bias": "model-00001-of-00002.safetensors", "model.layers.24.self_attn.q_proj.weight": "model-00001-of-00002.safetensors", "model.layers.24.self_attn.q_proj.bias": "model-00001-of-00002.safetensors", "model.layers.24.self_attn.dense.weight": "model-00001-of-00002.safetensors", "model.layers.24.self_attn.dense.bias": "model-00001-of-00002.safetensors", "model.layers.24.input_layernorm.weight": "model-00001-of-00002.safetensors", "model.layers.24.input_layernorm.bias": "model-00001-of-00002.safetensors", "model.layers.23.mlp.fc2.weight": "model-00001-of-00002.safetensors", "model.layers.23.mlp.fc2.bias": "model-00001-of-00002.safetensors", "model.layers.23.mlp.fc1.weight": "model-00001-of-00002.safetensors", "model.layers.23.mlp.fc1.bias": "model-00001-of-00002.safetensors", "model.layers.23.self_attn.v_proj.weight": "model-00001-of-00002.safetensors", "model.layers.23.self_attn.v_proj.bias": "model-00001-of-00002.safetensors", "model.layers.23.self_attn.k_proj.weight": "model-00001-of-00002.safetensors", "model.layers.23.self_attn.k_proj.bias": "model-00001-of-00002.safetensors", "model.layers.23.self_attn.q_proj.weight": "model-00001-of-00002.safetensors", "model.layers.23.self_attn.q_proj.bias": "model-00001-of-00002.safetensors", "model.layers.23.self_attn.dense.weight": "model-00001-of-00002.safetensors", "model.layers.23.self_attn.dense.bias": "model-00001-of-00002.safetensors", "model.layers.23.input_layernorm.weight": "model-00001-of-00002.safetensors", "model.layers.23.input_layernorm.bias": "model-00001-of-00002.safetensors", "model.layers.22.mlp.fc2.weight": "model-00001-of-00002.safetensors", "model.layers.22.mlp.fc2.bias": "model-00001-of-00002.safetensors", "model.layers.22.mlp.fc1.weight": "model-00001-of-00002.safetensors", "model.layers.22.mlp.fc1.bias": "model-00001-of-00002.safetensors", "model.layers.22.self_attn.v_proj.weight": "model-00001-of-00002.safetensors", "model.layers.22.self_attn.v_proj.bias": "model-00001-of-00002.safetensors", "model.layers.22.self_attn.k_proj.weight": "model-00001-of-00002.safetensors", "model.layers.22.self_attn.k_proj.bias": "model-00001-of-00002.safetensors", "model.layers.22.self_attn.q_proj.weight": "model-00001-of-00002.safetensors", "model.layers.22.self_attn.q_proj.bias": "model-00001-of-00002.safetensors", "model.layers.22.self_attn.dense.weight": "model-00001-of-00002.safetensors", "model.layers.22.self_attn.dense.bias": "model-00001-of-00002.safetensors", "model.layers.22.input_layernorm.weight": "model-00001-of-00002.safetensors", "model.layers.22.input_layernorm.bias": "model-00001-of-00002.safetensors", "model.layers.21.mlp.fc2.weight": "model-00001-of-00002.safetensors", "model.layers.21.mlp.fc2.bias": "model-00001-of-00002.safetensors", "model.layers.21.mlp.fc1.weight": "model-00001-of-00002.safetensors", "model.layers.21.mlp.fc1.bias": "model-00001-of-00002.safetensors", "model.layers.21.self_attn.v_proj.weight": "model-00001-of-00002.safetensors", "model.layers.21.self_attn.v_proj.bias": "model-00001-of-00002.safetensors", "model.layers.21.self_attn.k_proj.weight": "model-00001-of-00002.safetensors", "model.layers.21.self_attn.k_proj.bias": "model-00001-of-00002.safetensors", "model.layers.21.self_attn.q_proj.weight": "model-00001-of-00002.safetensors", "model.layers.21.self_attn.q_proj.bias": "model-00001-of-00002.safetensors", "model.layers.21.self_attn.dense.weight": "model-00001-of-00002.safetensors", "model.layers.21.self_attn.dense.bias": "model-00001-of-00002.safetensors", "model.layers.21.input_layernorm.weight": "model-00001-of-00002.safetensors", "model.layers.21.input_layernorm.bias": "model-00001-of-00002.safetensors", "model.layers.20.mlp.fc2.weight": "model-00001-of-00002.safetensors", "model.layers.20.mlp.fc2.bias": "model-00001-of-00002.safetensors", "model.layers.20.mlp.fc1.weight": "model-00001-of-00002.safetensors", "model.layers.20.mlp.fc1.bias": "model-00001-of-00002.safetensors", "model.layers.20.self_attn.v_proj.weight": "model-00001-of-00002.safetensors", "model.layers.20.self_attn.v_proj.bias": "model-00001-of-00002.safetensors", "model.layers.20.self_attn.k_proj.weight": "model-00001-of-00002.safetensors", "model.layers.20.self_attn.k_proj.bias": "model-00001-of-00002.safetensors", "model.layers.20.self_attn.q_proj.weight": "model-00001-of-00002.safetensors", "model.layers.20.self_attn.q_proj.bias": "model-00001-of-00002.safetensors", "model.layers.20.self_attn.dense.weight": "model-00001-of-00002.safetensors", "model.layers.20.self_attn.dense.bias": "model-00001-of-00002.safetensors", "model.layers.20.input_layernorm.weight": "model-00001-of-00002.safetensors", "model.layers.20.input_layernorm.bias": "model-00001-of-00002.safetensors", "model.layers.19.mlp.fc2.weight": "model-00001-of-00002.safetensors", "model.layers.19.mlp.fc2.bias": "model-00001-of-00002.safetensors", "model.layers.19.mlp.fc1.weight": "model-00001-of-00002.safetensors", "model.layers.19.mlp.fc1.bias": "model-00001-of-00002.safetensors", "model.layers.19.self_attn.v_proj.weight": "model-00001-of-00002.safetensors", "model.layers.19.self_attn.v_proj.bias": "model-00001-of-00002.safetensors", "model.layers.19.self_attn.k_proj.weight": "model-00001-of-00002.safetensors", "model.layers.19.self_attn.k_proj.bias": "model-00001-of-00002.safetensors", "model.layers.19.self_attn.q_proj.weight": "model-00001-of-00002.safetensors", "model.layers.19.self_attn.q_proj.bias": "model-00001-of-00002.safetensors", "model.layers.19.self_attn.dense.weight": "model-00001-of-00002.safetensors", "model.layers.19.self_attn.dense.bias": "model-00001-of-00002.safetensors", "model.layers.19.input_layernorm.weight": "model-00001-of-00002.safetensors", "model.layers.19.input_layernorm.bias": "model-00001-of-00002.safetensors", "model.layers.18.mlp.fc2.weight": "model-00001-of-00002.safetensors", "model.layers.18.mlp.fc2.bias": "model-00001-of-00002.safetensors", "model.layers.18.mlp.fc1.weight": "model-00001-of-00002.safetensors", "model.layers.18.mlp.fc1.bias": "model-00001-of-00002.safetensors", "model.layers.18.self_attn.v_proj.weight": "model-00001-of-00002.safetensors", "model.layers.18.self_attn.v_proj.bias": "model-00001-of-00002.safetensors", "model.layers.18.self_attn.k_proj.weight": "model-00001-of-00002.safetensors", "model.layers.18.self_attn.k_proj.bias": "model-00001-of-00002.safetensors", "model.layers.18.self_attn.q_proj.weight": "model-00001-of-00002.safetensors", "model.layers.18.self_attn.q_proj.bias": "model-00001-of-00002.safetensors", "model.layers.18.self_attn.dense.weight": "model-00001-of-00002.safetensors", "model.layers.18.self_attn.dense.bias": "model-00001-of-00002.safetensors", "model.layers.18.input_layernorm.weight": "model-00001-of-00002.safetensors", "model.layers.18.input_layernorm.bias": "model-00001-of-00002.safetensors", "model.layers.17.mlp.fc2.weight": "model-00001-of-00002.safetensors", "model.layers.17.mlp.fc2.bias": "model-00001-of-00002.safetensors", "model.layers.17.mlp.fc1.weight": "model-00001-of-00002.safetensors", "model.layers.17.mlp.fc1.bias": "model-00001-of-00002.safetensors", "model.layers.17.self_attn.v_proj.weight": "model-00001-of-00002.safetensors", "model.layers.17.self_attn.v_proj.bias": "model-00001-of-00002.safetensors", "model.layers.17.self_attn.k_proj.weight": "model-00001-of-00002.safetensors", "model.layers.17.self_attn.k_proj.bias": "model-00001-of-00002.safetensors", "model.layers.17.self_attn.q_proj.weight": "model-00001-of-00002.safetensors", "model.layers.17.self_attn.q_proj.bias": "model-00001-of-00002.safetensors", "model.layers.17.self_attn.dense.weight": "model-00001-of-00002.safetensors", "model.layers.17.self_attn.dense.bias": "model-00001-of-00002.safetensors", "model.layers.17.input_layernorm.weight": "model-00001-of-00002.safetensors", "model.layers.17.input_layernorm.bias": "model-00001-of-00002.safetensors", "model.layers.16.mlp.fc2.weight": "model-00001-of-00002.safetensors", "model.layers.16.mlp.fc2.bias": "model-00001-of-00002.safetensors", "model.layers.16.mlp.fc1.weight": "model-00001-of-00002.safetensors", "model.layers.16.mlp.fc1.bias": "model-00001-of-00002.safetensors", "model.layers.16.self_attn.v_proj.weight": "model-00001-of-00002.safetensors", "model.layers.16.self_attn.v_proj.bias": "model-00001-of-00002.safetensors", "model.layers.16.self_attn.k_proj.weight": "model-00001-of-00002.safetensors", "model.layers.16.self_attn.k_proj.bias": "model-00001-of-00002.safetensors", "model.layers.16.self_attn.q_proj.weight": "model-00001-of-00002.safetensors", "model.layers.16.self_attn.q_proj.bias": "model-00001-of-00002.safetensors", "model.layers.16.self_attn.dense.weight": "model-00001-of-00002.safetensors", "model.layers.16.self_attn.dense.bias": "model-00001-of-00002.safetensors", "model.layers.16.input_layernorm.weight": "model-00001-of-00002.safetensors", "model.layers.16.input_layernorm.bias": "model-00001-of-00002.safetensors", "model.layers.15.mlp.fc2.weight": "model-00001-of-00002.safetensors", "model.layers.15.mlp.fc2.bias": "model-00001-of-00002.safetensors", "model.layers.15.mlp.fc1.weight": "model-00001-of-00002.safetensors", "model.layers.15.mlp.fc1.bias": "model-00001-of-00002.safetensors", "model.layers.15.self_attn.v_proj.weight": "model-00001-of-00002.safetensors", "model.layers.15.self_attn.v_proj.bias": "model-00001-of-00002.safetensors", "model.layers.15.self_attn.k_proj.weight": "model-00001-of-00002.safetensors", "model.layers.15.self_attn.k_proj.bias": "model-00001-of-00002.safetensors", "model.layers.15.self_attn.q_proj.weight": "model-00001-of-00002.safetensors", "model.layers.15.self_attn.q_proj.bias": "model-00001-of-00002.safetensors", "model.layers.15.self_attn.dense.weight": "model-00001-of-00002.safetensors", "model.layers.15.self_attn.dense.bias": "model-00001-of-00002.safetensors", "model.layers.15.input_layernorm.weight": "model-00001-of-00002.safetensors", "model.layers.15.input_layernorm.bias": "model-00001-of-00002.safetensors", "model.layers.14.mlp.fc2.weight": "model-00001-of-00002.safetensors", "model.layers.14.mlp.fc2.bias": "model-00001-of-00002.safetensors", "model.layers.14.mlp.fc1.weight": "model-00001-of-00002.safetensors", "model.layers.14.mlp.fc1.bias": "model-00001-of-00002.safetensors", "model.layers.14.self_attn.v_proj.weight": "model-00001-of-00002.safetensors", "model.layers.14.self_attn.v_proj.bias": "model-00001-of-00002.safetensors", "model.layers.14.self_attn.k_proj.weight": "model-00001-of-00002.safetensors", "model.layers.14.self_attn.k_proj.bias": "model-00001-of-00002.safetensors", "model.layers.14.self_attn.q_proj.weight": "model-00001-of-00002.safetensors", "model.layers.14.self_attn.q_proj.bias": "model-00001-of-00002.safetensors", "model.layers.14.self_attn.dense.weight": "model-00001-of-00002.safetensors", "model.layers.14.self_attn.dense.bias": "model-00001-of-00002.safetensors", "model.layers.14.input_layernorm.weight": "model-00001-of-00002.safetensors", "model.layers.14.input_layernorm.bias": "model-00001-of-00002.safetensors", "model.layers.13.mlp.fc2.weight": "model-00001-of-00002.safetensors", "model.layers.13.mlp.fc2.bias": "model-00001-of-00002.safetensors", "model.layers.13.mlp.fc1.weight": "model-00001-of-00002.safetensors", "model.layers.13.mlp.fc1.bias": "model-00001-of-00002.safetensors", "model.layers.13.self_attn.v_proj.weight": "model-00001-of-00002.safetensors", "model.layers.13.self_attn.v_proj.bias": "model-00001-of-00002.safetensors", "model.layers.13.self_attn.k_proj.weight": "model-00001-of-00002.safetensors", "model.layers.13.self_attn.k_proj.bias": "model-00001-of-00002.safetensors", "model.layers.13.self_attn.q_proj.weight": "model-00001-of-00002.safetensors", "model.layers.13.self_attn.q_proj.bias": "model-00001-of-00002.safetensors", "model.layers.13.self_attn.dense.weight": "model-00001-of-00002.safetensors", "model.layers.13.self_attn.dense.bias": "model-00001-of-00002.safetensors", "model.layers.13.input_layernorm.weight": "model-00001-of-00002.safetensors", "model.layers.13.input_layernorm.bias": "model-00001-of-00002.safetensors", "model.layers.12.mlp.fc2.weight": "model-00001-of-00002.safetensors", "model.layers.12.mlp.fc2.bias": "model-00001-of-00002.safetensors", "model.layers.12.mlp.fc1.weight": "model-00001-of-00002.safetensors", "model.layers.12.mlp.fc1.bias": "model-00001-of-00002.safetensors", "model.layers.12.self_attn.v_proj.weight": "model-00001-of-00002.safetensors", "model.layers.12.self_attn.v_proj.bias": "model-00001-of-00002.safetensors", "model.layers.12.self_attn.k_proj.weight": "model-00001-of-00002.safetensors", "model.layers.12.self_attn.k_proj.bias": "model-00001-of-00002.safetensors", "model.layers.12.self_attn.q_proj.weight": "model-00001-of-00002.safetensors", "model.layers.12.self_attn.q_proj.bias": "model-00001-of-00002.safetensors", "model.layers.12.self_attn.dense.weight": "model-00001-of-00002.safetensors", "model.layers.12.self_attn.dense.bias": "model-00001-of-00002.safetensors", "model.layers.12.input_layernorm.weight": "model-00001-of-00002.safetensors", "model.layers.12.input_layernorm.bias": "model-00001-of-00002.safetensors", "model.layers.11.mlp.fc2.weight": "model-00001-of-00002.safetensors", "model.layers.11.mlp.fc2.bias": "model-00001-of-00002.safetensors", "model.layers.11.mlp.fc1.weight": "model-00001-of-00002.safetensors", "model.layers.11.mlp.fc1.bias": "model-00001-of-00002.safetensors", "model.layers.11.self_attn.v_proj.weight": "model-00001-of-00002.safetensors", "model.layers.11.self_attn.v_proj.bias": "model-00001-of-00002.safetensors", "model.layers.11.self_attn.k_proj.weight": "model-00001-of-00002.safetensors", "model.layers.11.self_attn.k_proj.bias": "model-00001-of-00002.safetensors", "model.layers.11.self_attn.q_proj.weight": "model-00001-of-00002.safetensors", "model.layers.11.self_attn.q_proj.bias": "model-00001-of-00002.safetensors", "model.layers.11.self_attn.dense.weight": "model-00001-of-00002.safetensors", "model.layers.11.self_attn.dense.bias": "model-00001-of-00002.safetensors", "model.layers.11.input_layernorm.weight": "model-00001-of-00002.safetensors", "model.layers.11.input_layernorm.bias": "model-00001-of-00002.safetensors", "model.layers.10.mlp.fc2.weight": "model-00001-of-00002.safetensors", "model.layers.10.mlp.fc2.bias": "model-00001-of-00002.safetensors", "model.layers.10.mlp.fc1.weight": "model-00001-of-00002.safetensors", "model.layers.10.mlp.fc1.bias": "model-00001-of-00002.safetensors", "model.layers.10.self_attn.v_proj.weight": "model-00001-of-00002.safetensors", "model.layers.10.self_attn.v_proj.bias": "model-00001-of-00002.safetensors", "model.layers.10.self_attn.k_proj.weight": "model-00001-of-00002.safetensors", "model.layers.10.self_attn.k_proj.bias": "model-00001-of-00002.safetensors", "model.layers.10.self_attn.q_proj.weight": "model-00001-of-00002.safetensors", "model.layers.10.self_attn.q_proj.bias": "model-00001-of-00002.safetensors", "model.layers.10.self_attn.dense.weight": "model-00001-of-00002.safetensors", "model.layers.10.self_attn.dense.bias": "model-00001-of-00002.safetensors", "model.layers.10.input_layernorm.weight": "model-00001-of-00002.safetensors", "model.layers.10.input_layernorm.bias": "model-00001-of-00002.safetensors", "model.layers.9.mlp.fc2.weight": "model-00002-of-00002.safetensors", "model.layers.9.mlp.fc2.bias": "model-00002-of-00002.safetensors", "model.layers.9.mlp.fc1.weight": "model-00002-of-00002.safetensors", "model.layers.9.mlp.fc1.bias": "model-00002-of-00002.safetensors", "model.layers.9.self_attn.v_proj.weight": "model-00002-of-00002.safetensors", "model.layers.9.self_attn.v_proj.bias": "model-00002-of-00002.safetensors", "model.layers.9.self_attn.k_proj.weight": "model-00002-of-00002.safetensors", "model.layers.9.self_attn.k_proj.bias": "model-00002-of-00002.safetensors", "model.layers.9.self_attn.q_proj.weight": "model-00002-of-00002.safetensors", "model.layers.9.self_attn.q_proj.bias": "model-00002-of-00002.safetensors", "model.layers.9.self_attn.dense.weight": "model-00002-of-00002.safetensors", "model.layers.9.self_attn.dense.bias": "model-00002-of-00002.safetensors", "model.layers.9.input_layernorm.weight": "model-00002-of-00002.safetensors", "model.layers.9.input_layernorm.bias": "model-00002-of-00002.safetensors", "model.layers.8.mlp.fc2.weight": "model-00002-of-00002.safetensors", "model.layers.8.mlp.fc2.bias": "model-00002-of-00002.safetensors", "model.layers.8.mlp.fc1.weight": "model-00002-of-00002.safetensors", "model.layers.8.mlp.fc1.bias": "model-00002-of-00002.safetensors", "model.layers.8.self_attn.v_proj.weight": "model-00002-of-00002.safetensors", "model.layers.8.self_attn.v_proj.bias": "model-00002-of-00002.safetensors", "model.layers.8.self_attn.k_proj.weight": "model-00002-of-00002.safetensors", "model.layers.8.self_attn.k_proj.bias": "model-00002-of-00002.safetensors", "model.layers.8.self_attn.q_proj.weight": "model-00002-of-00002.safetensors", "model.layers.8.self_attn.q_proj.bias": "model-00002-of-00002.safetensors", "model.layers.8.self_attn.dense.weight": "model-00002-of-00002.safetensors", "model.layers.8.self_attn.dense.bias": "model-00002-of-00002.safetensors", "model.layers.8.input_layernorm.weight": "model-00002-of-00002.safetensors", "model.layers.8.input_layernorm.bias": "model-00002-of-00002.safetensors", "model.layers.7.mlp.fc2.weight": "model-00002-of-00002.safetensors", "model.layers.7.mlp.fc2.bias": "model-00002-of-00002.safetensors", "model.layers.7.mlp.fc1.weight": "model-00002-of-00002.safetensors", "model.layers.7.mlp.fc1.bias": "model-00002-of-00002.safetensors", "model.layers.7.self_attn.v_proj.weight": "model-00002-of-00002.safetensors", "model.layers.7.self_attn.v_proj.bias": "model-00002-of-00002.safetensors", "model.layers.7.self_attn.k_proj.weight": "model-00002-of-00002.safetensors", "model.layers.7.self_attn.k_proj.bias": "model-00002-of-00002.safetensors", "model.layers.7.self_attn.q_proj.weight": "model-00002-of-00002.safetensors", "model.layers.7.self_attn.q_proj.bias": "model-00002-of-00002.safetensors", "model.layers.7.self_attn.dense.weight": "model-00002-of-00002.safetensors", "model.layers.7.self_attn.dense.bias": "model-00002-of-00002.safetensors", "model.layers.7.input_layernorm.weight": "model-00002-of-00002.safetensors", "model.layers.7.input_layernorm.bias": "model-00002-of-00002.safetensors", "model.layers.6.mlp.fc2.weight": "model-00002-of-00002.safetensors", "model.layers.6.mlp.fc2.bias": "model-00002-of-00002.safetensors", "model.layers.6.mlp.fc1.weight": "model-00002-of-00002.safetensors", "model.layers.6.mlp.fc1.bias": "model-00002-of-00002.safetensors", "model.layers.6.self_attn.v_proj.weight": "model-00002-of-00002.safetensors", "model.layers.6.self_attn.v_proj.bias": "model-00002-of-00002.safetensors", "model.layers.6.self_attn.k_proj.weight": "model-00002-of-00002.safetensors", "model.layers.6.self_attn.k_proj.bias": "model-00002-of-00002.safetensors", "model.layers.6.self_attn.q_proj.weight": "model-00002-of-00002.safetensors", "model.layers.6.self_attn.q_proj.bias": "model-00002-of-00002.safetensors", "model.layers.6.self_attn.dense.weight": "model-00002-of-00002.safetensors", "model.layers.6.self_attn.dense.bias": "model-00002-of-00002.safetensors", "model.layers.6.input_layernorm.weight": "model-00002-of-00002.safetensors", "model.layers.6.input_layernorm.bias": "model-00002-of-00002.safetensors", "model.layers.5.mlp.fc2.weight": "model-00002-of-00002.safetensors", "model.layers.5.mlp.fc2.bias": "model-00002-of-00002.safetensors", "model.layers.5.mlp.fc1.weight": "model-00002-of-00002.safetensors", "model.layers.5.mlp.fc1.bias": "model-00002-of-00002.safetensors", "model.layers.5.self_attn.v_proj.weight": "model-00002-of-00002.safetensors", "model.layers.5.self_attn.v_proj.bias": "model-00002-of-00002.safetensors", "model.layers.5.self_attn.k_proj.weight": "model-00002-of-00002.safetensors", "model.layers.5.self_attn.k_proj.bias": "model-00002-of-00002.safetensors", "model.layers.5.self_attn.q_proj.weight": "model-00002-of-00002.safetensors", "model.layers.5.self_attn.q_proj.bias": "model-00002-of-00002.safetensors", "model.layers.5.self_attn.dense.weight": "model-00002-of-00002.safetensors", "model.layers.5.self_attn.dense.bias": "model-00002-of-00002.safetensors", "model.layers.5.input_layernorm.weight": "model-00002-of-00002.safetensors", "model.layers.5.input_layernorm.bias": "model-00002-of-00002.safetensors", "model.layers.4.mlp.fc2.weight": "model-00002-of-00002.safetensors", "model.layers.4.mlp.fc2.bias": "model-00002-of-00002.safetensors", "model.layers.4.mlp.fc1.weight": "model-00002-of-00002.safetensors", "model.layers.4.mlp.fc1.bias": "model-00002-of-00002.safetensors", "model.layers.4.self_attn.v_proj.weight": "model-00002-of-00002.safetensors", "model.layers.4.self_attn.v_proj.bias": "model-00002-of-00002.safetensors", "model.layers.4.self_attn.k_proj.weight": "model-00002-of-00002.safetensors", "model.layers.4.self_attn.k_proj.bias": "model-00002-of-00002.safetensors", "model.layers.4.self_attn.q_proj.weight": "model-00002-of-00002.safetensors", "model.layers.4.self_attn.q_proj.bias": "model-00002-of-00002.safetensors", "model.layers.4.self_attn.dense.weight": "model-00002-of-00002.safetensors", "model.layers.4.self_attn.dense.bias": "model-00002-of-00002.safetensors", "model.layers.4.input_layernorm.weight": "model-00002-of-00002.safetensors", "model.layers.4.input_layernorm.bias": "model-00002-of-00002.safetensors", "model.layers.3.mlp.fc2.weight": "model-00002-of-00002.safetensors", "model.layers.3.mlp.fc2.bias": "model-00002-of-00002.safetensors", "model.layers.3.mlp.fc1.weight": "model-00002-of-00002.safetensors", "model.layers.3.mlp.fc1.bias": "model-00002-of-00002.safetensors", "model.layers.3.self_attn.v_proj.weight": "model-00002-of-00002.safetensors", "model.layers.3.self_attn.v_proj.bias": "model-00002-of-00002.safetensors", "model.layers.3.self_attn.k_proj.weight": "model-00002-of-00002.safetensors", "model.layers.3.self_attn.k_proj.bias": "model-00002-of-00002.safetensors", "model.layers.3.self_attn.q_proj.weight": "model-00002-of-00002.safetensors", "model.layers.3.self_attn.q_proj.bias": "model-00002-of-00002.safetensors", "model.layers.3.self_attn.dense.weight": "model-00002-of-00002.safetensors", "model.layers.3.self_attn.dense.bias": "model-00002-of-00002.safetensors", "model.layers.3.input_layernorm.weight": "model-00002-of-00002.safetensors", "model.layers.3.input_layernorm.bias": "model-00002-of-00002.safetensors", "model.layers.2.mlp.fc2.weight": "model-00002-of-00002.safetensors", "model.layers.2.mlp.fc2.bias": "model-00002-of-00002.safetensors", "model.layers.2.mlp.fc1.weight": "model-00002-of-00002.safetensors", "model.layers.2.mlp.fc1.bias": "model-00002-of-00002.safetensors", "model.layers.2.self_attn.v_proj.weight": "model-00002-of-00002.safetensors", "model.layers.2.self_attn.v_proj.bias": "model-00002-of-00002.safetensors", "model.layers.2.self_attn.k_proj.weight": "model-00002-of-00002.safetensors", "model.layers.2.self_attn.k_proj.bias": "model-00002-of-00002.safetensors", "model.layers.2.self_attn.q_proj.weight": "model-00002-of-00002.safetensors", "model.layers.2.self_attn.q_proj.bias": "model-00002-of-00002.safetensors", "model.layers.2.self_attn.dense.weight": "model-00002-of-00002.safetensors", "model.layers.2.self_attn.dense.bias": "model-00002-of-00002.safetensors", "model.layers.2.input_layernorm.weight": "model-00002-of-00002.safetensors", "model.layers.2.input_layernorm.bias": "model-00002-of-00002.safetensors", "model.layers.1.mlp.fc2.weight": "model-00002-of-00002.safetensors", "model.layers.1.mlp.fc2.bias": "model-00002-of-00002.safetensors", "model.layers.1.mlp.fc1.weight": "model-00002-of-00002.safetensors", "model.layers.1.mlp.fc1.bias": "model-00002-of-00002.safetensors", "model.layers.1.self_attn.v_proj.weight": "model-00002-of-00002.safetensors", "model.layers.1.self_attn.v_proj.bias": "model-00002-of-00002.safetensors", "model.layers.1.self_attn.k_proj.weight": "model-00002-of-00002.safetensors", "model.layers.1.self_attn.k_proj.bias": "model-00002-of-00002.safetensors", "model.layers.1.self_attn.q_proj.weight": "model-00002-of-00002.safetensors", "model.layers.1.self_attn.q_proj.bias": "model-00002-of-00002.safetensors", "model.layers.1.self_attn.dense.weight": "model-00002-of-00002.safetensors", "model.layers.1.self_attn.dense.bias": "model-00002-of-00002.safetensors", "model.layers.1.input_layernorm.weight": "model-00002-of-00002.safetensors", "model.layers.1.input_layernorm.bias": "model-00002-of-00002.safetensors", "model.layers.0.mlp.fc2.weight": "model-00002-of-00002.safetensors", "model.layers.0.mlp.fc2.bias": "model-00002-of-00002.safetensors", "model.layers.0.mlp.fc1.weight": "model-00002-of-00002.safetensors", "model.layers.0.mlp.fc1.bias": "model-00002-of-00002.safetensors", "model.layers.0.self_attn.v_proj.weight": "model-00002-of-00002.safetensors", "model.layers.0.self_attn.v_proj.bias": "model-00002-of-00002.safetensors", "model.layers.0.self_attn.k_proj.weight": "model-00002-of-00002.safetensors", "model.layers.0.self_attn.k_proj.bias": "model-00002-of-00002.safetensors", "model.layers.0.self_attn.q_proj.weight": "model-00002-of-00002.safetensors", "model.layers.0.self_attn.q_proj.bias": "model-00002-of-00002.safetensors", "model.layers.0.self_attn.dense.weight": "model-00002-of-00002.safetensors", "model.layers.0.self_attn.dense.bias": "model-00002-of-00002.safetensors", "model.layers.0.input_layernorm.weight": "model-00002-of-00002.safetensors", "model.layers.0.input_layernorm.bias": "model-00002-of-00002.safetensors", "model.embed_tokens.weight": "model-00002-of-00002.safetensors"}}

special_tokens_map.json ADDED Viewed

	@@ -0,0 +1,23 @@

+{
+  "bos_token": {
+    "content": "<|endoftext|>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "eos_token": {
+    "content": "<|endoftext|>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "unk_token": {
+    "content": "<|endoftext|>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  }
+}

tokenizer.json ADDED Viewed

The diff for this file is too large to render. See raw diff

tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,323 @@

+{
+  "add_prefix_space": false,
+  "added_tokens_decoder": {
+    "50256": {
+      "content": "<|endoftext|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "50257": {
+      "content": "                               ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50258": {
+      "content": "                              ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50259": {
+      "content": "                             ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50260": {
+      "content": "                            ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50261": {
+      "content": "                           ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50262": {
+      "content": "                          ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50263": {
+      "content": "                         ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50264": {
+      "content": "                        ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50265": {
+      "content": "                       ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50266": {
+      "content": "                      ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50267": {
+      "content": "                     ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50268": {
+      "content": "                    ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50269": {
+      "content": "                   ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50270": {
+      "content": "                  ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50271": {
+      "content": "                 ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50272": {
+      "content": "                ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50273": {
+      "content": "               ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50274": {
+      "content": "              ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50275": {
+      "content": "             ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50276": {
+      "content": "            ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50277": {
+      "content": "           ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50278": {
+      "content": "          ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50279": {
+      "content": "         ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50280": {
+      "content": "        ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50281": {
+      "content": "       ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50282": {
+      "content": "      ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50283": {
+      "content": "     ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50284": {
+      "content": "    ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50285": {
+      "content": "   ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50286": {
+      "content": "  ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50287": {
+      "content": "\t\t\t\t\t\t\t\t\t",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50288": {
+      "content": "\t\t\t\t\t\t\t\t",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50289": {
+      "content": "\t\t\t\t\t\t\t",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50290": {
+      "content": "\t\t\t\t\t\t",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50291": {
+      "content": "\t\t\t\t\t",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50292": {
+      "content": "\t\t\t\t",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50293": {
+      "content": "\t\t\t",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50294": {
+      "content": "\t\t",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    }
+  },
+  "bos_token": "<|endoftext|>",
+  "clean_up_tokenization_spaces": true,
+  "eos_token": "<|endoftext|>",
+  "model_max_length": 2048,
+  "tokenizer_class": "CodeGenTokenizer",
+  "unk_token": "<|endoftext|>"
+}

vocab.json ADDED Viewed

The diff for this file is too large to render. See raw diff