adding model weights

Browse files

Files changed (9) hide show

README.md +140 -0
added_tokens.json +5 -0
config.json +21 -0
merges.txt +0 -0
pytorch_model.bin +3 -0
special_tokens_map.json +35 -0
tokenizer.json +0 -0
tokenizer_config.json +198 -0
vocab.json +0 -0

README.md CHANGED Viewed

@@ -1,3 +1,143 @@
 ---
 license: apache-2.0
 ---

 ---
 license: apache-2.0
 ---
+## Description
+This model is intended to be used as an accelerator for [Granite-3.0-8b-instruct](https://huggingface.co/ibm-granite/granite-3.0-8b-instruct) and takes inspiration from the Medusa speculative decoding architecture.
+This accelerator modifies the MLP into a multi-stage MLP, where each stage predicts
+a single token in the draft based on both a state vector and sampled token
+from the prior stage (the base model can be considered stage 0).
+The state vector from the base model provides contextual information to the accelerator,
+while conditioning on prior sampled tokens allows it to produce higher-quality draft n-grams.
+Note: The underlying MLP speculator is a generic architecture that can be trained with any generative model to accelerate inference.
+Training is light-weight and can be completed in only a few days depending on base model size and speed.
+## Repository Links
+1. [Paged Attention KV-Cache / Speculator](https://github.com/foundation-model-stack/fms-extras)
+2. [Production Server with speculative decoding](https://github.com/IBM/text-generation-inference.git)
+3. [Speculator training](https://github.com/foundation-model-stack/fms-fsdp.git)
+## Samples
+_Note: For all samples, your environment must have access to cuda_
+### Use in IBM Production TGIS
+*To try this out running in a production-like environment, please use the pre-built docker image:*
+#### Setup
+```bash
+HF_HUB_CACHE=/hf_hub_cache
+chmod a+w $HF_HUB_CACHE
+HF_HUB_TOKEN="your huggingface hub token"
+TGIS_IMAGE=quay.io/wxpe/text-gen-server:main.ddc56ee
+docker pull $TGIS_IMAGE
+# optionally download granite-3.0-8b-instruct if the weights do not already exist
+docker run --rm \
+    -v $HF_HUB_CACHE:/models \
+    -e HF_HUB_CACHE=/models \
+    -e TRANSFORMERS_CACHE=/models \
+    $TGIS_IMAGE \
+    text-generation-server download-weights \
+    ibm-granite/granite-3.0-8b-instruct \
+    --token $HF_HUB_TOKEN
+# optionally download the speculator model if the weights do not already exist
+docker run --rm \
+    -v $HF_HUB_CACHE:/models \
+    -e HF_HUB_CACHE=/models \
+    -e TRANSFORMERS_CACHE=/models \
+    $TGIS_IMAGE \
+    text-generation-server download-weights \
+    ibm-granite/granite-3.0-8b-instruct-accelerator \
+    --token $HF_HUB_TOKEN
+# note: if the weights were downloaded separately (not with the above commands), please place them in the HF_HUB_CACHE directory and refer to them with /models/<model_name>
+docker run -d --rm --gpus all \
+    --name my-tgis-server \
+    -p 8033:8033 \
+    -v $HF_HUB_CACHE:/models \
+    -e HF_HUB_CACHE=/models \
+    -e TRANSFORMERS_CACHE=/models \
+    -e MODEL_NAME=ibm-granite/granite-3.0-8b-instruct \
+    -e SPECULATOR_NAME=ibm-granite/granite-3.0-8b-instruct-accelerator \
+    -e FLASH_ATTENTION=true \
+    -e PAGED_ATTENTION=true \
+    -e DTYPE=float16 \
+    $TGIS_IMAGE
+# check logs and wait for "gRPC server started on port 8033" and "HTTP server started on port 3000"
+docker logs my-tgis-server -f
+# get the client sample (Note: The first prompt will take longer as there is a warmup time)
+conda create -n tgis-client-env python=3.11
+conda activate tgis-client-env
+git clone --branch main --single-branch https://github.com/IBM/text-generation-inference.git
+cd text-generation-inference/integration_tests
+make gen-client
+pip install . --no-cache-dir
+```
+#### Run Sample
+```bash
+python sample_client.py
+```
+_Note: first prompt may be slower as there is a slight warmup time_
+### Use in Huggingface TGI
+#### start the server
+```bash
+model=ibm-granite/granite-3.0-8b-instruct-accelerator
+volume=$PWD/data # share a volume with the Docker container to avoid downloading weights every run
+docker run --gpus all --shm-size 1g -p 8080:80 -v $volume:/data ghcr.io/huggingface/text-generation-inference:latest --model-id $model
+```
+_note: for tensor parallel, add --num-shard_
+#### make a request
+```bash
+curl 127.0.0.1:8080/generate_stream \
+    -X POST \
+    -d '{"inputs":"What is Deep Learning?","parameters":{"max_new_tokens":20}}' \
+    -H 'Content-Type: application/json'
+```
+### Use in vLLM
+```from vllm import LLM, SamplingParams
+# Sample prompts.
+prompts = [
+    "The president of the United States is",
+]
+# Create a sampling params object.
+sampling_params = SamplingParams(temperature=0.0)
+# Create an LLM.
+llm = LLM(
+    model="/path/to/granite-3.0-8b-instruct",
+    tensor_parallel_size=4,
+    speculative_model="/path/to/granite-3.0-8b-instruct-accelerator",
+    speculative_draft_tensor_parallel_size=1,
+    use_v2_block_manager=True,
+)
+# Generate texts from the prompts. The output is a list of RequestOutput objects
+# that contain the prompt, generated text, and other information.
+outputs = llm.generate(prompts, sampling_params)
+# Print the outputs.
+for output in outputs:
+    prompt = output.prompt
+    generated_text = output.outputs[0].text
+    print(f"Prompt: {prompt!r}, Generated text: {generated_text!r}")
+```

added_tokens.json ADDED Viewed

	@@ -0,0 +1,5 @@

+{
+  "<|end_of_role|>": 49153,
+  "<|start_of_role|>": 49152,
+  "<|tool_call|>": 49154
+}

config.json ADDED Viewed

	@@ -0,0 +1,21 @@

+{
+  "architectures": [
+    "MLPSpeculatorPreTrainedModel"
+  ],
+  "emb_dim": 4096,
+  "inner_dim": 4096,
+  "model_type": "mlp_speculator",
+  "n_candidates": 4,
+  "n_predict": 4,
+  "scale_input": true,
+  "tie_weights": true,
+  "top_k_tokens_per_head": [
+    4,
+    3,
+    2,
+    2
+  ],
+  "torch_dtype": "bfloat16",
+  "transformers_version": "4.41.2",
+  "vocab_size": 49155
+}

merges.txt ADDED Viewed

The diff for this file is too large to render. See raw diff

pytorch_model.bin ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:be9dd56e96508c2592ac857ceea862ea29f674a6cba6910bb23ba4462bb2ee13
+size 3355712106

special_tokens_map.json ADDED Viewed

	@@ -0,0 +1,35 @@

+{
+  "additional_special_tokens": [
+    "<|start_of_role|>",
+    "<|end_of_role|>",
+    "<|tool_call|>"
+  ],
+  "bos_token": {
+    "content": "<|end_of_text|>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "eos_token": {
+    "content": "<|end_of_text|>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "pad_token": {
+    "content": "<|end_of_text|>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "unk_token": {
+    "content": "<|end_of_text|>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  }
+}

tokenizer.json ADDED Viewed

The diff for this file is too large to render. See raw diff

tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,198 @@

+{
+  "add_bos_token": false,
+  "add_prefix_space": false,
+  "added_tokens_decoder": {
+    "0": {
+      "content": "<|end_of_text|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "1": {
+      "content": "<fim_prefix>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "2": {
+      "content": "<fim_middle>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "3": {
+      "content": "<fim_suffix>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "4": {
+      "content": "<fim_pad>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "5": {
+      "content": "<filename>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "6": {
+      "content": "<gh_stars>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "7": {
+      "content": "<issue_start>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "8": {
+      "content": "<issue_comment>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "9": {
+      "content": "<issue_closed>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "10": {
+      "content": "<jupyter_start>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "11": {
+      "content": "<jupyter_text>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "12": {
+      "content": "<jupyter_code>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "13": {
+      "content": "<jupyter_output>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "14": {
+      "content": "<empty_output>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "15": {
+      "content": "<commit_before>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "16": {
+      "content": "<commit_msg>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "17": {
+      "content": "<commit_after>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "18": {
+      "content": "<reponame>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "49152": {
+      "content": "<|start_of_role|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "49153": {
+      "content": "<|end_of_role|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "49154": {
+      "content": "<|tool_call|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    }
+  },
+  "additional_special_tokens": [
+    "<|start_of_role|>",
+    "<|end_of_role|>",
+    "<|tool_call|>"
+  ],
+  "bos_token": "<|end_of_text|>",
+  "chat_template": "{%- if tools %}\n    {{- '<|start_of_role|>available_tools<|end_of_role|>\n' }}\n    {%- for tool in tools %}\n    {{- tool | tojson(indent=4) }}\n    {%- if not loop.last %}\n        {{- '\n\n' }}\n    {%- endif %}\n    {%- endfor %}\n    {{- '<|end_of_text|>\n' }}\n{%- endif %}\n{%- for message in messages %}\n    {%- if message['role'] == 'system' %}\n    {{- '<|start_of_role|>system<|end_of_role|>' + message['content'] + '<|end_of_text|>\n' }}\n    {%- elif message['role'] == 'user' %}\n    {{- '<|start_of_role|>user<|end_of_role|>' + message['content'] + '<|end_of_text|>\n' }}\n    {%- elif message['role'] == 'assistant' %}\n    {{- '<|start_of_role|>assistant<|end_of_role|>'  + message['content'] + '<|end_of_text|>\n' }}\n    {%- elif message['role'] == 'assistant_tool_call' %}\n    {{- '<|start_of_role|>assistant<|end_of_role|><|tool_call|>' + message['content'] + '<|end_of_text|>\n' }}\n    {%- elif message['role'] == 'tool_response' %}\n    {{- '<|start_of_role|>tool_response<|end_of_role|>' + message['content'] + '<|end_of_text|>\n' }}\n    {%- endif %}\n    {%- if loop.last and add_generation_prompt %}\n    {{- '<|start_of_role|>assistant<|end_of_role|>' }}\n    {%- endif %}\n{%- endfor %}",
+  "clean_up_tokenization_spaces": true,
+  "eos_token": "<|end_of_text|>",
+  "errors": "replace",
+  "model_max_length": 9223372036854775807,
+  "pad_token": "<|end_of_text|>",
+  "padding_side": "left",
+  "tokenizer_class": "GPT2Tokenizer",
+  "unk_token": "<|end_of_text|>",
+  "vocab_size": 49152
+}

vocab.json ADDED Viewed

The diff for this file is too large to render. See raw diff