jphme
/

vicuna-13b-v1.3-ger

+---
+language:
+- de
+- en
+pipeline_tag: text-generation
+inference: false
+---
+# Vicuna 13b v1.3 Ger
+vicuna-13b-v1.3-ger is a variant of [LMSYS](https://huggingface.co/lmsys)´s [Vicuna 13b v1.3](https://huggingface.co/lmsys/vicuna-13b-v1.3) model, finetuned on an additional dataset in German language. The original model has been trained on explain tuned datasets, created using instructions and input from WizardLM, Alpaca & Dolly-V2 datasets and applying Orca Research Paper dataset construction approaches.
+This model is optimized for German text, providing proficiency in understanding, generating, and interacting with German language content. However the model is not yet fully optimized for German language, as it has been trained on a small, experimental dataset and has limited capabilities due to the small parameter count.
+I am working on improving the model´s capabilities and will update the model if there is sufficient interest.
+## Results
+I did only evaluate the output on a small, handcrafted sample on test prompts in German, confirming that the model's ability to understand and generate German text is well above the base model.
+## Problems
+There might be inconsistencies in multi-turn chat applications, as there was a small problem with the <eos> tokens during preparation of the finetuning dataset.
+Please report any problems so I can fix this for the next version.
+# Original Vicuna Model Card
+## Model Details
+Vicuna is a chat assistant trained by fine-tuning LLaMA on user-shared conversations collected from ShareGPT.
+- **Developed by:** [LMSYS](https://lmsys.org/)
+- **Model type:** An auto-regressive language model based on the transformer architecture.
+- **License:** Non-commercial license
+- **Finetuned from model:** [LLaMA](https://arxiv.org/abs/2302.13971).
+### Model Sources
+- **Repository:** https://github.com/lm-sys/FastChat
+- **Blog:** https://lmsys.org/blog/2023-03-30-vicuna/
+- **Paper:** https://arxiv.org/abs/2306.05685
+- **Demo:** https://chat.lmsys.org/
+## Uses
+The primary use of Vicuna is research on large language models and chatbots.
+The primary intended users of the model are researchers and hobbyists in natural language processing, machine learning, and artificial intelligence.
+## How to Get Started with the Model
+- Command line interface: https://github.com/lm-sys/FastChat#vicuna-weights.
+- APIs (OpenAI API, Huggingface API): https://github.com/lm-sys/FastChat/tree/main#api.
+## Training Details
+Vicuna v1.3 is fine-tuned from LLaMA with supervised instruction fine-tuning.
+The training data is around 140K conversations collected from ShareGPT.com.
+See more details in the "Training Details of Vicuna Models" section in the appendix of this [paper](https://arxiv.org/pdf/2306.05685.pdf).
+## Evaluation
+Vicuna is evaluated with standard benchmarks, human preference, and LLM-as-a-judge. See more details in this [paper](https://arxiv.org/pdf/2306.05685.pdf) and [leaderboard](https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboard).
+## Difference between different versions of Vicuna
+See [vicuna_weights_version.md](https://github.com/lm-sys/FastChat/blob/main/docs/vicuna_weights_version.md)