ZeroXClem commited on
Commit
f369629
1 Parent(s): cd87a9d

Adding Evaluation Results

Browse files

This is an automated PR created with https://huggingface.co/spaces/Weyaxi/open-llm-leaderboard-results-pr

The purpose of this PR is to add evaluation results from the Open LLM Leaderboard to your model card.

If you encounter any issues, please report them to https://huggingface.co/spaces/Weyaxi/open-llm-leaderboard-results-pr/discussions

Files changed (1) hide show
  1. README.md +112 -4
README.md CHANGED
@@ -1,5 +1,8 @@
1
  ---
 
 
2
  license: apache-2.0
 
3
  tags:
4
  - merge
5
  - mergekit
@@ -13,8 +16,6 @@ tags:
13
  - nerd
14
  - homer
15
  - Qandora
16
- language:
17
- - en
18
  base_model:
19
  - bunnycore/Qandora-2.5-7B-Creative
20
  - allknowingroger/HomerSlerp1-7B
@@ -23,7 +24,101 @@ base_model:
23
  - jeffmeloy/Qwen2.5-7B-nerd-uncensored-v1.0
24
  - newsbang/Homer-v0.5-Qwen2.5-7B
25
  pipeline_tag: text-generation
26
- library_name: transformers
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
27
  ---
28
 
29
  # ZeroXClem/Qwen2.5-7B-HomerAnvita-NerdMix
@@ -205,4 +300,17 @@ This model is open-sourced under the **Apache-2.0 License**.
205
  - `jeffmeloy/Qwen2.5-7B-nerd-uncensored-v1.0`
206
  - `newsbang/Homer-v0.5-Qwen2.5-7B`
207
 
208
- ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ language:
3
+ - en
4
  license: apache-2.0
5
+ library_name: transformers
6
  tags:
7
  - merge
8
  - mergekit
 
16
  - nerd
17
  - homer
18
  - Qandora
 
 
19
  base_model:
20
  - bunnycore/Qandora-2.5-7B-Creative
21
  - allknowingroger/HomerSlerp1-7B
 
24
  - jeffmeloy/Qwen2.5-7B-nerd-uncensored-v1.0
25
  - newsbang/Homer-v0.5-Qwen2.5-7B
26
  pipeline_tag: text-generation
27
+ model-index:
28
+ - name: Qwen2.5-7B-HomerAnvita-NerdMix
29
+ results:
30
+ - task:
31
+ type: text-generation
32
+ name: Text Generation
33
+ dataset:
34
+ name: IFEval (0-Shot)
35
+ type: HuggingFaceH4/ifeval
36
+ args:
37
+ num_few_shot: 0
38
+ metrics:
39
+ - type: inst_level_strict_acc and prompt_level_strict_acc
40
+ value: 77.08
41
+ name: strict accuracy
42
+ source:
43
+ url: https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard?query=ZeroXClem/Qwen2.5-7B-HomerAnvita-NerdMix
44
+ name: Open LLM Leaderboard
45
+ - task:
46
+ type: text-generation
47
+ name: Text Generation
48
+ dataset:
49
+ name: BBH (3-Shot)
50
+ type: BBH
51
+ args:
52
+ num_few_shot: 3
53
+ metrics:
54
+ - type: acc_norm
55
+ value: 36.58
56
+ name: normalized accuracy
57
+ source:
58
+ url: https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard?query=ZeroXClem/Qwen2.5-7B-HomerAnvita-NerdMix
59
+ name: Open LLM Leaderboard
60
+ - task:
61
+ type: text-generation
62
+ name: Text Generation
63
+ dataset:
64
+ name: MATH Lvl 5 (4-Shot)
65
+ type: hendrycks/competition_math
66
+ args:
67
+ num_few_shot: 4
68
+ metrics:
69
+ - type: exact_match
70
+ value: 29.53
71
+ name: exact match
72
+ source:
73
+ url: https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard?query=ZeroXClem/Qwen2.5-7B-HomerAnvita-NerdMix
74
+ name: Open LLM Leaderboard
75
+ - task:
76
+ type: text-generation
77
+ name: Text Generation
78
+ dataset:
79
+ name: GPQA (0-shot)
80
+ type: Idavidrein/gpqa
81
+ args:
82
+ num_few_shot: 0
83
+ metrics:
84
+ - type: acc_norm
85
+ value: 9.28
86
+ name: acc_norm
87
+ source:
88
+ url: https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard?query=ZeroXClem/Qwen2.5-7B-HomerAnvita-NerdMix
89
+ name: Open LLM Leaderboard
90
+ - task:
91
+ type: text-generation
92
+ name: Text Generation
93
+ dataset:
94
+ name: MuSR (0-shot)
95
+ type: TAUR-Lab/MuSR
96
+ args:
97
+ num_few_shot: 0
98
+ metrics:
99
+ - type: acc_norm
100
+ value: 14.41
101
+ name: acc_norm
102
+ source:
103
+ url: https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard?query=ZeroXClem/Qwen2.5-7B-HomerAnvita-NerdMix
104
+ name: Open LLM Leaderboard
105
+ - task:
106
+ type: text-generation
107
+ name: Text Generation
108
+ dataset:
109
+ name: MMLU-PRO (5-shot)
110
+ type: TIGER-Lab/MMLU-Pro
111
+ config: main
112
+ split: test
113
+ args:
114
+ num_few_shot: 5
115
+ metrics:
116
+ - type: acc
117
+ value: 38.13
118
+ name: accuracy
119
+ source:
120
+ url: https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard?query=ZeroXClem/Qwen2.5-7B-HomerAnvita-NerdMix
121
+ name: Open LLM Leaderboard
122
  ---
123
 
124
  # ZeroXClem/Qwen2.5-7B-HomerAnvita-NerdMix
 
300
  - `jeffmeloy/Qwen2.5-7B-nerd-uncensored-v1.0`
301
  - `newsbang/Homer-v0.5-Qwen2.5-7B`
302
 
303
+ ---
304
+ # [Open LLM Leaderboard Evaluation Results](https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard)
305
+ Detailed results can be found [here](https://huggingface.co/datasets/open-llm-leaderboard/details_ZeroXClem__Qwen2.5-7B-HomerAnvita-NerdMix)
306
+
307
+ | Metric |Value|
308
+ |-------------------|----:|
309
+ |Avg. |34.17|
310
+ |IFEval (0-Shot) |77.08|
311
+ |BBH (3-Shot) |36.58|
312
+ |MATH Lvl 5 (4-Shot)|29.53|
313
+ |GPQA (0-shot) | 9.28|
314
+ |MuSR (0-shot) |14.41|
315
+ |MMLU-PRO (5-shot) |38.13|
316
+