基准介绍

BIG-Bench 的 23 个最难任务子集,覆盖逻辑推理、符号操作、常识推理等领域,传统模型表现不佳,用于评估模型的复杂推理能力。

评测指标

指标名单位方向
accuracy%↑ 越高越好

数据来源

模型得分排名

#模型厂商得分
1Gemini 2.0 Flash Thinkinggoogle87.9
2Gemini 1.5 Pro 002google87.2
3GPT-4o (2024-05-13)openai86.7
4DeepSeek V3deepseek84.9
5GPT-4o (2024-08-06)openai84.9
6Claude 3.5 Sonnetanthropic84.5
7o1 Previewopenai84.5
8Claude 3.5 Sonnet (2024-10-22)anthropic84.2
9Gemini 1.5 Progoogle84
10Grok-2xai84
11Llama 3.3 70Bmeta83.9
12Gemini 1.0 Ultragoogle83.8
13GLM-4 Plusother83.6
14Gemini 2.0 Flashgoogle83.4
15Mistral Mediummistral83.4
16Qwen2 72Balibaba83.4
17Sonar Reasoningother83.4
18GPT-4oopenai83.1
19o1openai83.1
20Llama 3.1 405Bmeta82.9
21Jamba 1.5 Largeother82.8
22Command Nightlycohere82.7
23Qwen2.5 72Balibaba82.4
24Qwen1.5 110Balibaba82.3
25Command R+ (08-2024)cohere82.1
26Hermes 3 Llama 3.1 405Bother82.1
27Claude 3 Opusanthropic81.6
28Sonar Largeother81.6
29Mistral Large 2mistral81
30GPT-4openai80.9
31GPT-4 Visionopenai80.9
32GPT-4 1106 Previewopenai80
33WizardLM Team WizardLM 2 8x22Bother80
34Gemini 1.5 Flash 002google79.8
35Claude 3 Opus (2024-02-29)anthropic79.5
36Phi-3.5 MoEother79.4
37GPT-4 32Kopenai79.2
38Gemini 1.0 Progoogle79.1
39GLM-4 Airother79
40Command R (08-2024)cohere78.8
41Sonar Hugeother78.6
42GPT-4 Turboopenai78.5
43Command Rcohere78.4
44Mixtral 8x7Bmistral78.4
45Yi Visionother78.3
46GPT-4 0125 Previewopenai78.2
47Llama 3 70Bmeta77.6
48Jamba 1.5other77.3
49o1 miniopenai77.2
50Yi Largeother77.1
51Qwen2.5 32Balibaba76.8
52Mistral Largemistral76.7
53NVIDIA Llama 3.1 Nemotron 70Bother76.6
54Qwen1.5 72Balibaba76.5
55Falcon 180Bother76.4
56Yi Large Turboother76.4
57Claude 3 Sonnet (2024-02-29)anthropic76.2
58Claude 3 Sonnetanthropic76.1
59Llama 3.2 90B Visionmeta76.1
60Hermes 3 Llama 3.1 70Bother75.9
61Claude 3 Haikuanthropic75.7
62Jamba 1.5 Miniother75.6
63Nous Hermes 2 Yi 34Bother75.6
64DeepSeek V2 Chatdeepseek75.4
65Grok-2 Visionxai75.4
66GPT-4 Vision Previewopenai75
67Gemini 1.0 Flashgoogle74.8
68Gemini 1.5 Flash-8Bgoogle74.8
69Grok-2 Minixai74.6
70Mixtral 8x22Bmistral74.5
71Microsoft WizardMath 7B v1other74.4
72Yi 1.5 34Bother74.1
73Gemini 1.5 Flash-8B 002google73.6
74Microsoft WizardLM 2 8x22Bother73.6
75Llemma 7Bother73.1
76Qwen1.5 32Balibaba73
77DeepSeek LLM 67Bdeepseek72.6
78Nous Hermes 2 Mixtral 8x7Bother72.4
79Qwen2 57Balibaba72.4
80DBRX Baseother72.3
81DeepSeek Coder V2deepseek72.3
82Claude 3 Haiku (2024-03-07)anthropic72.2
83Qwen2.5 14Balibaba72.2
84Code Llama 70Bmeta72.1
85Jamba Instructother71.8
86Nous Hermes 2 Solar 10.7Bother71.8
87Phi-3 Smallother71.8
88DBRX Instructother71.7
89Gemini 1.5 Flashgoogle71.7
90Sonar Smallother71.7
91Mistral Smallmistral71.6
92Qwen1.5 14Balibaba71.6
93OLMo 7B Instructother71.5
94Code Bisongoogle71.4
95Phi-3 Mediumother71.4
96Mistral 7B v0.2mistral71.3
97Mistral Small 3mistral71.3
98Gemma 2 27Bgoogle71.2
99Mistral Nemomistral71.2
100Code Llama 13Bmeta71
101OpenChat 3.6 8Bother71
102Orca 2 13Bother70.8
103StableLM 2 12Bother70.6
104StarChat2 15B v0.1other70.6
105DeepSeek V2deepseek70.5
106GLM-4 Flashother70.5
107Yi 1.5 6Bother70.5
108GPT-4o miniopenai70.3
109Llama 3.1 Nemotron 70Bmeta70.3
110Zephyr ORPO 141B Alphaother70.3
111Claude 3.5 Haikuanthropic70.2
112Llama 3.1 70Bmeta70.2
113Phi-1other69.9
114Mathstral 7Bmistral69.8
115Microsoft WizardLM 2 7Bother69.4
116Llama 3.2 11B Visionmeta69.1
117OLMo 1.7 7Bother68.9
118OLMo 7Bother67.7
119GLM-4 9B Chatother67.2
120Microsoft WizardCoder Python 34Bother67.2
121Llama 3.1 8Bmeta67.1
122Code Llama 7Bmeta66.9
123Phi-4other66.9
124Gemma 2 9Bgoogle66.7
125Command R+cohere66.1
126Command R7Bcohere66.1
127Codestral Mambamistral66
128DeepSeek Math 7Bdeepseek65.9
129Gemma 7Bgoogle65.7
130Mistral 7B v0.1mistral65.4
131Yi 1.5 9Bother65.3
132GLM-4V 9Bother65.1
133Llama 3 8Bmeta65
134Baichuan2 13B Chatother64.2
135OLMo 7B SFTother64
136GPT-3.5openai63.9
137Claude 2.1anthropic63.8
138Open-Platypusother63.7
139Qwen2 7Balibaba63.3
140Claude Instant 1anthropic63.1
141OLMo 2 1124 7Bother63.1
142StarCoder2 15Bother62.8
143Hermes 3 Llama 3.1 8Bother62.6
144Llama 2 70Bmeta62.5
145Argilla Notus 7B v1other62.4
146Llama 2 13Bmeta62.1
147StarCoder2 3Bother62.1
148Qwen2.5 7Balibaba61.9
149DeepSeek Coder 33Bdeepseek61.6
150Flan-T5 XXLother60.5
151Mistral 7B v0.3mistral60.5
152OpenChat 3.5 1210other60.4
153Claude 2anthropic60.2
154Ministral 8Bmistral60
155StarCoder2 7Bother60
156MPT 7Bother59.9
157DeepSeek Coder 7Bdeepseek59.7
158Code Llama 34Bmeta59.1
159StableCode 3Bother58.4
160Codestralmistral58.3
161Grok Vision Betaxai56.8
162Nous Capybara 34Bother56.1
163Phi-3 Miniother55.9
164ChatGLM3 6Bother55.7
165Command Lightcohere55.5
166Yi 6Bother55.3
167StableLM 3 4Bother54.9
168Llama 2 7Bmeta54.7
169Chat Bisongoogle54.6
170Baichuan2 7B Chatother53
171Gemma 2Bgoogle53
172Jurassic-2 Ultraother52.9
173PaLM 2google52.9
174Flan-T5 XLother52.5
175TigerBot 70B Chatother51.9
176Jurassic-2 Midother51.8
177Qwen2 1.5Balibaba51.5
178Qwen2.5 1.5Balibaba51.5
179Mistral Tinymistral51.3
180Qwen2.5 0.5Balibaba51
181Zephyr 7B Alphaother50.7
182Falcon 40Bother50.6
183GPT-3.5 Turbo 16Kopenai50.2
184Grok Betaxai50.2
185Ministral 3Bmistral50.1
186Zephyr 7B Betaother50
187Nous Capybara 7Bother49.6
188Phi-1.5other49.5
189Flan-UL2other49.2
190MPT 30Bother49.2
191Qwen2.5 3Balibaba48.8
192Yi 34Bother48.8
193Capybara 1.5Bother48.3
194GPT-3.5 Turboopenai48.2
195Text Bisongoogle48.2
196Phi-2other47.8
197Llama Guard 2 8Bmeta47.6
198Phi-3.5 Miniother47.5
199Llama 3.2 3Bmeta46
200Llama Guard 3 8Bmeta44.1
201StableLM Zephyr 3Bother43.9
202StableLM 2 1.6Bother40.7
203Phi-3 Visionother40.1
204Llama 3.2 1Bmeta38.8
205Embed English v3cohere0

BBH (BIG-Bench Hard)

描述

BIG-Bench 的 23 个最难任务子集,覆盖逻辑推理、符号操作、常识推理等领域,传统模型表现不佳,用于评估模型的复杂推理能力。

核心规格

分类许可证最后更新
reasoningMIT2022-10-01

基准

单位
%

数据来源

官方地址