IFEval: Benchmark Detailed Guide
Detailed guide to IFEval: category, metrics, sources and applicable models.
Overview
评估大型语言模型严格遵循指令的能力,包含500个可验证的指令
Metrics
| Metric | Unit | Direction |
|---|---|---|
| prompt_strict_acc | % | ↑ Higher is better |
| prompt_loose_acc | % | ↑ Higher is better |
Sources
Model Score Ranking
IFEval
Description
评估大型语言模型严格遵循指令的能力,包含500个可验证的指令
Core Specifications
| Category | License | Last Updated |
|---|---|---|
| reasoning | Apache-2.0 | 2024-02-01 |
Benchmark
| Unit |
|---|
| % |
| % |