UNVERIFIED literature results. Source: Qwen2 Technical Report, Tables 2 and 4.
Commonsense completion and disambiguation evaluations.
hellaswag_10shot vs. estimated ISO FLOPs
9 matching runs · 6ND; missing tokens estimated at 20× parameters
All 10 filtered runs
| Run | Scale (B) | ISO FLOPs | Training data | hellaswag_10shot | winogrande_5shot |
|---|---|---|---|---|---|
| literature/mixtral_8x22b | — | Unknown | 88.70 | 85.00 | |
| literature/llama_3_70b | ~70.000 | ~5.9e23 | 88.00 | 85.30 | |
| literature/qwen2_72b | ~72.000 | ~6.2e23 | 87.60 | 85.10 | |
| literature/qwen1.5_110b | ~110.000 | ~1.5e24 | 87.50 | 83.50 | |
| literature/qwen1.5_72b | ~72.000 | ~6.2e23 | 86.00 | 83.00 | |
| literature/mistral_7b_v0.2 | ~7.000 | ~5.9e21 | 83.20 | 78.40 | |
| literature/gemma_7b | ~7.000 | ~5.9e21 | 82.20 | 79.00 | |
| literature/llama_3_8b | ~8.000 | ~7.7e21 | 82.10 | 77.40 | |
| literature/qwen2_7b | ~7.000 | ~5.9e21 | 80.70 | 77.00 | |
| literature/qwen1.5_7b | ~7.000 | ~5.9e21 | 78.50 | 71.30 |