UNVERIFIED literature results. Source: Qwen2 Technical Report, Tables 2 and 4.
Multi-step and scientific reasoning evaluations.
bbh_3shot vs. estimated ISO FLOPs
9 matching runs · 6ND; missing tokens estimated at 20× parameters
All 10 filtered runs
| Run | Scale (B) | ISO FLOPs | Training data | bbh_3shot | arc_challenge_25shot |
|---|---|---|---|---|---|
| literature/qwen1.5_7b | ~7.000 | ~5.9e21 | 40.20 | 54.20 | |
| literature/gemma_7b | ~7.000 | ~5.9e21 | 55.10 | 61.10 | |
| literature/mistral_7b_v0.2 | ~7.000 | ~5.9e21 | 56.10 | 60.00 | |
| literature/llama_3_8b | ~8.000 | ~7.7e21 | 57.70 | 59.30 | |
| literature/qwen2_7b | ~7.000 | ~5.9e21 | 62.60 | 60.60 | |
| literature/qwen1.5_72b | ~72.000 | ~6.2e23 | 65.50 | 65.90 | |
| literature/qwen1.5_110b | ~110.000 | ~1.5e24 | 74.80 | 69.60 | |
| literature/mixtral_8x22b | — | Unknown | 78.90 | 70.70 | |
| literature/llama_3_70b | ~70.000 | ~5.9e23 | 81.00 | 68.80 | |
| literature/qwen2_72b | ~72.000 | ~6.2e23 | 82.40 | 68.90 |