UNVERIFIED literature results. Source: Qwen2 Technical Report, Tables 2 and 4.
Multilingual exam, understanding, mathematics, and translation evaluations.
multilingual_understanding vs. estimated ISO FLOPs
9 matching runs · 6ND; missing tokens estimated at 20× parameters
All 10 filtered runs
| Run | Scale (B) | ISO FLOPs | Training data | multilingual_understanding | multilingual_exam | multilingual_mathematics | multilingual_translation |
|---|---|---|---|---|---|---|---|
| literature/qwen2_72b | ~72.000 | ~6.2e23 | 80.70 | 76.60 | 76.00 | 37.80 | |
| literature/llama_3_70b | ~70.000 | ~5.9e23 | 79.90 | 70.00 | 67.10 | 38.00 | |
| literature/qwen1.5_110b | ~110.000 | ~1.5e24 | 78.20 | 75.60 | 64.40 | 36.20 | |
| literature/qwen1.5_72b | ~72.000 | ~6.2e23 | 78.20 | 66.40 | 61.70 | 35.60 | |
| literature/mixtral_8x22b | — | Unknown | 77.70 | 63.50 | 62.90 | 23.30 | |
| literature/qwen2_7b | ~7.000 | ~5.9e21 | 72.00 | 59.20 | 57.50 | 31.50 | |
| literature/llama_3_8b | ~8.000 | ~7.7e21 | 68.60 | 52.30 | 36.30 | 31.90 | |
| literature/qwen1.5_7b | ~7.000 | ~5.9e21 | 67.60 | 47.70 | 37.30 | 28.40 | |
| literature/mistral_7b_v0.2 | ~7.000 | ~5.9e21 | 63.30 | 47.10 | 26.30 | 23.30 | |
| literature/gemma_7b | ~7.000 | ~5.9e21 | 58.30 | 42.70 | 39.10 | 31.20 |