UNVERIFIED literature results. Source: Qwen2 Technical Report, Tables 2 and 4.
Code generation evaluations.
humaneval_0shot vs. estimated ISO FLOPs
9 matching runs · 6ND; missing tokens estimated at 20× parameters
All 10 filtered runs
| Run | Scale (B) | ISO FLOPs | Training data | humaneval_0shot | evalplus_0shot | mbpp_0shot | multipl_e_0shot |
|---|---|---|---|---|---|---|---|
| literature/qwen2_72b | ~72.000 | ~6.2e23 | 64.60 | 65.40 | 76.90 | 59.60 | |
| literature/qwen1.5_110b | ~110.000 | ~1.5e24 | 54.30 | 57.70 | 70.90 | 52.70 | |
| literature/qwen2_7b | ~7.000 | ~5.9e21 | 51.20 | 54.20 | 65.90 | 46.30 | |
| literature/llama_3_70b | ~70.000 | ~5.9e23 | 48.20 | 54.80 | 70.40 | 46.30 | |
| literature/mixtral_8x22b | — | Unknown | 46.30 | 54.10 | 71.70 | 46.70 | |
| literature/qwen1.5_72b | ~72.000 | ~6.2e23 | 46.30 | 52.90 | 66.90 | 41.80 | |
| literature/gemma_7b | ~7.000 | ~5.9e21 | 37.20 | 39.60 | 50.60 | 29.70 | |
| literature/qwen1.5_7b | ~7.000 | ~5.9e21 | 36.00 | 40.00 | 51.60 | 28.10 | |
| literature/llama_3_8b | ~8.000 | ~7.7e21 | 33.50 | 40.30 | 53.90 | 22.60 | |
| literature/mistral_7b_v0.2 | ~7.000 | ~5.9e21 | 29.30 | 36.40 | 51.10 | 29.40 |