UNVERIFIED literature results. Source: Qwen2 Technical Report, Tables 2 and 4.
Truthful response evaluation.
truthfulqa_0shot vs. estimated ISO FLOPs
9 matching runs · 6ND; missing tokens estimated at 20× parameters
All 10 filtered runs
| Run | Scale (B) | ISO FLOPs | Training data | truthfulqa_0shot |
|---|---|---|---|---|
| literature/qwen1.5_72b | ~72.000 | ~6.2e23 | 59.60 | |
| literature/qwen2_72b | ~72.000 | ~6.2e23 | 54.80 | |
| literature/qwen2_7b | ~7.000 | ~5.9e21 | 54.20 | |
| literature/qwen1.5_7b | ~7.000 | ~5.9e21 | 51.10 | |
| literature/mixtral_8x22b | — | Unknown | 51.00 | |
| literature/qwen1.5_110b | ~110.000 | ~1.5e24 | 49.60 | |
| literature/llama_3_70b | ~70.000 | ~5.9e23 | 45.60 | |
| literature/gemma_7b | ~7.000 | ~5.9e21 | 44.80 | |
| literature/llama_3_8b | ~8.000 | ~7.7e21 | 44.00 | |
| literature/mistral_7b_v0.2 | ~7.000 | ~5.9e21 | 42.20 |