DVBench / Measures
arc_challenge_25shot
ARC-Challenge, 25-shot score (%).
bbh_3shot
BIG-Bench Hard, 3-shot score (%).
evalplus_0shot
EvalPlus, 0-shot score (%).
gpqa_5shot
GPQA, 5-shot score (%).
gsm8k_5shot
GSM8K, 5-shot score (%).
hellaswag_10shot
HellaSwag, 10-shot score (%).
humaneval_0shot
HumanEval, 0-shot score (%).
math_4shot
MATH, 4-shot score (%).
mbpp_0shot
MBPP, 0-shot score (%).
mean_confidence
Mean maximum class probability over the evaluation set.
mean_logit
Mean checkpoint logit over the evaluation set.
mmlu_5shot
MMLU, 5-shot score (%).
mmlu_pro_5shot
MMLU-Pro, 5-shot score (%).
multilingual_exam
Multilingual exam aggregate score.
multilingual_mathematics
Multilingual MGSM aggregate score.
multilingual_translation
Multilingual Flores-101 translation score.
multilingual_understanding
Multilingual understanding aggregate score.
multipl_e_0shot
MultiPL-E, 0-shot score (%).
perplexity
Evaluation perplexity; lower is better.
smoke_accuracy
Smoke test: fraction of parity labels predicted correctly.
tape_validation/20260905_201353/next_token_mse
Tiny next-token regression MSE.
tape_validation/jax_20260906_005256/next_token_mse
Tiny next-token regression MSE.
tape_validation/torch_20260906_004839/next_token_mse
Tiny next-token regression MSE.
tape_validation/torch_20260906_004900/next_token_mse
Tiny next-token regression MSE.
tape_validation/torch_20260906_004936/next_token_mse
Tiny next-token regression MSE.
tape_validation/torch_20260906_005049/next_token_mse
Tiny next-token regression MSE.
tape_validation/torch_20260906_005141/next_token_mse
Tiny next-token regression MSE.
tape_validation/torch_20260906_171053/next_token_mse
Tiny next-token regression MSE.
tape_validation/torch_20260906_171228/next_token_mse
Tiny next-token regression MSE.
tape_validation/torch_20260906_171316/next_token_mse
Tiny next-token regression MSE.
theorem_qa_5shot
TheoremQA, 5-shot score (%).
truthfulqa_0shot
TruthfulQA, 0-shot score (%).
winogrande_5shot
WinoGrande, 5-shot score (%).