DVBench
/
Benchmarks
/
New
Identity
Name
A unique, descriptive name.
Description
A short summary shown in listings.
README
Markdown
Detailed instructions and context, formatted as Markdown.
Measures
Required measurements
Measurements a run must provide to appear on this benchmark. The key measurement is included automatically.
arc_challenge_25shot
bbh_3shot
evalplus_0shot
gpqa_5shot
gsm8k_5shot
hellaswag_10shot
humaneval_0shot
math_4shot
mbpp_0shot
mean_confidence
mean_logit
mmlu_5shot
mmlu_pro_5shot
multilingual_exam
multilingual_mathematics
multilingual_translation
multilingual_understanding
multipl_e_0shot
perplexity
smoke_accuracy
tape_validation/20260905_201353/next_token_mse
tape_validation/jax_20260906_005256/next_token_mse
tape_validation/torch_20260906_004839/next_token_mse
tape_validation/torch_20260906_004900/next_token_mse
tape_validation/torch_20260906_004936/next_token_mse
tape_validation/torch_20260906_005049/next_token_mse
tape_validation/torch_20260906_005141/next_token_mse
tape_validation/torch_20260906_171053/next_token_mse
tape_validation/torch_20260906_171228/next_token_mse
tape_validation/torch_20260906_171316/next_token_mse
theorem_qa_5shot
truthfulqa_0shot
winogrande_5shot
Key measurement
The measurement used to rank runs on this benchmark.
Select…
arc_challenge_25shot
bbh_3shot
evalplus_0shot
gpqa_5shot
gsm8k_5shot
hellaswag_10shot
humaneval_0shot
math_4shot
mbpp_0shot
mean_confidence
mean_logit
mmlu_5shot
mmlu_pro_5shot
multilingual_exam
multilingual_mathematics
multilingual_translation
multilingual_understanding
multipl_e_0shot
perplexity
smoke_accuracy
tape_validation/20260905_201353/next_token_mse
tape_validation/jax_20260906_005256/next_token_mse
tape_validation/torch_20260906_004839/next_token_mse
tape_validation/torch_20260906_004900/next_token_mse
tape_validation/torch_20260906_004936/next_token_mse
tape_validation/torch_20260906_005049/next_token_mse
tape_validation/torch_20260906_005141/next_token_mse
tape_validation/torch_20260906_171053/next_token_mse
tape_validation/torch_20260906_171228/next_token_mse
tape_validation/torch_20260906_171316/next_token_mse
theorem_qa_5shot
truthfulqa_0shot
winogrande_5shot
Default sort
Whether lower or higher values of the key measurement rank first.
Lower is better
Higher is better
Cancel
Add benchmark