DVBench
/
Runs
/
Ingest
Corpus
Dataset or source name used to identify this text.
Tokenizer
Tokenizer used to match text against training transcripts.
cl100k_base
o200k_base
o200k_harmony
p50k_base
p50k_edit
r50k_base
gpt2
llama3
deepseek_v3
qwen2
mistral_v3
Minimum window
Smallest token window to index, from 1 to 65536.
Maximum window
Largest token window to index, at least the minimum and at most 65536.
Text
Source text to tokenize and index.
Tokenize and index