- bytes (utf-8)
- –
- byte length
- –
- decodes to
- –
Pre-tokenizer regex pieces (0)
The regex splits text before any merging happens. Hover the “How it works” tab for the full pattern and the possessive-quantifier detective story.
Watch byte-pair encoding happen
Type anything. The pre-tokenizer splits it into pieces, each piece starts as raw UTF-8 bytes, and the pair with the lowest merge rank merges first — step by step, exactly like the real algorithm.
Train a BPE tokenizer from scratch
This is the real 1990s Sennrich algorithm: count every adjacent pair, merge the most frequent, repeat. It runs entirely in this page — no server, no libraries. Keep the corpus modest and it trains in seconds.
Vocabulary browser
Every mergeable token, searchable. Token ids 0–199997 are learned merges, 199999 is <|endoftext|>, 200018 is <|endofprompt|>, and 19 ids are unassigned.
Token byte-length distribution
The pipeline
<|endoftext|> (199999) and <|endofprompt|> (200018). Everything between them is ordinary text.The possessive-quantifier detective story
OpenAI publishes the o200k regex with possessive quantifiers (?+, *+, ++) — matchers that never give characters back. But tiktoken's core is written in Rust, and Rust's regex crate has no possessive quantifiers: it silently parses a++ as nested (a+)+, i.e. plain greedy.
That changes real output. Try it:
This implementation matches tiktoken's actual behavior — verified byte-identical on 3,019 adversarial inputs — not the paper regex. Even tiktoken's own Python encode_ordinary path disagrees with its Rust path here.
Why byte-level BPE won
Character-level vocabularies break on unseen scripts; word-level ones explode in size and can't handle typos. Byte-level BPE starts from 256 byte values and learns 199,998 merges from real text, so the vocabulary is finite, compact, and total: every Unicode string has exactly one encoding. The regex exists to keep linguistically sensible units together — that's why world (leading space) is usually one token, and why code and CJK text tokenize the way they do.
From-scratch, honestly
The engine — regex splitter, UTF-8 handling, merge loop, special-token logic, trainer — is written from zero in this repo's tokenizer.js (browser) and python/gpt_tokenizer.py, with no tiktoken import anywhere. The vocabulary data (token bytes → ranks) is OpenAI's public o200k_base.tiktoken file — you can't learn GPT-4o's merges without GPT-4o's training data, and pretending otherwise would be dishonest. Verification: python/verify.py diffs this implementation against real tiktoken on thousands of inputs.