Loading the tokenizer…
Fetching 200,000 token ranks (1.6 MB). Everything runs locally in your browser.

Byte-level BPE · from scratch · no dependencies

o200k, rebuilt by hand.

The GPT-4o tokenizer — 200,019 token ids, a 7-branch multilingual pre-tokenizer regex, and 199,998 byte-pair merges — reimplemented from zero in dependency-free JavaScript. Verified byte-identical to OpenAI's tiktoken on thousands of adversarial inputs. Paste text, watch it tokenize, train your own BPE, and inspect the vocabulary.

200,019 token ids 199,998 merge ranks 0 dependencies byte-identical to tiktoken
0
tokens
0
characters
0
utf-8 bytes
–
chars / token
–
≈ gpt-4o input cost
–
bytes (utf-8)
–
byte length
–
decodes to
–
Pre-tokenizer regex pieces (0)

The regex splits text before any merging happens. Hover the “How it works” tab for the full pattern and the possessive-quantifier detective story.

Watch byte-pair encoding happen

Type anything. The pre-tokenizer splits it into pieces, each piece starts as raw UTF-8 bytes, and the pair with the lowest merge rank merges first — step by step, exactly like the real algorithm.

Train a BPE tokenizer from scratch

This is the real 1990s Sennrich algorithm: count every adjacent pair, merge the most frequent, repeat. It runs entirely in this page — no server, no libraries. Keep the corpus modest and it trains in seconds.

Vocabulary browser

Every mergeable token, searchable. Token ids 0–199997 are learned merges, 199999 is <|endoftext|>, 200018 is <|endofprompt|>, and 19 ids are unassigned.

Token byte-length distribution

The pipeline

1 · special tokensScan for <|endoftext|> (199999) and <|endofprompt|> (200018). Everything between them is ordinary text.
↓
2 · pre-tokenizer regexSplit the text with a 7-branch multilingual regex: words (with optional leading punctuation and contractions), numbers, punctuation runs, newlines, whitespace. This decides the maximum span any token may cover — merges never cross a regex boundary.
↓
3 · utf-8 bytesEach regex piece becomes raw bytes. The vocabulary's first entries are the 256 single bytes, so any text is encodable — even emoji, even binary garbage.
↓
4 · byte-pair mergesRepeatedly merge the adjacent pair with the lowest rank (rank = merge priority learned during training). Stop when no adjacent pair exists in the rank table.
↓
5 · idsLook every merged byte-chunk up in the rank table → token ids. Decoding is the reverse: ids → bytes → UTF-8 text.

The possessive-quantifier detective story

OpenAI publishes the o200k regex with possessive quantifiers (?+, *+, ++) — matchers that never give characters back. But tiktoken's core is written in Rust, and Rust's regex crate has no possessive quantifiers: it silently parses a++ as nested (a+)+, i.e. plain greedy.

That changes real output. Try it:

tiktoken (greedy — real)
–
possessive (published)
–

This implementation matches tiktoken's actual behavior — verified byte-identical on 3,019 adversarial inputs — not the paper regex. Even tiktoken's own Python encode_ordinary path disagrees with its Rust path here.

Why byte-level BPE won

Character-level vocabularies break on unseen scripts; word-level ones explode in size and can't handle typos. Byte-level BPE starts from 256 byte values and learns 199,998 merges from real text, so the vocabulary is finite, compact, and total: every Unicode string has exactly one encoding. The regex exists to keep linguistically sensible units together — that's why world (leading space) is usually one token, and why code and CJK text tokenize the way they do.

From-scratch, honestly

The engine — regex splitter, UTF-8 handling, merge loop, special-token logic, trainer — is written from zero in this repo's tokenizer.js (browser) and python/gpt_tokenizer.py, with no tiktoken import anywhere. The vocabulary data (token bytes → ranks) is OpenAI's public o200k_base.tiktoken file — you can't learn GPT-4o's merges without GPT-4o's training data, and pretending otherwise would be dishonest. Verification: python/verify.py diffs this implementation against real tiktoken on thousands of inputs.

github.com/absolukie/tokenizer-from-scratch