7 ms·
Show HN: TokenDagger – A tokenizer faster than OpenAI's Tiktoken
TokenDagger is a drop-in replacement for OpenAI’s Tiktoken (the tokenizer behind Llama 3, Mistral, GPT-3.*, etc.). It’s written in C++ 17 with thin Python bindings, keeps the exact same BPE vocab/special-token rules, and focuses on raw speed.
I’m teaching myself LLM internals by re-implementing the stack from first principles. Profiling TikToken’s Python/Rust implementation showed a lot of time was spent doing regex matching. Most of my perf gains come from a) using a faster jit-compiled regex engine; and b) simplifying the algorithm to forego regex matching special tokens at all.
Benchmarking code is included. Notable results show:
- 4x faster code sample tokenization on a single thread.
- 2-3x higher throughput when tested on a 1GB natural language text file.
- isjustintime 1y agoVery cool. We use Tiktoken and I'd love to see the performance impact. Pretty great decision to make it drop-in compatible.
- justinhj 1y agoI've been playing with tokenization too. Starting from Kaparthy's Python minbpe I set myself the task of training a tokenizer on wikitext (500mb) in a reasonable time. I got the C++ version down to about 50 minutes compared to the original Python code (estimated) several months. Haven't really spent much time looking at encode and decode but I plan to incorporate these regex modifications when I do! https://github.com/justinhj/minbpe-cc https://github.com/justinhj/minbpe-cc
- chrismustcode 1y agoThere’s something beautiful about creating a drop in replacement for something that improves performance substantially. ScyllaDB comes to mind
- matthewolfe 1y agoAgreed. I figured nobody would use it otherwise.
- parhamn 1y agoPut it in there readme & description. It's a big selling point.
- matthewolfe 1y agoThanks, I clarified it.
- pvg 1y agoTo be fair, many people have token stabbing needs.
- npalli 1y agoKudos, I think (in the short term at least) there is a large amount of perf. optimization to be found by coding parts of the whole AI/ML infrastructure in C++ like this one, not as a rewrite (god no!) but drop in and fix key bottlenecks. Anytime I see someone (seems Chinese engineers are good at this) put something out in C++, good chance some solid engineering tradeoffs have been made and dramatic improvement will be seen.
- matthewolfe 1y agoAgreed. A former mentor of mine told me a nice way of viewing software development: 1. Make it work. 2. Make it fast. 3. Make it pretty. Transformers & LLMs have been developed to a point where they work quite well. I feel as though we're at a stage where most substantial progress is being made on the performance side.
- diggan 1y agoHeh, seems people I've been learning from been biased away from beauty, as I know that as "Make It Work, Make It Right, Make It Fast".
- abybaddi009 1y agoWhat's the difference between make it work and make it right? Aren't they the same thing?
- konsalexee 1y ago> simplifying the algorithm to forego regex matching special tokens at all Does that mean there could be cases with less quality in terms of tokenization?
- matthewolfe 1y agoThe output should be identical, assuming no bugs. The Tiktoken implementation takes a collection of all special tokens upon initialization and compiles them into a regex by joining them with `|` [0]. Then the actual encoding process checks for matches on this expression. Models like Llama 4 define a list of 1,135 special tokens. Notably, 1,115 of those are "reserved" special tokens! So this yields a huge regexp of special tokens that shouldn't be considered at all. TokenDagger does not do this. Instead, simple string matching is used. This works because we don't need to consider the entire special vocabulary every time. The caller of `encode` must explicitly define which special tokens should be considered [1]. So it's faster to check against the much smaller list we _know_ is being used. [0] https://github.com/openai/tiktoken/blob/main/src/lib.rs#L476 https://github.com/openai/tiktoken/blob/main/src/lib.rs#L476 [1] https://github.com/openai/tiktoken/blob/main/tiktoken/core.py#L79 https://github.com/openai/tiktoken/blob/main/tiktoken/core.p...
- anonymoushn 1y agoIsn't this incorrect? If the user doesn't specify what to do with almost all of the special tokens, you still must detect them so you can raise an error.
- manishsharan 1y agoIs there a tokenizer someone can recommend for code ? I have tried CodeBert but maybe I am using it wrong as my results with it were pretty bad.
- fkyoureadthedoc 1y agoWould be cool to see WASM bindings for this here https://github.com/dqbd/tiktoken https://github.com/dqbd/tiktoken Or maybe even your speedups from "b" in the pure js implementation
- p0 1y agoHow does this compare to the BPE crate [1]? Its main selling point is support for incrementally re-tokenising text, but it's also faster than tiktoken. [1] https://crates.io/crates/bpe https://crates.io/crates/bpe
- matthewolfe 1y agoI'm working on incremental re-tokenizing next. Then I'll run some benchmarks against this crate too.
- frabcus 1y agoIs there any way we can get local tokenizers for other LLMs? e.g. Gemini only offer a remote API for their tokenizer. Is it proprietary? Could we infer the token mapping somehow efficiently by making lots of calls?
- weberer 1y agoI thought Gemini used SentencePiece https://github.com/google/sentencepiece https://github.com/google/sentencepiece
- Deathmax 1y agoGemini uses SentencePiece [1], and the proprietary Gemini models share the same tokenizer vocabulary as Gemma [2, 3, 4]. Out of the large proprietary western AI labs (OpenAI, Anthropic, Google), only Anthropic with Claude 3 and newer lack local tokenizers. [1] https://github.com/google/sentencepiece https://github.com/google/sentencepiece [2] https://github.com/googleapis/python-aiplatform/blob/main/vertexai/tokenization/_tokenizer_loading.py https://github.com/googleapis/python-aiplatform/blob/main/ve... [3] https://storage.googleapis.com/deepmind-media/gemma/gemma-2-report.pdf https://storage.googleapis.com/deepmind-media/gemma/gemma-2-...: "We inherit from the large Gemini vocabulary (256k entries)." [4] https://storage.googleapis.com/deepmind-media/gemma/Gemma3Report.pdf https://storage.googleapis.com/deepmind-media/gemma/Gemma3Re...: "We use the same tokenizer as Gemini 2.0."
- matthewolfe 1y agoA lot of model-specific tokenizers have reference implementations ([0], [1]). Underlying them is a core algorithm like SentencePiece or Byte-pair encoding (BPE). Tiktoken and TokenDagger are BPE implementations. The wrapping "tokenizer" mostly deals with the quirks of the vocabulary and handling special tokens. For this project, I think there is value in building some of these model-specific quirks into the library. Could see some minor performance gains and generally make it easier to integrate with. It's probably not too much work to keep up with newer models. Tokenizers change much less frequently. [0] https://github.com/meta-llama/llama-models/blob/01dc8ce46fecf06b639598f715efbb4ab981fb4c/models/llama4/tokenizer.py https://github.com/meta-llama/llama-models/blob/01dc8ce46fec... [1] https://github.com/mistralai/mistral-common/tree/main/src/mistral_common/tokens/tokenizers https://github.com/mistralai/mistral-common/tree/main/src/mi...
- pama 1y agoCool. Would it be possible to eliminate that little vocab format conversion requirement for the vocab I see in the test against tiktoken? It would be nice to have a fully compatible drop in replacement without having to think about details. It also would be nice to have examples that work the other way around: initialize tiktoken as you normally would, including any specialized extension of standard tokenizers, and then use that initialized tokenizer to initialize a new tokendagger and test identity of results.
- matthewolfe 1y agoAh good catch. Updating this right now.
- matthewolfe 1y agoAlright, 0.1.1 should now be a true drop-in replacement. I'll write up some examples soon.
- janwilmake 1y agoYou know what's also faster to roughly get the amount of tokens? string.length/5
- deleted 1y ago[deleted]
- _flux 1y agoIt is not helpful in actual tokenization, though.
- EGreg 1y agoWhat about pairing this with BigBird and Mamba?
- pamelafox 1y agoJust curious whether it's possible to push any of your performance improvements to tiktoken itself?
- matthewolfe 1y agoI probably will. Was hesitant initially, because adding PCRE2 as a dependency might cause issues to existing projects. I believe this was discussed briefly in a closed PR with other performance improvements.
- b0a04gl 1y ago[dead]
- kevmo314 1y agoNice work! I tried something similar a while back ago: https://github.com/kevmo314/tokie https://github.com/kevmo314/tokie The takeaway I also found was that the running cost was really dominated by pretokenization (the regex). It's cool to see that you found a faster way to run the regex, but have you tried comparing the performance of just swapping out the regex engine and leaving the actual BPE to tiktoken? I wonder if that is upstreamable?
- 22c 1y agoThere is at least some awareness already when it comes to the performance of the regex engine: https://github.com/openai/tiktoken/blob/main/src/lib.rs#L95-L110 https://github.com/openai/tiktoken/blob/main/src/lib.rs#L95-...
- matthewolfe 1y agoCool! I've reached out to the guy who maintains Tiktoken to talk about this.
- polynomial 1y agoJust to note that Tiktoken is still the tokenizer behind the GPT-4x series, it just uses a different token model. (Post only says GPT-3, implying they were using something else for subsequent iterations.)
- silentsea90 1y ago"I’m teaching myself LLM internals by re-implementing the stack from first principles." - curious what resources you're using? Any books or courses, or just building it straight up? Great work!
- deleted 1y ago[deleted]
- matthewolfe 1y agoModal's GPU glossary is a good overview about how GPUs work [0]. Karpathy's LLM overview is a good high level overview on LLMs [1]. 3b1b's video (and subsequent videos) on transformers was excellent at helping me understand the math at a high level [2]. This matrix multiplication optimization worklog helped me understand writing better CUDA (not for beginner intro though) [3]. During this process I also asked ChatGPT a lot of questions. I'm definitely open to suggestions about "how to learn" with all the new tools we have. I felt this has not been straightforward to figure out. [0] https://modal.com/gpu-glossary https://modal.com/gpu-glossary [1] https://www.youtube.com/watch?v=7xTGNNLPyMI https://www.youtube.com/watch?v=7xTGNNLPyMI [2] https://www.youtube.com/watch?v=wjZofJX0v4M https://www.youtube.com/watch?v=wjZofJX0v4M [3] https://siboehm.com/articles/22/CUDA-MMM https://siboehm.com/articles/22/CUDA-MMM
- matrix2596 1y agois is possible for your tokenizer to give different tokenization ever then openai tokenizer? i am asking because there are multiple ways to tokenize the same string?? sry if i am mistaken
- matthewolfe 1y agoShould be the same. Both use Byte-Pair Encoding (BPE) as underlying algo.
- Tiberium 1y agoCan you also compare the performance with https://github.com/huggingface/tokenizers/ https://github.com/huggingface/tokenizers/? Would be helpful, since the benchmark in the tiktoken readme seems to be very outdated.
- binarymax 1y agoAnecdotally I've always found tiktoken to be far slower than huggingface tokenizers. I'm not sure why, as I haven't dug into tiktoken, but I'm a heavy user of HF's rust tokenizers
- superlopuh 1y agoCan someone familiar with performance of LLMs please tell me how important this is to the overall perf? I'm interested in looking into optimizing tokenizers, and have not yet run the measurements. I would have assumed that the cost is generally dominated by matmuls but am encouraged by the reception of this post in the comments.
- serjester 1y agoTokenizing text is ridiculously small part of the overall computation that goes into serving a request. With that said if you’re doing this on petabytes of data, never hurts to have something faster.
- odyssey7 1y agoA language that isn’t memory-safe can definitely hurt. AI needs more security, not less.
- refibrillator 1y agoTokenization is typically done on CPU and is rarely (if ever) a bottleneck for training or inference. GPU kernels typically dominate in terms of wall clock time, the only exception might be very small models. Thus the latency of tokenization can essentially be “hidden”, by having the CPU prepare the next batch while the GPU finishes the current batch.
- benreesman 1y agoTokenization performance is complicated, but your guidepost is that the institutions with the resources and talent to do so choose to write extremely fast tokenizers: sentencepiece and tiktoken both pay dearly in complexity (particularly complexity of deployment because now you've got another axis of architecture-specific build/bundle/dylib to manage in addition to whatever your accelerator burden always was: its now aarch64 cross x86_64 cross CUDA capability...) Sometimes it can overlap with accelerator issue, but pros look at flame graphs: a CPU core running the AVX lanes hard isn't keeping the bus fed, million things. People pre-tokenize big runs all the time. I don't know why this thread is full of "nothing to see here", this obliterates the SOTA from the money is no object status quo: I'd like to think better of the community than the obvious which is that C++ is threatening a modest mindshare comeback against a Rust narrative that's already under pressure from the explosion of interest in Zig. Maybe there's a better reason.
- luppy47474 1y ago[flagged]
- sheerun 1y agoNow that byte-patch-level embeddings are discovered?
- semiinfinitely 1y agoI'm relieved to see that its not written in rust
- matthewolfe 1y agohaha, I thought about it.
- singularity2001 1y agothis is still the outdated architecture without special tokens for numbers like out-of-vocab tokens like NUM_FLOAT(3.1415) right?