Inspiration

https://platform.openai.com/tokenizer https://gpt-tokenizer.dev

What it does

The user can input content into an advanced VSCode-like text editor (Monaco, so literally the same base). Every keystroke will trigger the content to get retokenized with multiple tokenizers in parallel, depending on which one the user has selected. It’s great for comparing tokenization efficiency between models and becoming a better tokenization-aware prompt engineer in general.

How we built it

It’s a modern single-page React app, nothing exotic.

I started with a hand-written web app scaffold I actually spent a few days on. Then I handed it off to a bunch of models to flesh out core functionality. GPT, DeepSeek, Kimi, GLM. Then another GPT agent reviewed every candidate’s submission and merged them into a definitive edition that incorporated the best parts of each. Then I added another few days of handwriting (polishing, cleaning up, UI design) and finally gave the project into GPT Sol’s hands to add minor additional features and keep the supported tokenizer list up to date.

Challenges we ran into

The web app is massive in size because of the tokenizer vocabularies. No good approach to tackle this found so far.

Accomplishments that we're proud of

It has a unique and opinionated look and feel directed toward power users (which is rarely chosen target audience these days), the tokenizers run fast and reliably and this actually turned out to be an S-tier tool in regards to productivity.

I like it so much myself that I write my prompts in it and then copy them into the chat inputs they are meant for.

What we learned

Tokenizer vocabularies are really odd.

What's next for TokShow

more models, better model selection interface and better saving and loading of UI state via URL search parameters and local storage

Built With

Share this project:

Updates

Submission history