Inspiration
I could have spent the hackathon building something more practical, like a corporate dashboard. Instead, I wanted to build something that captures what the modern software engineering experience is actually like: dependency hell, git panic, and crushing imposter syndrome.
So, I made a programming language for it. All variables are npm dependencies, assignment is a git push --force, every program has to start with a i use arch btw.
I wanted there to be a contrast. A stupid idea, but executed seriously. Joke languages usually consists of regex, but I wanted btw to have a functioning compiler underneath.
What it does
btw is a joke programming language that is built from dev memes + a real toolchain. You can try it with nothing installed at the playground linked in this submission.
- A checker that roasts you. Every error code is an HTTP status. Forget :wq and you get E408: "Error: program never exited. Classic Vim user." Type if and you get "if is a boomer conditional. Use vibe check." There are 51 exact messages under 31 codes.
- A Big O checker. Annotate a microservice with O(n) and it counts your nested loops. Lie, and it says "You said O(n), but this is O(n²). Skill issue."
- A load tester that checks the checker. btw loadtest runs a function for n = 8 to 1024, counts steps, and fits the slope. When the static checker is wrong, it says so: "The static checker was being optimistic."
- An interpreter (btw run) and a native x86-64 compiler (btw build). The same program becomes a 14 KB binary that prints the same bytes about 1,000 times faster.
- Git for variables. git revert x, git log x and git blame x work on every variable, in both backends.
- A language server for VS Code and Neovim, with live diagnostics, hovers, completion, go to definition, inlay hints and quick fixes like "Exit Vim" and "Install Arch".
- A browser playground running the real compiler on Pyodide, including a Two Sum challenge where passing 6/6 tests earns Accepted, but what percentile did you beat?!
How we built it
I wrote the spec before any code. I designed the syntax, the keywords and every error message myself, and pinned them down in a Language Spec, character for character, backticks included. A second Implementation Spec fixed the repo layout, the CLI, the core types and how each component fits together. Then I created the contract: there were 109 golden programs, each with its expected diagnostics, stdout, stderr and exit code. Those tests were the definition of done. With that in place, I used AI coding agents for most of the implementation. I split the work into task cards (lexer, parser, checker, Big O, interpreter, codegen, LSP, editors), and each agent owned one card, one branch and one notes file. A card wasn't finished until its golden tests passed. When an agent hit an ambiguity in the spec, the rule was to stop and write the question down, not invent behavior. I reviewed and merged every pull request. The stack:
- Python 3.12 and uv, with one runtime dependency (pygls, for the language server).
- A hand-written lexer and a recursive descent + Pratt parser that recovers at the next line or }, so half-typed code gets one squiggle, not a red file.
- A stack-machine code generator that writes x86-64 assembly as text (GNU as, Intel syntax, System V ABI). No LLVM. gcc only assembles it and links a 160-line C runtime.
- Pyodide in a Web Worker for the playground, so the real compiler runs in the browser and an infinite loop can't freeze the tab.
- GitHub Actions running every test on every pull request, and releases that ship btw and btw-lsp as standalone Linux executables (ELFs).
Challenges we ran into
- Making two backends agree to the byte. The interpreter is the reference, and every native binary has to match its stdout, stderr and exit code exactly. Python integers never overflow and Python's // and % round differently from x86, so every arithmetic result had to be wrapped to signed 64 bits and division rewritten to truncate toward zero.
- The one division x86 refuses to do. The minimum 64-bit number divided by -1 doesn't wrap on x86; idiv crashes the process with SIGFPE. I had to define that case in the spec and special-case it in the codegen.
- **Stack alignment. **System V requires rsp aligned to 16 bytes before every call. The code generator counts its own pushes, asserts that every expression leaves exactly one value and every statement leaves none, and uses the count to align. A test probe fails if any call is made misaligned.
- **Errors that don't cascade. **A single typo shouldn't produce 30 errors. Getting the parser to recover cleanly, and stopping errors like E503 from piling up, took several rounds.
- Keeping AI agents honest. Agents like to fill gaps with plausible behavior. The fix was structural: a spec precise enough to test against, golden tests nobody was allowed to edit, and a rule to ask instead of guessing. The early spec still had ambiguities that only surfaced once the lexer and parser existed.
- A stale playground. After one deploy, browsers kept running the previous build for up to 10 minutes because of GitHub Pages caching, which turned every new curl into an E404. Every file the page loads now carries a hash of the build in its URL.
Accomplishments that we're proud of
- A real native compiler with no LLVM. btw writes its own x86-64 assembly. The native FizzBuzz is 14 KB including the runtime, and a prime-counting benchmark drops from 14.9 s in the interpreter to 12.4 ms as a binary, about 1,200 times faster.
- Zero bytes of difference. Every golden program that compiles runs on both backends and must match exactly. Seeded differential fuzzing generates random well-typed programs and holds both backends to the same rule.
- 1,778 passing tests, including 109 hand-written golden programs, run by CI on every pull request in about 25 seconds.
- A Big O checker that doubts itself. Static inference handles nested loops, O(log n) halving loops, calls and recursion, and btw loadtest measures what it can't prove. It catches the static checker being wrong in both directions.
- Funny all the way down. The jokes go into places nobody has to look: hovers, quick fixes, exit codes (128 like git, 52 like curl), and comments in the generated assembly like jz .Lloop1_end # 404: stop scrolling. Even the 6-parameter limit is a joke about the ABI: a seventh parameter is E413, "that's a monolith."
- Four releases in two days, from first commit on October 3 to v0.4.0 on October 4, with a changelog, standalone executables and a live playground.
What we learned
- With AI agents, the spec is the product. The agents were only as good as the spec and tests they worked against. Hours spent pinning down exact messages and writing golden tests by hand saved far more hours of review. Where the spec was vague, the code was wrong.
- A second implementation is the best test. Running every program through the interpreter and the native backend caught bugs neither would have found alone, like the min / -1 crash.
- How a compiler fits together end to end: lexing, Pratt parsing, error recovery, scope and type checking, stack-machine codegen and the System V calling convention, all in a project small enough to hold in my head.
- Static analysis has limits, and you can say so. The Big O checker is a heuristic, not a proof. Pairing it with a load test that measures instead of guessing was more honest, and funnier.
- Comedy needs precision too. A joke error message only works if it fires at exactly the right moment, on exactly the right span. Treating the jokes as a spec made them land.
What's next for BTW
- Show the hover in the playground under the cursor in a line below the editor.
- A second challenge: Maybe something like the power of two.
- Load-testing inside the browser: It's already pretty slow in the terminal, so I'll have to figure out a way to optimize this in the browser.
- A CTF challenge based on
btw. If you're interested in CTFs and I do post a challenge, check https://home.sebastianyeo.dev/challenge/ !!
Do I even use arch? btw
Built With
- github-actions
- pylsp
- pyodide
- python
Log in or sign up for Devpost to join the conversation.