Inspiration
We'd all used chatbots to practice a language, and they all had the same hole: you make the same mistake in week three that you made on day one, and the bot cheerfully corrects it again as if it were the first time. Nothing accumulates. Meanwhile the tools that do track what you know — Anki and friends — need you to build decks by hand, and the cards are somebody else's idea of what's hard. We wanted the conversation itself to be the input, and a real model of the learner to be the output.
What it does
Lingua is a chat-based language tutor that builds a persistent model of your weaknesses and schedules them for review. You pick a language at sign-up and the tutor speaks only that one. As you chat, it corrects you inline. At the end of a lesson, a background pipeline reads the whole transcript, extracts the errors you actually made, and merges them into a durable learner model in SQLite — where an FSRS-6 scheduler assigns each weakness a due date.
The crucial part is that errors are deduped by meaning, not by text. "mit der Bus" and "nach die Schule" are the same weakness wearing different words. Each error is embedded from its underlying rule and matched against existing items by cosine similarity, so a recurrence bumps an existing item's lapse count instead of creating a duplicate. That's what turns a log of mistakes into a model of a learner.
Practice is built from that model. Five tools render as inline widgets mid-conversation — speed rounds, single MCQs, correct-the-sentence drills, reading passages, and writing prompts — and the tutor picks one when you ask to practice, or you force one from the composer's mode pill. Answers feed straight back: a miss inserts or relapses an item, a hit is an FSRS review. Writing submissions go through the same error extractor as end-of-session transcripts, so free-form writing becomes tracked weaknesses too. A metrics dashboard plots errors per 100 words over time.
How we built it
A Next.js frontend talks to a FastAPI backend that holds all the tutor logic and streams from Qwen (Alibaba DashScope, via its OpenAI-compatible API). The API key never leaves the backend.
The backend runs two loops. The live loop streams chat over Server-Sent Events, where the model can emit either text deltas or a tool call; a tool call is executed server-side and its result is pushed down the same stream as a widget frame the React client renders inline. The consolidation loop fires on session end as a background task — extract errors, embed them, dedupe-merge, schedule — so the learner never waits on it.
Auth is stdlib-only: PBKDF2-hashed passwords and HMAC-signed bearer tokens, no session table. FSRS-6 is implemented as pure functions, which made it straightforward to test. The store is deliberately written behind a thin interface so SQLite can be swapped for Postgres + pgvector without touching callers.
Challenges we ran into
Deduping by meaning needed calibration, not just cosine. Too low a threshold and unrelated errors merge into one useless mega-item; too high and every recurrence looks new and the learner model never converges. Getting MERGE_THRESHOLD right, and deciding to embed the rule (concept: note) rather than the example sentence, took real iteration.
The tutor kept doing the tools' job. Given a tool for quizzing, the model would happily write the quiz questions itself in prose instead of calling it. Fixing that was prompt surgery — explicit "do NOT write the questions yourself, this tool builds and renders its own content" instructions per tool.
Repeats inside a single session were the subtle one. Every generator is a stateless one-shot LLM call, and the correct-the-sentence builder always picks the first usable due item — so asking twice in one lesson returned a byte-identical exercise. We built a per-session ledger tracking concepts, item ids, topics, and normalized surface text. It gets consulted three ways: filtering generated content, injecting already-used topics into the author prompt as a "don't write about these" clause, and writing a short note into the chat history so the tutor stops asking for a repeat in the first place. Normalization mattered more than expected — the model writes blanks as runs of underscores and the count varies between generations, so the same question slipped through as "new" on a different underscore count until we collapsed them.
Streaming plus tool calls plus grading is a lot of state. Recording served content only after the widget frame is out, so a client disconnect mid-yield doesn't mark an exercise as seen when the learner never saw it, is the kind of bug that only shows up as a mysteriously skipped question.
Accomplishments that we're proud of
The learner model genuinely works — you can watch an item's lapse count climb across sessions as you keep making the same mistake in different words, then watch its due date stretch out as you stop. Semantic dedupe is the technical core and it earns its complexity.
We're also proud of the anti-repeat architecture. Rather than bolting a check onto each generator, the ledger keeps writers and readers in one module so both share a single normalization rule by construction, and an import-time assertion forces every tool to have a recorder and an exhaustion message. A tool that runs dry says so and suggests a different activity instead of dead-ending the learner.
What we learned
We learned that embeddings are a design decision, not a library call: what you embed determines what "the same thing" means downstream. We learned how much of agent reliability is prompt-level contract-setting, and how much of the rest is deterministic validation around the model rather than trust in it — nearly every generator here has a filter behind it that throws away malformed or repeated output. And implementing FSRS-6 by hand taught us why the state machine and the interval math are worth separating.
What's next for AILanguage tutor
Postgres + pgvector so the learner model scales past a local file, and per-user FSRS parameter optimization — the algorithm supports fitting the schedule to an individual's actual review history, which is the natural next step once there's enough of it. Beyond that: persistent conversation memory (it's currently in-process and resets on restart), history trimming for long sessions, speaking and listening exercises
Log in or sign up for Devpost to join the conversation.