"Chalk: The Self-Hosted AI Tutor That Refuses to Do Your Homework"


description: Chalk is a class-scoped AI tutor that won't hand over answers, and turns anonymous student questions into a week-by-week digests for professor on where the class is stuck — self-hosted so the data never leaves campus. Chalk goal is to reduce costs for students, increase AI tool fluency, and data privacy, professor IP ownership.

TL;DR — Students use Ai to due class why not have models that protect privacy, have guard rails, and overall

Chalk is a class-scoped tutor that refuses to do hand over answers, then turns every anonymous question into a weekly digest the professor can act on. We didn't call an API — we self-hosted the entire stack: our own model, our own database, our own deployment. The data never leaves campus. Built in 24 hours for, live right now.

Inspiration

Two things break every semester, and they're the same thing. Students ask an AI for help and get the answer handed to them — so they learn nothing, and they learn it privately, at 2am, where no one can help. Professors fly blind and discover the class was lost on the exam, weeks after the confusion started. The question a student is too embarrassed to ask in class is the most valuable signal a professor could ever get. Right now it disappears into a chatbot.

So we asked: what if the arrow ran the other way? Every feed pushes content at students. Chalk reverses it — questions flow back to the person who desires the course. That's the whole idea. Beyond the Feed.

What it does

Student view — a class-scoped chat with a hard guardrail: no handing AI questions to get fully finished assignments, only Socratic coaching (hints, leading questions, next steps). A token meter and per-student limits keep it fair.

Professor view — Professor digest clusters and representative prompts. It updates as questions arrive. Feed back over the last 7 days.

The gateway — A single sever the only thing that talks to the model and the database. It injects the guardrail, enforces limits, anonymizes every turn, and logs it.

The clever part is the memory pipeline. When a conversation ends, Chalk writes a compact two-line summary of what the student was actually asking. Each summary is attached to the class, never to the student. The professor can then generate a 7-day digest of the most common questions — and tell the difference between "homework 3, question 2," "the lecture didn't land," and "this topic was never covered." A chat log becomes a curriculum signal.

How we built it

Almost everything is JavaScript, front and back. One language across the stack.

  • Front end: React, component-based, two personas (student chat / professor dashboard), with the Canvas API for the interactive surface.
  • Back end: a JavaScript API server handles auth, class scope, token limits, and anonymization. The browser talks to the model and the DB only through POST requests — it never touches either directly.
  • Database: Postgres + TimescaleDB. Every question is a time-series event:
SELECT create_hypertable('prompt_events', 'created_at');

CREATE MATERIALIZED VIEW struggle_weekly
WITH (timescaledb.continuous) AS
SELECT time_bucket('1 week', created_at) AS week,
       topic_cluster_id, count(*) AS question_count
FROM prompt_events
GROUP BY week, topic_cluster_id;

The model: we host our own LLM with vLLM. PagedAttention manages the KV cache like virtual memory in an OS — memory paging for attention — which is what keeps throughput up when many students hit it at once.
Deployment: everything is infrastructure-as-a-service, on Vultr. Our own database, our own front-end deployment, our own model image, nginx in front, and a .tech domain whose DNS record points straight at our instance. No API we don't control. The only thing we don't own is the literal hardware — which is true of any large company too.
Privacy by architecture
Students never touch the model or the database. The university owns the gateway, the database, and the GPU. The model is a black box behind the gateway. Identity is stored separately from prompt_events — the event table has no foreign key back to a student. Clusters are computed on content, never on who.

This isn't a privacy policy — it's the network diagram.

Challenges we ran into
Self-hosting every layer is hard. We could have called Gemini and gotten good answers in an afternoon. But could we guarantee what happens to the data? Could we log every conversation the way we do? Could we set the context window the way we do? No. So we took the harder path: DNS, TLS, the database, and the GPU image are all ours. Deploying the Vultr instance and wiring it end to end was the single most complex part of the build — and the most worth doing.

Anonymity is the real design problem. We wanted to help students without the professor ever seeing an individual. The moment you log anything, students read "we measure engagement" as "we're watching you." Our rule became: the system can measure the room, not the person. Every metric has to be useful to the professor and impossible to reduce to a single student. If a signal can't pass that test, it doesn't ship.

Accomplishments that we're proud of
It's live. You can log in right now from your own computer with your own Google account and use it — real deployment, not a mockup.
The refusal guardrail actually refuses. On an eval set of {{N}} prompts, Chalk refuses to do the homework {{X}}% of the time across {{M}} subject areas — and every attempt is still logged as a signal.
The digest works. A question a student asks shows up on the professor's map anonymously in {{T}} seconds.
No data leaves campus. The privacy property isn't a promise; it's the topology.
What we learned
A data pipeline beats a prompt. The guardrail was the easy part. Turning a chat log into an anonymous, clustered, time-series curriculum signal was the real engineering.
Human-centered design matters more here. The person using the tool isn't always the person it measures.
Managing your own infra teaches you what the abstractions hide. We know what nginx is doing, what the context window is, and where every record lives — because we set all of it up ourselves.
AI must combine with human judgment, not replace it. The professor stays the decision-maker. Chalk just widens the view.
What's next for Chalk
Pilot one real class for a full semester.
Sharper clustering so topics map directly to the syllabus.
Week-by-week curriculum recommendations, not just the heatmap.
Stronger AI-fluency tooling for students — teaching them to learn with AI, not lean on it.
The point
Edtech's default move is to point AI at the student. Chalk points it at the classroom — a better view for the professor, more room for the student, and none of the surveillance. AI should widen the professor's view, not watch the student.


Built With

Share this project:

Updates

Submission history