Inspiration

Every semester, students at UMBC (and everywhere else) ask the same three questions and get vague answers: Am I actually on track? What happens if I take this course instead of that one? What does someone with my exact record usually end up doing? Academic advisors are stretched thin, degree audits are dense PDFs nobody enjoys reading, and most "AI advisor" tools just have a language model improvise plausible-sounding numbers.

We wanted to build the opposite of that: an agent that speaks to you like an advisor, but where every number it says out loud was computed by a trained model or pulled from a real record — never invented.

What it does

COOKED reads a student's real degree audit (upload a PDF/image, paste the text, or pick a sample profile) and answers, by voice or text:

  • Am I cooked? — a trained risk model's score, with the decision line it's measured against
  • What if I work 10 more hours a week, or take 15 credits instead of 12? — a live what-if simulator
  • When do I actually graduate? — a time-to-degree model with a p25–p75 range, not a single fake-precise number
  • Who are my "twins"? — real alumni matched on the student's own trajectory (course load, withdrawals, work hours, major), with their real graduation outcomes
  • How do I get un-cooked? — the smallest realistic change (credits, work hours, course swap) that moves the needle, computed by re-running the actual model
  • Show me the Watchtower — an advisor-facing view scoring every current student in the cohort against the same model, for a bird's-eye risk picture

Every answer is voice-driven end to end (ElevenLabs conversational agent), and every number on screen carries a tool_result_id back to the exact computation that produced it — the UI's own provenance layer literally refuses to render a sentence with an untraced digit in it.

How we built it

  • Data: HackUMBC's synthetic UMBC dataset — 3,200 alumni, 1,800 current students, 140,000+ course transcripts, a 72-course catalog — loaded into TimescaleDB on Tiger Cloud.
  • Models: for each prediction task (risk, time-to-degree, career outcome, salary), we ran a real "model arena" — four candidate families (baseline, logistic regression, random forest, gradient boosting) trained on identical temporal train/test splits, with a champion picked by a pre-registered validation rule. The full arena report (every candidate's held-out AUROC, not just the winner's) is served live and rendered as a real bar chart, so a judge can see the model was actually tested, not just claimed.
  • Audit parsing: a deterministic PDF/text parser reads UMBC's real Oracle Analytics audit export format directly from the text layer, with a vision-model fallback (Gemini/Claude) for scanned documents, and an honest failure mode — if a parse can't produce anything the model would actually score, it says so instead of showing a confident "0 credits, cooked" result.
  • Voice: a live ElevenLabs conversational agent (Claude Haiku-routed tool calls) is the primary interface everywhere in the app, not a bolted-on chatbot.
  • Backend: FastAPI + scikit-learn, with a §8.4 "provenance rule" enforced in code: narration is built from segments (plain text, which may hold no digits, and value tokens tied to a recorded tool result) — a language model is structurally incapable of writing a number into what the user hears.
  • Frontend: Next.js 16 / React 19, a scene-deck UI (one idea per screen, voice-first, no page scroll), real charts (bar, range, scatter, ridgeline) driven directly by API payloads.

Challenges we ran into

  • Keeping the provenance rule airtight. Partway through, we found narration text with real-looking but hardcoded numbers (a fake "$75,000 median salary," fake feature-importance percentages) baked directly into voice scripts — copy that had drifted in from an earlier draft. Our own anti-fabrication guard caught it and silently blocked those answers in production, which took real debugging to trace back to the actual root cause instead of just monkey-patching the symptom.
  • Two ElevenLabs voice identities. A "coach" narration mode used a different ElevenLabs voice ID than the default "narrator" for one specific answer type — invisible in code review, obvious the moment you actually listened to two different questions back to back.
  • Real vs. synthetic-realistic audits. UMBC's real degree-audit export format (Oracle Analytics) is nothing like a clean CSV — reconciling transfer credits, in-progress courses, and repeated courses against the model's expected feature shape took several honest-failure iterations before we trusted it not to silently misread a real upload.

Accomplishments we're proud of

  • A concrete, code-enforced rule against AI hallucination in a domain (a student's own academic risk) where a confidently wrong number is actually harmful — not just embarrassing.
  • A model evaluation that shows its work: four real trained candidates per task, held-out test metrics for all of them, a documented promotion rule, not a black box "trust us."
  • A voice interface that's the actual primary way to use the product, not a demo gimmick layered on top of a dashboard.

What we learned

Provenance discipline is easy to state and hard to keep — every new feature is a new place for a plausible-sounding fabricated number to sneak back in, and the only real defense is enforcing it in code (a runtime check that blocks output), not just in code review.

What's next

  • Real per-feature importance (SHAP) instead of the placeholder we deliberately removed rather than fabricate.
  • Extending the twin-matching model to more majors and a richer feature set.
  • A staff-facing alerting flow off the Watchtower's real-time cohort scoring.

Built With

Share this project:

Updates

Submission history