RavenMap
An AI academic advisor that never lies about your degree
Carleton University students plan their entire academic career on top of two things: a static PDF program map and a chatbot that will confidently make up an answer if it does not know one. RavenMap replaces both with a system where every number is computed, every rule is cited, and the AI assistant is structurally forbidden from inventing either one.
Inspiration
I am a Software Engineering student at Carleton, and this project started from a genuinely irritating personal problem. Carleton publishes real degree requirements, real prerequisite chains, and a real co-op work-study calendar, but none of it is connected. To answer a question as simple as "can I still graduate on time if I take a lighter course load next term," a student has to manually cross-reference a static program map PDF, their own unofficial transcript, and a separate co-op sequence page, then do the arithmetic themselves and hope they did not miss a prerequisite buried three terms back.
The obvious modern answer is to ask an AI. So I tried it. General-purpose chatbots handled the conversational part fine and then quietly failed at the one part that actually mattered: when asked to calculate a CGPA or confirm a prerequisite, they produced a plausible, confident, wrong number. For a tool that a real student might use to decide whether to drop a course or apply for co-op, a confidently wrong answer is worse than no answer at all.
That gap, real academic rules with no computation behind them on one side, and a computational AI with no grounding in real rules on the other, is what RavenMap was built to close.
What it does
- Academic Health dashboard: real, computed CGPA, degree completion percentage, and a list of structural risk factors, all derived deterministically, never estimated by a model.
- Degree audit ingestion: uploads a real Carleton uAchieve PDF export and parses it deterministically, not through an AI reading the document.
- Interactive degree tree: a node-based visualization of a student's entire program, laid out exactly as Carleton's own official program map positions it by year and term.
- Autonomous multi-year planner: a genuine constraint solver that proves an optimal schedule to graduation or proves that none exists, respecting prerequisites, real term offerings, credit limits, and workload balance.
- Term eligibility checker: what a student can actually register for next term, respecting live prerequisite data and real co-op status.
- Instant grade simulation: exact CGPA impact of a hypothetical grade change or a hypothetical added course.
- AI Advisor: a conversational agent that answers in plain language but is never allowed to calculate anything itself; every fact it states is either a deterministic tool result or a cited calendar quote.
- Policy grounding: hybrid retrieval over the real Undergraduate Calendar and Academic Regulations, so a policy answer is a quotation, not a recollection.
- Real co-op modeling: Carleton's actual, often irregular, work-study sequences sourced per program and per admission term directly from Carleton's own co-op office.
Underneath all of it sits one rule that shaped every design decision in this system: the AI layer narrates, it never calculates.
Academic Health dashboard
An Academic Health Report showing a student's real, computed CGPA (calculated from their actual course completions, not just restated from an uploaded document), degree completion percentage, and a list of structural risk factors: real blocking issues such as a mandatory course whose prerequisite chain has not been started yet, computed the same deterministic way every time, never guessed by a model.
Degree audit ingestion
Students upload their real Carleton uAchieve degree audit export, and it is parsed deterministically with a regex-based parser tuned to that document's exact fixed-width course-attempt format, not handed to an AI to "read." The parser also understands Carleton's own real edge cases: repeated attempts, forfeited courses, and a PDF text extraction process that does not preserve the audit's visual table order.
Degree Progress tree
An interactive, node-based visualization of a student's entire program, built with a real graph-rendering library, showing mandatory requirements, elective groups, and breadth requirements laid out exactly as Carleton's own official program map positions them by year and term.
Next Term and Simulator
Two connected tools. The first checks real term eligibility: what a student can register for in a specific upcoming term, respecting live prerequisite data, real course offering history, and whether that term is actually a co-op work term for that student. The second is an autonomous multi-year planner backed by a genuine constraint solver (Google OR-Tools CP-SAT), which models the student's entire remaining degree as a set of hard constraints, prerequisite graphs, real term offerings, credit limits, and workload balance, then either mathematically proves an optimal schedule to graduation or proves that no valid schedule exists under the given assumptions. It does not guess. It solves.
Grade simulation
Instant, exact recalculation of CGPA impact for a hypothetical grade change or a hypothetical added course, using the same deterministic engine that powers the dashboard.
AI Advisor
A conversational agent, built on LangChain and Google Gemini, that answers questions in plain language but is never allowed to do the actual thinking itself. Every CGPA, every prerequisite check, every multi-year schedule is produced by a deterministic tool call underneath; the model's only job is choosing which tool to call and narrating the result honestly, labeling every claim as high confidence (grounded in a real tool result) or low confidence (general knowledge), and citing the exact calendar section behind any policy statement.
Policy grounding
Behind that advisor sits a hybrid retrieval system over the real Carleton Undergraduate Calendar and Academic Regulations: exact course-code matching for precise lookups, combined with Gemini embedding similarity search over a vector database for open-ended policy questions. When the advisor states a rule, it is quoting real calendar text, not recalling it from general training data.
Real co-op modeling
Carleton's co-op work-study sequences are far stranger than "work every summer." Some programs split work terms non-consecutively across a student's degree in ways that look almost arbitrary until you see the department's own published sequence. RavenMap sources these sequences directly from Carleton's own co-op office, per program and per admission term, rather than assuming a single generic pattern for every student.
How we built it
We designed RavenMap with a real production-style architecture, split cleanly across a frontend, a backend, an AI orchestration layer, and a data layer that is treated as seriously as the code.
Frontend
- Framework: React 18, built and served with Vite
- Language: TypeScript throughout
- Styling: Tailwind CSS
- Server state: TanStack Query
- Routing: React Router
- Visualization: a dedicated graph-rendering library for the interactive degree tree
- Deployment: Vercel
The frontend is built to make a genuinely complex domain feel calm: a student should be able to glance at the dashboard and understand their real standing in seconds, not parse a spreadsheet.
Backend
- Framework: FastAPI, fully asynchronous end to end
- ORM: SQLAlchemy 2.0's async engine
- Database: PostgreSQL, with the pgvector extension enabled directly on the same database
- Migrations: Alembic
- Deployment: Railway
Running embedding search on the same PostgreSQL instance as the relational degree data, instead of standing up a separate vector store, keeps a student's academic record and the calendar data grounding their AI advisor's answers in one consistent, transactional place.
Authentication
- Provider: Supabase Auth
- Verification: the backend independently verifies every JWT against Supabase's own published JWKS endpoint, with no shared secret stored anywhere in the backend
AI orchestration
- Framework: LangChain's agent framework, wrapping Google's Gemini models
- Reliability: a genuine multi-model fallback chain, so a quota-exhausted or briefly overloaded Gemini tier automatically fails over to the next tier instead of failing the student's turn
- Security: every tool exposed to the model is scoped server-side to the authenticated student, so there is no parameter the model could fill in, correctly or by mistake, that would let it see another student's record
Retrieval-augmented generation
- Embeddings: Google's Gemini embedding model
- Storage and search: pgvector, using cosine similarity
- Retrieval strategy: exact course-code matching first, semantic similarity search as the fallback, both scoped to whichever calendar year actually governs the asking student
Constraint solving
- Engine: Google OR-Tools CP-SAT
- Model: courses, terms, prerequisite edges, credit caps, and workload rules are all encoded as hard constraints
- Objective: a strategy-weighted blend of predicted grade performance and a student's stated topic or career interest
- Output: a proof, either a concrete optimal schedule or a formal statement that none exists given the current constraints, never a heuristic dressed up as one
Data
Every seeded program's course catalog, prerequisites, preclusions, and requirement structure was built from two independently authoritative sources: Carleton's official visual program map for that exact program and calendar year, cross-checked line by line against the real Undergraduate Calendar's own course descriptions and prerequisite text, catching real discrepancies between the two along the way.
Infrastructure and deployment strategy
RavenMap deliberately does not run its own Kubernetes cluster. Both Vercel and Railway already build and run this application as containers behind the scenes, and already provide managed autoscaling, zero-downtime deploys, and automatic TLS, which is exactly what a self-managed Kubernetes control plane exists to provide. Standing up and operating our own cluster, control plane, node pools, ingress, service mesh, on top of that only pays for itself once there are enough independently scaling services, or high enough sustained traffic, to justify owning the orchestration layer directly instead of renting it through a managed platform. At this project's current scale and budget, that trade-off runs the other way: the operational cost of a self-managed cluster would be pure overhead with no corresponding benefit.
The system is still architected as two independently deployable services with a clean HTTP boundary between them and no shared state outside the database, specifically so that decision stays reversible. If traffic, team size, or a move to a multi-service architecture ever justified the operational cost, migrating either service into a container orchestrated by Kubernetes would be a deployment target change, not a re-architecture.
Challenges we ran into
Making hallucination structurally impossible, not just unlikely. A well-worded prompt is not enough to guarantee an AI advisor never fabricates a number. We built an architecture where the model has no path to a number except through a deterministic tool call underneath it, then stress-tested that boundary adversarially, deliberately forcing tool failures, to confirm the model reports an honest error instead of inventing a plausible answer to fill the silence.
Reverse-engineering Carleton's real co-op work-study sequences. Co-op scheduling at Carleton is not "work every summer." Several programs split work terms non-consecutively in ways that only make sense against the department's own published sequence. We sourced and encoded the real sequence for every program and admission term directly from Carleton's co-op office, so the planner reflects each program's actual structure instead of one generic assumption applied to all of them.
Deriving a real per-program course load instead of a fixed constant. Course load naturally differs by how a program distributes its own requirements across a degree. We built a function that computes each program's real typical per-term load directly from its own published requirement data, so the constraint solver models a front-loaded program differently from one that spreads its workload evenly, grounded in that program's own real map rather than a single assumed number.
Generalizing structural patterns across every seeded program. A two-term capstone project exists in nearly every engineering program, each under a different course code. We built detection based on the real structural shape of a capstone requirement rather than any single course number, so the same logic correctly scales across every program in the system without a special case per program.
Verifying elective category correctness against real Carleton data. A generic department-prefix match can look structurally reasonable and still be topically wrong, for instance resolving a biomedical elective slot against an unrelated economics course. Catching this required direct testing against real seeded data for every program, and led to rebuilding every elective category to resolve against a verified course list or an explicit departmental rule instead of a loose text match.
Treating third-party AI provider failure as a first-class design case. Quota exhaustion and transient server overload are genuinely different failure modes and need different handling. We built a real multi-model fallback chain so either kind of failure fails over quickly to another Gemini tier, and identified a redundant internal retry layer that could otherwise silently add up to a minute of latency before that fallback ever ran.
Optimizing conversation context for speed and reliability at the same time. A growing chat history resent on every turn slows every later turn and dilutes the model's own adherence to its system prompt the longer a conversation runs. Capping how much history gets resent measurably improved both response time and correctness in a single change.
Accomplishments that we're proud of
- Building a real constraint solver into a student-facing product, one that proves a multi-year schedule optimal or proves it infeasible across thousands of real variables and constraints, instead of shipping a heuristic and calling it intelligent.
- Building an AI advisor where hallucination is not mitigated, it is architecturally impossible for the specific class of error that matters most: the model is never given a path to invent a number, only to call a tool that computes one for real.
- Sourcing real, specific Carleton data rather than a plausible approximation of it, down to the level of individual co-op work-study sequences published by Carleton's own co-op office and cross-verified program maps for every seeded degree.
- Keeping the system honest under pressure: when the underlying AI provider has a bad moment, quota exhaustion or a genuine outage, the product degrades gracefully instead of failing or, worse, guessing.
What we learned
In an academic planning tool, correctness is not a feature, it is the entire product. A wrong CGPA or a missed prerequisite has a real consequence for a real student's graduation timeline, so every shortcut that traded correctness for convenience had to be treated as a defect, never as an acceptable rounding error.
Grounding a language model in deterministic tools changes more than accuracy, it changes the model's entire job. Once the model can no longer be the source of a fact, it becomes an interpreter and a narrator instead, and the system prompt has to be written with that boundary treated as load-bearing, not aspirational.
Real institutional data is messy in specific, unglamorous ways: a PDF that scrambles a table's visual reading order on extraction, two calendar labels that describe the same term in a different word order from each other. Handling that correctly took more careful engineering than either the AI layer or the constraint solver.
What we want to add next (and why)
Computer Science coverage, including all seven streams. Why: Engineering is fully modeled today, but Computer Science students have none of this yet. Unlike Engineering, Science does not publish an equivalent visual program map, so this means constructing a defensible course sequence from Carleton's own real course-numbering convention and the real prerequisite graph, clearly labeled as a constructed estimate rather than an official map, the same honest distinction the system already draws elsewhere between confirmed and estimated data.
Coverage for additional calendar years as Carleton publishes them. Why: A student's governing calendar year does not change, but the years available to new students entering the system does, and stale coverage quietly turns into wrong coverage.
A grade-prediction model trained on genuine Carleton outcomes. Why: The predictor already refuses to load a model trained on synthetic data and runs on a transparent heuristic instead. Growing the real completion dataset far enough to train responsibly replaces that heuristic with something more accurate, without ever compromising on where the data came from.
Extending the same architecture to other universities. Why: Nothing about the core idea, deterministic domain tools grounding an AI that is never allowed to calculate on its own, is specific to Carleton. Any university with a structured degree audit and a published calendar could be mapped the same way.
RavenMap is not trying to make an AI sound like it understands your degree. It is trying to make sure that when it tells you something about your degree, it actually does.
Built With
- agentic-ai
- ci/cd
- docker
- fastapi
- jwt
- kubernetes
- langchain
- machine-learning
- pgvector
- postgresql
- python
- react
- scikit-learn
- typescript
Log in or sign up for Devpost to join the conversation.