WhyRight — Where reasoning wins

Inspiration

A student selects the correct answer.

The teacher sees a green checkmark and moves on.

But the student’s explanation reveals something very different:

“The heavier object falls faster because gravity pulls harder on it.”

The student happened to choose the correct option, but the mental model behind that choice is still wrong.

Another student selects an incorrect answer, yet their explanation shows that they understand most of the concept and made only one small reasoning error. A traditional quiz scores the first student as correct and the second as wrong—even though the second student may be much closer to real understanding.

That gap inspired WhyRight.

Most classroom assessment tools are excellent at recording outcomes:

  • Which option was selected
  • Whether the answer was correct
  • How quickly the student responded
  • What percentage of the class got the question right

But they often cannot answer the questions teachers actually need answered:

  • Did the student understand the concept or guess correctly?
  • What mental model produced the wrong answer?
  • Are multiple students making the same mistake for the same reason?
  • Which misconception is held with high confidence?
  • What counterexample would challenge that misconception?
  • Did the student actually change their reasoning after receiving feedback?
  • What should the teacher ask next to separate the remaining misunderstandings?

Teachers rarely have enough time to manually inspect every written explanation, identify the misconception behind it, design a targeted intervention, and then verify whether the student’s thinking changed. This becomes especially difficult in a classroom with dozens of learners.

We wanted to build something that treats an incorrect answer not merely as a failure, but as evidence of a mental model.

That led to the central idea behind WhyRight:

Traditional quizzes measure which answer students choose. WhyRight reveals the reasoning that produced it—and whether that reasoning can change.

WhyRight was designed as a live classroom reasoning system where students are rewarded not only for being correct, but for explaining, questioning, revising, and improving their thinking.

The goal is not to replace teachers or allow AI to make unquestionable judgments about students. The goal is to give teachers a clearer, inspectable picture of classroom reasoning so they can make better instructional decisions.


What it does

WhyRight is a reasoning-first classroom assessment platform.

It helps teachers create diagnostic questions, collect student explanations, identify competing mental models, deliver targeted challenges, measure how reasoning changes, and determine the most informative question to ask next.

Instead of stopping at:

“Eleven students selected the wrong answer.”

WhyRight helps the teacher understand:

“Six students believe heavier objects fall faster, four believe continuous force is required to maintain motion, two selected the correct answer with unsupported reasoning, and three remain uncertain.”

It then helps the teacher act on that information.


The teacher experience

A teacher begins by creating a reasoning set around a learning objective.

For example:

“Diagnose Grade 8 misconceptions about falling objects and gravity.”

WhyRight generates a structured diagnostic package containing:

  • A question aligned with the learning objective
  • One correct answer
  • Three plausible misconception-based distractors
  • The specific flawed mental model represented by each distractor
  • An explanation of why each misconception is educationally plausible
  • A reasoning rubric
  • A targeted counterexample or Socratic challenge for each misconception
  • A follow-up question to verify whether understanding transfers to a new context
  • Ambiguity warnings for teacher review

The teacher does not have to accept the generated material blindly. They can:

  • Edit the question
  • Rewrite answer options
  • Change misconception labels
  • Modify the reasoning rubric
  • Review each planned challenge
  • Approve the package before launching it

Once approved, WhyRight creates a classroom session with a unique participant link and class code.


The student experience

Students join through a separate participant interface.

For each question, they complete three important steps:

1. Choose

The student selects an answer.

2. Explain

The student writes why they believe their answer is correct.

3. Calibrate

The student reports how confident they are.

This combination helps WhyRight distinguish between:

  • Demonstrated understanding
  • A correct but weakly supported answer
  • A lucky guess
  • Partial understanding
  • A stable misconception
  • Mixed or unclear reasoning
  • A confidently held incorrect mental model

Correctness matters, but it does not dominate the experience.

A student can earn meaningful credit for:

  • Providing relevant evidence
  • Explaining causal relationships
  • Identifying flaws in another argument
  • Recognising uncertainty
  • Revising a belief after encountering contradictory evidence
  • Applying a corrected principle to a new situation

Reasoning analysis

When a response is submitted, WhyRight analyses the student’s answer, explanation, and confidence together.

The system produces:

  • A suggested mental-model classification
  • Reasoning-quality assessment
  • Evidence-quality assessment
  • Confidence-calibration assessment
  • Exact phrases from the student response supporting the classification
  • An uncertainty estimate
  • An alternative possible interpretation
  • A recommendation for teacher review when the evidence is ambiguous

WhyRight never presents these classifications as unquestionable truths.

Every AI-generated assessment remains inspectable and teacher-overridable.

The teacher can see:

  • The learner’s original words
  • The relevant rubric criterion
  • Why the classification was suggested
  • The system’s confidence
  • Alternative interpretations
  • Whether the response should be manually reviewed

This keeps the teacher in control.


The Mental Model Map

The teacher dashboard does not reduce the class to a correct-versus-incorrect chart.

WhyRight groups responses according to the reasoning patterns behind them.

For a lesson on falling objects, the map might show:

  • Scientific model: gravity produces equal acceleration in controlled conditions
  • Heavier falls faster: greater mass is assumed to create greater acceleration
  • Force must continue: motion is believed to require a continuous push
  • Correct answer, weak reasoning: correct selection without stable conceptual support
  • Mixed or unclear: multiple conflicting ideas or insufficient explanation

Teachers can inspect individual learners within each group and review the exact evidence behind the classification.

This gives the teacher a live view of how the class is thinking, not only how the class scored.


Targeted counterexample challenges

WhyRight does not immediately reveal the correct answer after detecting a misconception.

Instead, it attempts to challenge the rule the student is currently using.

For a student who believes heavier objects always fall faster, WhyRight might present:

“A bowling ball and a smaller metal ball are released at the same time inside a vacuum chamber where air resistance has been removed. Which object reaches the bottom first, and why?”

The challenge may take several forms:

  • Counterexample
  • Edge case
  • Contradiction
  • Comparison
  • Socratic question
  • Evidence request
  • Mini simulation
  • Transfer scenario

The purpose is to create productive cognitive conflict.

Rather than saying:

“You are wrong. Here is the answer.”

WhyRight asks:

“Would your current rule still predict the correct result in this new situation?”

The student then receives an opportunity to:

  • Keep their original answer
  • Change their answer
  • Revise their explanation
  • Defend their reasoning
  • Express lower or higher confidence

Measuring reasoning change

After the student responds to the challenge, WhyRight compares the original and revised reasoning.

It evaluates:

  • Whether the answer changed
  • Whether the explanation changed
  • Whether the underlying mental model changed
  • Which misconception may have been repaired
  • Whether the student merely repeated expected language
  • Whether uncertainty remains
  • Whether the revised reasoning transfers to a new example

The goal is not simply to record that a student clicked a different option.

The goal is to determine whether the evidence suggests a meaningful change in reasoning.


Misconception Migration

One of WhyRight’s signature features is the Misconception Migration view.

Instead of showing only:

“61% of students answered correctly after feedback.”

WhyRight visualises movement between mental models.

For example:

Before the challenge After the challenge
Heavier falls faster: 6 students Heavier falls faster: 2 students
Force must continue: 4 students Force must continue: 1 student
Scientific model: 4 students Scientific model: 9 students
Correct answer, weak reasoning: 2 students Correct answer, weak reasoning: 3 students
Mixed or unclear: 2 students Mixed or unclear: 3 students

Teachers can follow individual reasoning trails and see:

  • Where a student started
  • Which challenge they received
  • How they revised their explanation
  • Which evidence supported the updated classification
  • Whether the system believes the misconception was repaired
  • Whether teacher review is still required

The migration view makes learning change visible.


Confidence calibration

WhyRight treats confidence as meaningful evidence.

Four students may produce very different instructional signals:

Answer Confidence Possible interpretation
Correct High Likely stable understanding
Correct Low Fragile knowledge or uncertainty
Incorrect Low Recognised uncertainty
Incorrect High Strong misconception requiring targeted intervention

Confidence is not used by itself. It is interpreted alongside the selected answer and written reasoning.

This helps teachers identify students who are confidently wrong, quietly uncertain, or correct for unstable reasons.


Reasoning-first scoring

WhyRight is designed so that correctness and speed do not dominate the score.

A possible ten-point structure is:

Dimension Maximum points
Initial answer 1
Reasoning quality 3
Relevant evidence 2
Critiquing an argument 1
Improvement after challenge 2
Confidence calibration 1
Total 10

This enables educationally meaningful outcomes:

  • A student who guessed correctly should not automatically outrank a student who reasoned thoughtfully and repaired a misconception.
  • A student who was initially wrong but revised their belief using evidence should receive meaningful credit.
  • A student who is correct and can defend the answer with strong evidence should receive the highest score.

WhyRight rewards productive reasoning, not only the first response.


Next Best Question

After the class completes a question, WhyRight identifies unresolved mental models and recommends what the teacher should ask next.

The system generates candidate diagnostic questions and estimates how students holding different misconceptions are likely to respond.

It then ranks those candidates using expected information gain.

This allows WhyRight to answer:

What is the most informative question this teacher can ask next?

For example, the system may recommend a question because it separates:

  • Students who believe heavier objects fall faster
  • Students who believe continuous force is required for motion
  • Students who selected correctly but still lack transferable understanding

The interface explains the recommendation in teacher-friendly language before displaying the mathematical score.

For example:

“This question is recommended because it separates the two largest remaining misconception groups and tests whether students can apply the scientific model in a new context.”

Then:

Expected information gain: 0.72 bits

This makes the recommendation understandable rather than presenting a mysterious AI-generated score.


Transparent hybrid classroom experience

A completely empty classroom does not immediately demonstrate the value of class-level reasoning analytics. At the same time, representing synthetic students as live participants would be misleading.

We therefore created a transparent hybrid evaluation experience.

WhyRight includes a clearly labelled sample classroom containing anonymised prepared responses. Judges can immediately explore:

  • Multiple student reasoning patterns
  • Mental-model classification
  • Targeted challenges
  • Before-and-after revision
  • Misconception migration
  • Next-question ranking

A judge can then join the same session as a genuine new participant.

Their submission is:

  • Marked Live participant
  • Timestamped
  • Processed through the real response-analysis pipeline
  • Added to the Mental Model Map
  • Reflected in the class counts
  • Sent through a targeted challenge
  • Included in Misconception Migration
  • Used to recalculate the next-question ranking

Prepared responses, live participant responses, live AI outputs, cached fallbacks, and provider sources are labelled separately.

The sample cohort provides immediate context. The live judge interaction proves that the product is real.


How we built it

WhyRight was built as a focused, locally runnable application rather than an overengineered production platform.

We intentionally avoided unnecessary infrastructure so that we could invest our time in:

  • The complete teacher-to-student workflow
  • Structured AI reasoning
  • Trust and explainability
  • Real application state
  • Demo reliability
  • Automated testing
  • Product-quality interaction design

Core technology

We used:

  • React 19 for the interactive teacher and participant experiences
  • TypeScript for type-safe application logic
  • Vite for frontend tooling and local API middleware
  • Zod for validating structured AI responses
  • Recharts for reasoning distributions and learning-change visualisations
  • Vitest for schema, unit, and integration tests
  • Playwright for browser-level workflow testing
  • Local JSON persistence for reasoning sets, sessions, participant responses, challenges, and revised explanations

The application is designed to run locally with minimal setup and does not require public deployment.


AI provider architecture

WhyRight supports two model providers:

  • OpenAI gpt-5.6-luna through the Responses API
  • NVIDIA NIM openai/gpt-oss-120b for repeated development, testing, and alternative-provider execution

Provider and model provenance are visible inside the product.

The application distinguishes between:

  • Live OpenAI output
  • Live NVIDIA output
  • Prepared sample content
  • Cached model output
  • Deterministic fallback behaviour

We never represent cached or prepared evidence as a live model result.


The structured AI pipeline

The AI system performs five main tasks.

1. Diagnostic question generation

Given a subject, learning level, topic, learning objective, teacher material, and desired difficulty, the model generates:

  • The diagnostic question
  • Correct answer
  • Misconception-based distractors
  • Misconception definitions
  • Reasoning rubric
  • Counterexample plans
  • Verification question
  • Ambiguity warnings

The result must satisfy a strict typed schema before it can enter the application.


2. Student reasoning analysis

The system analyses:

  • Selected answer
  • Written explanation
  • Confidence level
  • Teacher-approved misconception definitions
  • Teacher-approved rubric

It returns:

  • Suggested mental model
  • Supporting phrases
  • Reasoning score
  • Evidence score
  • Confidence calibration
  • Classification confidence
  • Alternative interpretation
  • Teacher-review recommendation

The system is instructed to avoid making claims about intelligence, permanent ability, or student personality.

It analyses only the current response and current reasoning pattern.


3. Targeted challenge generation

Based on the student’s current mental model, WhyRight creates or selects a challenge that tests the student’s rule without immediately revealing the expected answer.

The challenge is constrained by:

  • Learning objective
  • Teacher-approved concept
  • Misconception definition
  • Original student explanation
  • Appropriate challenge style
  • Age and difficulty level

4. Revision comparison

After the student responds again, the system compares the original and revised explanations.

It identifies:

  • New concepts introduced
  • Unsupported claims removed
  • Changes in causal reasoning
  • Changes in confidence
  • Whether the misconception remains
  • Whether the revised explanation merely copies challenge language
  • What uncertainty still requires teacher review

5. Candidate next-question generation

The model generates multiple follow-up diagnostic questions designed to distinguish unresolved mental models.

However, the final ranking is not delegated entirely to the model.


Deterministic information-gain ranking

One of the most important architectural decisions was separating probabilistic AI generation from deterministic mathematical ranking.

The model proposes:

  • Candidate follow-up questions
  • Possible answers
  • Which mental models each answer is expected to reveal
  • Predicted response distributions for each misconception group

TypeScript then independently calculates:

  • Current classroom entropy
  • Expected entropy after each candidate question
  • Expected information gain
  • Candidate ranking

Conceptually:

Information Gain =
Current uncertainty
−
Expected uncertainty after asking the question

This means the model can generate educational possibilities, but transparent code determines which candidate is mathematically most diagnostic.

The teacher can inspect:

  • Which misconceptions a question targets
  • Why the question was recommended
  • The expected distribution of responses
  • The resulting information-gain calculation

This separation makes the recommendation system more explainable and testable.


Structured-output reliability

Model-generated JSON can be incomplete, malformed, or inconsistent.

To make the pipeline reliable enough for a live workflow, we implemented:

  • Strict Zod schemas
  • Required enums and bounded values
  • Model-response validation
  • Repair attempts
  • Bounded retries
  • Request timeouts
  • Provider fallback handling
  • Visible failure states
  • Clearly labelled cached results
  • Deterministic fallback fixtures for demonstration continuity
  • Teacher review for low-confidence classifications

The interface never silently converts malformed output into an authoritative educational judgment.


Real application state

Every major screen is connected to actual state.

The application supports:

  • Creating reasoning sets
  • Editing generated questions
  • Approving questions
  • Launching sessions
  • Creating invitation links
  • Joining through the participant interface
  • Submitting answers
  • Saving explanations and confidence
  • Analysing responses
  • Delivering challenges
  • Recording revised explanations
  • Updating class distributions
  • Updating reasoning trails
  • Recalculating next-question rankings
  • Teacher overrides
  • Persistence across page refreshes

The prepared sample class uses the same underlying response structures as new participant submissions.


User-interface design

We explored three visual directions before selecting a spacious light-theme design focused on the reasoning journey.

The final design emphasises:

  • Clear teacher and student role separation
  • Generous whitespace
  • High readability
  • Large, accessible type
  • Consistent semantic colours
  • Progressive disclosure
  • Evidence-first AI explanations
  • Clear loading, success, uncertainty, and fallback states

The central visual sequence is:

Initial reasoning
      ↓
Targeted challenge
      ↓
Revised reasoning

Our primary colours have consistent meanings:

  • Green: scientifically supported or repaired reasoning
  • Red: unresolved misconception
  • Amber: uncertainty or teacher review
  • Grey: mixed or unclear reasoning
  • Indigo: primary action and diagnostic guidance
  • Deep navy: targeted challenge

The most important visual moment is the transition from the class’s initial mental models to its revised mental models.


How Codex helped

Codex was used as an active product and engineering collaborator throughout development.

It helped us:

  • Critique the original concept
  • Identify overlap with conventional quiz products
  • Strengthen the product around mental-model change
  • Compare architecture options
  • Select a focused hackathon architecture
  • Define structured AI schemas
  • Build teacher and participant workflows
  • Implement deterministic scoring and information-gain logic
  • Explore three visual design directions
  • Refine the selected user-interface system
  • Connect screens to real application state
  • Build local persistence
  • Generate evaluation fixtures
  • Create and run tests
  • Diagnose implementation issues
  • Validate the final production build
  • Document trade-offs and limitations

We did not ask Codex merely to generate isolated components. We used it to move from product critique through architecture, implementation, testing, and refinement.

Human decisions remained central to:

  • Product scope
  • Educational principles
  • Trust and teacher-control requirements
  • What not to build
  • Which visual direction to choose
  • How sample and live evidence should be represented
  • Which parts of the system should remain deterministic
  • Which claims the product should and should not make

Challenges we ran into

Moving beyond answer correctness

The first major challenge was defining what “understanding” means in a way that could be represented responsibly.

A student response cannot safely be reduced to one binary label.

We needed to distinguish:

  • Correct answer with strong reasoning
  • Correct answer with weak reasoning
  • Correct answer produced through an incorrect mental model
  • Incorrect answer with useful partial understanding
  • Stable misconception
  • Mixed reasoning
  • Unclear explanation
  • High-confidence misunderstanding
  • Low-confidence uncertainty

This required the system to analyse answer choice, explanation, evidence, and confidence together.

It also required careful language.

WhyRight never labels a student as inherently weak or incapable. It uses temporary, evidence-based terms such as:

  • “Current reasoning pattern”
  • “Possible misconception”
  • “Evidence suggests”
  • “Mixed or unclear”
  • “Needs teacher review”

Making AI judgments inspectable

A raw model label is not enough for a classroom product.

If the system suggests that a learner holds a misconception, the teacher needs to know why.

We therefore designed every classification around inspectable evidence:

  • Original student phrase
  • Matched misconception
  • Rubric evidence
  • Confidence level
  • Alternative interpretation
  • Uncertainty
  • Teacher override

This added complexity, but it made the product more trustworthy.


Generating meaningful distractors

Generating random incorrect answers is easy.

Generating distractors that correspond to plausible and instructionally useful misconceptions is much harder.

A poor distractor may be:

  • Obviously wrong
  • Ambiguous
  • Unrelated to the learning objective
  • Based on wording tricks
  • Impossible to interpret diagnostically

We addressed this by requiring every distractor to include:

  • Explicit misconception label
  • Explanation of why a student might believe it
  • Relevant counterexample
  • Follow-up verification plan
  • Teacher approval

The teacher can modify or reject any generated option.


Challenging without immediately revealing

Most automated feedback tells the student what the correct answer is.

WhyRight needed to do something more difficult: generate an intervention that exposes a weakness in the student’s rule without simply supplying the solution.

The challenge had to be:

  • Relevant to the misconception
  • Understandable at the learner’s level
  • Short enough for a classroom
  • Different from the original question
  • Capable of producing cognitive conflict
  • Safe from introducing a new misconception

We implemented multiple challenge styles and kept teacher review available.


Reliable structured AI output

The application depends on model-generated structured data.

In development, responses could contain:

  • Missing fields
  • Invalid enum values
  • Extra prose
  • Incorrect nesting
  • Inconsistent scoring
  • Unsupported confidence values
  • Incomplete counterexamples

We addressed this through:

  • Strict schemas
  • Validation
  • Repair prompts
  • Retry limits
  • Provider-aware error handling
  • Cached fallback packages
  • Clear provenance labels

The interface is designed to fail visibly and recover gracefully rather than silently displaying invalid information.


Separating AI generation from mathematical ranking

The next-question feature initially risked becoming another arbitrary AI recommendation.

A model could say:

“Information gain: 0.84”

without any transparent basis.

We therefore separated the system:

  • AI generates candidate diagnostic questions and likely misconception responses.
  • Deterministic TypeScript code calculates entropy and expected information gain.
  • The UI shows the teacher both the human-readable rationale and the calculated value.

This made the feature more credible and technically meaningful.


Demonstrating a classroom honestly

A new session with zero responses cannot showcase mental-model clustering or migration.

However, presenting synthetic responses as a live classroom would be misleading.

We solved this with a clearly labelled prepared cohort.

The interface distinguishes:

  • Sample response
  • Live participant response
  • Live AI analysis
  • Prepared AI analysis
  • Cached fallback
  • Teacher override

A judge can add a genuine response through the participant experience and watch the application update.

The sample data establishes context. The live interaction demonstrates authenticity.


Keeping the scope disciplined

The product could easily have expanded into:

  • A full learning-management system
  • Student accounts
  • School administration
  • Parent reporting
  • Curriculum marketplaces
  • Real-time infrastructure
  • Complex role management
  • Mobile applications
  • Institutional analytics

We deliberately excluded these.

For the hackathon, we prioritised:

  1. Working vertical slice
  2. Coherent user experience
  3. Real state transitions
  4. AI transparency
  5. Demo reliability
  6. Visible technical depth

That discipline allowed us to build a connected product rather than a collection of incomplete features.


Accomplishments that we’re proud of

We are proud that WhyRight is not a static prototype or a set of disconnected demonstration screens.

It supports a complete reasoning workflow:

Teacher creates a diagnostic question
→ Teacher reviews misconception-based options
→ Teacher launches a reasoning set
→ Student joins through a participant link
→ Student submits an answer, explanation, and confidence
→ WhyRight analyses the reasoning
→ Teacher inspects evidence and uncertainty
→ Student receives a targeted challenge
→ Student revises or defends the explanation
→ WhyRight compares the reasoning
→ The class mental-model map updates
→ Misconception Migration becomes visible
→ The next diagnostic question is recalculated

The completed product includes:

  • A working teacher experience
  • A separate participant experience
  • Multi-question reasoning sets
  • Unique participant invitation links
  • Classroom join codes
  • Teacher-generated and AI-generated question packages
  • Misconception-based distractors
  • Teacher editing and approval
  • Student answer selection
  • Student confidence ratings
  • Written reasoning collection
  • Evidence-backed mental-model classification
  • Alternative interpretation and uncertainty
  • Teacher review flags
  • Teacher classification override
  • Targeted counterexample challenges
  • Revised explanation capture
  • Before-and-after reasoning comparison
  • Class-level mental-model grouping
  • Individual reasoning trails
  • Misconception Migration visualisation
  • Deterministic entropy calculation
  • Deterministic expected-information-gain ranking
  • Ranked next-question recommendations
  • Persistent local reasoning-set storage
  • Persistent participant-response storage
  • Transparent provider provenance
  • Transparent sample/live/fallback labelling
  • A prepared three-question classroom
  • Eight consistent anonymised learners
  • Twenty-four validated question-level reasoning records
  • Thirty-two passing automated tests
  • A successful production TypeScript build

Most importantly, major screens are connected to genuine state and user actions.

Counts change when responses are added.

Classifications affect the Mental Model Map.

Revisions affect Misconception Migration.

Unresolved reasoning affects the next-question recommendation.

Teacher overrides affect the stored classification.

The interface is not merely illustrating what the product might do. It is executing the workflow.


What we learned

Educational AI is more useful when it supports decisions

Generating a question is helpful.

Explaining an answer is helpful.

But teachers receive greater value when AI helps them decide:

  • Which misconception is most common?
  • Which students need review?
  • Which intervention should be used?
  • Which mental model remains unresolved?
  • What question should be asked next?

WhyRight became stronger when we moved from content generation to decision support.


Correctness is not the same as understanding

The project reinforced a simple but important lesson:

A correct answer is evidence, not proof, of understanding.

The explanation behind the answer often reveals more than the selected option.

A student who revises their belief after evidence may demonstrate more meaningful learning than a student who guessed correctly and cannot defend the response.

This shaped the scoring, analytics, and challenge design.


Misconceptions should be treated as models, not defects

Students are often not responding randomly.

They are applying a rule that seems internally reasonable based on their current understanding.

For example:

“A heavier object experiences more gravitational force, so it must fall faster.”

This is not meaningless. It combines true information with an incorrect conclusion.

Treating the response as a mental model makes it possible to design a challenge that targets the exact reasoning error.


AI assessments must remain contestable

A classroom AI system should not silently label students.

Teachers need:

  • Evidence
  • Confidence
  • Alternatives
  • Overrides
  • Review states
  • Auditability

The more consequential the interpretation, the more important transparency becomes.

We learned to present model outputs as suggestions supported by evidence—not final declarations.


Probabilistic and deterministic components should be separated

Language models are useful for:

  • Generating educational possibilities
  • Interpreting natural-language explanations
  • Producing counterexamples
  • Suggesting follow-up questions

Deterministic code is better for:

  • Scoring rules
  • Counts
  • State transitions
  • Entropy
  • Expected information gain
  • Candidate ranking

Combining both made the system more understandable and testable.


Sample data is acceptable when provenance is honest

Prepared data can make a complex workflow easier to evaluate.

The important questions are:

  • Is it labelled?
  • Is its source visible?
  • Can real input be added?
  • Does live input use the actual pipeline?
  • Does the system avoid claiming that prepared responses are live?

We learned that transparency is more valuable than pretending every visible classroom event occurred during the demonstration.


A polished demo requires real product behaviour

Visual design alone cannot make a product credible.

The application began to feel real only when:

  • Buttons changed state
  • Teacher edits persisted
  • Student submissions appeared in the teacher view
  • Charts derived from stored responses
  • Revisions changed classifications
  • Rankings recalculated
  • Empty states existed
  • Failure states were handled
  • Refreshes preserved work

The distinction between a mockup and a product is not the number of screens. It is whether user actions have consequences.


Scope is a product decision

We learned that avoiding overengineering is not the same as avoiding technical depth.

We did not need microservices, Kubernetes, multiple databases, or production authentication to demonstrate meaningful engineering.

The technical depth came from:

  • Typed domain models
  • Structured AI orchestration
  • Validation and recovery
  • Explainable classification
  • Deterministic information gain
  • State-connected visualisation
  • Automated testing
  • Honest fallback behaviour

A focused system can be technically serious without being unnecessarily large.


What’s next for WhyRight

The next step is not immediately adding dozens of new features.

The next step is validating whether WhyRight improves real teaching decisions.

We plan to begin with a small teacher pilot measuring:

  • Agreement between teacher classifications and WhyRight suggestions
  • Time saved when reviewing written reasoning
  • Accuracy of misconception identification
  • Frequency of teacher overrides
  • Quality of generated counterexamples
  • Whether students revise answers without merely copying expected language
  • Whether repaired reasoning transfers to new contexts
  • Whether next-question recommendations reduce classroom uncertainty
  • Teacher confidence in the evidence and explanations shown
  • Student perception of fairness and usefulness

Near-term product improvements

Real-time classroom synchronisation

The current application is locally persistent. A future version would support real-time multi-device classroom updates using production-grade storage and synchronisation.

Secure teacher accounts

Teachers would receive authenticated workspaces, private classroom sessions, and controlled invitation links.

Curriculum-aligned misconception libraries

Teachers could use reviewed misconception packages aligned with:

  • Grade level
  • Subject
  • Curriculum
  • Learning standard
  • Language
  • Difficulty

AI-generated content would complement, not replace, teacher-reviewed educational resources.

Teacher-created rubrics

Teachers would be able to define:

  • Required concepts
  • Acceptable evidence
  • Common misconceptions
  • Scoring priorities
  • Subject-specific terminology

Longitudinal reasoning profiles

WhyRight could track how a learner’s explanations evolve across lessons.

This would focus on:

  • Reasoning patterns
  • Evidence use
  • Confidence calibration
  • Transfer
  • Revision behaviour

It would not create permanent intelligence or ability labels.

Multilingual support

Students should be able to explain reasoning in the language in which they think most clearly.

Future work would include:

  • Multilingual questions
  • Multilingual explanations
  • Cross-language teacher summaries
  • Language-aware rubric evaluation

Accessibility

Planned accessibility work includes:

  • Keyboard navigation
  • Screen-reader improvements
  • Colour-independent classification indicators
  • Voice response support
  • Adjustable text size
  • Reduced-motion mode
  • High-contrast mode

Learning-management-system integrations

WhyRight could integrate with existing classroom environments rather than requiring teachers to replace them.

Potential integrations include:

  • Roster imports
  • Assignment links
  • Grade export
  • Class-level reasoning reports
  • Single sign-on

Privacy and school data controls

A production version would require:

  • Data-retention controls
  • Student anonymity settings
  • Consent workflows
  • Regional compliance
  • School-managed deletion
  • Encryption
  • Access logs
  • Clear limitations on model-provider data usage

Bias and error evaluation

We plan to build teacher-reviewed datasets covering:

  • Different writing styles
  • Short and long explanations
  • Multilingual reasoning
  • Learners with varying literacy levels
  • Ambiguous responses
  • Correct answers with flawed reasoning
  • Incorrect answers with strong partial understanding

This will help measure:

  • Classification agreement
  • False misconception detection
  • Overconfidence
  • Subject-specific failure modes
  • Differences across language and writing style

Long-term vision

Our long-term vision is for WhyRight to become a reasoning layer for classroom assessment.

It could work alongside quizzes, assignments, discussions, simulations, and classroom-response systems.

The purpose would not be to automate teaching.

The purpose would be to help teachers see what ordinary scores hide:

  • The ideas students are using
  • The misconceptions they share
  • The evidence they trust
  • The confidence with which they hold a belief
  • The interventions that change their thinking
  • The next question most likely to deepen understanding

Today, most classroom software answers:

“Who was right?”

WhyRight is built to answer a more useful sequence of questions:

Why did they believe it? What evidence could challenge it? Did their reasoning change? What should we ask next?

That is the future we want to build.

WhyRight — Where reasoning wins.

Built With

  • codex
  • gpt-5.6
Share this project:

Updates