Plant the cheat, then catch it

Clever Hans Lab turns shortcut learning into an experiment students can run themselves. Generate shapes whose background colour predicts the label. Train a real convolutional neural network in your browser. It aces the familiar test; reverse the cue and see what survives. Then redesign the training data and try again.

The problem and audience

Test accuracy can conceal a model's reliance on a spurious cue. Lecturing about someone else's failure does not let a student reproduce it. This free, zero-install game is for students meeting ML, teachers running a 20-minute lesson, and anyone evaluating models by accuracy alone. A cited teacher lesson is included.

What it does

  • Real TensorFlow.js convolutional-network training, gradients, and live loss curves, with no canned result
  • Matched, reversed-cue, and neutral exams generated separately from the training set
  • A paired intervention preserving identical shapes and changing only the cue, starting with a fully cue-agreeing evaluation set
  • A redesign level for cue strength, neutral fraction, and sample budget
  • Evidence-derived reports for shortcut, partial reliance, shape-consistent, undertrained, and inconclusive outcomes
  • Labelled secondary diagnostics: neutral-fill ablation and patch-occlusion heatmaps
  • Keyboard controls, aria-live results, canvas text alternatives, reduced motion, and colour-independent cues

How it was built

Vite, strict TypeScript, and TensorFlow.js with WebGL and CPU fallback. The fully static app needs no accounts, backend, or API keys. The evaluation harness uses tfjs-node. Playwright covers rendered workflows and the captioned demo. GitHub Actions runs lint, typecheck, tests, the full measured harness, results verification, browser E2E, and the production build; GitHub Pages hosts the app.

What was measured

The committed harness covers 25 seeds × 8 training designs, plus 25 shuffled-label controls, with bootstrap intervals and explicit gates. At perfect cue correlation, reversed-cue accuracy falls to approximately 0% in all 25 tested seeds. Paired reliance changes from 1.000 at rho=1.0 to approximately 0.006 at rho=0.5 in this test system.

The paired swap reports aggregate evidence for a controlled intervention, not a complete causal explanation of every prediction. Exact train/test images are deduplicated, although shape-label combinations recur in the finite synthetic space. The null is a sanity check, not proof of chance performance; failed/advisory results and unrecovered training collapses remain disclosed. Results and methodology are in the Lab page and repository.

Challenges and lessons

Independent review caught dilution in a mixed-baseline cue swap, inconsistent quiz/verdict thresholds, stale navigation state, and overly strong null-control claims. The diagnostic now starts from a cue-agreeing set; one shared classifier powers exam, quiz, and verdict; lifecycle tests cover state consistency. A collapsed run is detected, retried transparently, and disclosed.

The lesson is to test the diagnostic itself. An attractive heatmap or a high score is not enough. The intervention, uncertainty, failure cases, and reproducibility limits belong alongside the result.

Honest limits

This is a synthetic educational system, not an audit of a production model. Intermediate cue strengths can be bistable across seeds. Reliance cutoffs are educational design choices, not universally validated thresholds. Reproducibility is within a backend; browser and headless backends can classify borderline designs differently. Pure-JavaScript CPU training was measured at about 164 seconds, or 639 seconds under simulated 4× slowdown. Real-phone and GPU timing were not measured. The demo uses a disclosed reduced training configuration: 480 images and 5 epochs.

Team and AI disclosure

Sharon Basovich, University of Waterloo, authorized this project and its submission campaign. AI-led assistance from dot and Devin (Cognition) produced ideation, implementation, tests, evaluation harness, documentation, and iterative independent review. This does not claim Sharon personally coded or reviewed the system. External work is cited in-app: Pfungst (1907), Lapuschkin et al. (2019), Geirhos et al. (2020), Ribeiro et al. (2016), and Zech et al. (2018).

Social impact

A student can watch a real model succeed for the wrong reason, intervene, and measure the change. Making that failure accessible and reproducible can help learners ask better questions about AI systems. No measured learning-outcome or real-world harm-reduction claim is made.

Live game: https://sharonbasovich.github.io/clever-hans-lab/ Source, measured results, and limitations: https://github.com/sharonbasovich/clever-hans-lab Corrected demo (training waits shortened and labelled): https://youtu.be/GRavNWVOEC4

Built With

Share this project:

Updates

Submission history