Inspiration

Paper tests are still the default in most classrooms, and for blind and low-vision students they usually mean one thing: a human scribe. A scribe costs the student independence and privacy, and it means someone else decides what gets written down. We wanted a student to be able to pick up a pen, take a paper test, and write their own answer in the right box, without anyone else involved. Think about how a student with a sighted classmate next to them never has to ask where the answer box is. We wanted to give that same confidence to a student who can't see the page.

Implementation

Axyz is a pen probe and an overhead camera that work together, with a voice that knows when to stay quiet.

The student holds the pen still and hears the question read aloud. As they move their hand, left and right buzzers on the pen pulse and short spoken cues ("Move 2 inches down") steer the pen to the answer box. When the pen settles inside, a double tone says it's locked in. While the student writes, the voice goes silent. Drifting past the margin triggers a harsh buzz. Three taps on the pen means "I'm done," and Gemini checks a photo of the box before we confirm "Answer recorded."

Our core rule is: the camera knows where the pen is, the accelerometer knows what the pen is doing, and the brain acts on the combination. A camera can't tell hovering from writing, can't see the pen under the student's hand, and can't receive "I'm done." The pen can do all three.

To build it we used:

  • Pen probe: an Arduino 101 with an accelerometer, a light sensor and two buzzers. The accelerometer classifies the pen as still, moving, writing or lifted, and detects one, two and three taps on the device.
  • Overhead camera: marker tracking plus a four-corner homography, $\mathbf{p}{\text{page}} = H\,\mathbf{p}{\text{image}}$, which gives the pen tip in page coordinates. It keeps the last position for a short time when the hand covers the pen.
  • Brain: a state machine (idle, reading, navigating, writing, auditing) that never speaks while the student is writing.
  • Voice: ElevenLabs, so the questions and cues sound natural.
  • Auditor: Gemini 2.5 Flash with structured output on the answer box. If the network is down or slow, it falls back to a pixel-based ink check.
  • Live view: a browser UI shows the camera, the voice waveform, the buzzers and the pen's state, so everyone watching can see what the student feels.

Future Steps

  • Detect the box edge directly with the pen's light sensor, for a faster margin buzz than the camera path allows.
  • Support multi-page tests.
  • Transcribe handwriting so answers can be reviewed digitally.
  • Add more languages.

Greatest Challenges

  • Separating writing from travelling. Accelerometer data is noisy, and both motions look alike. We tuned the thresholds and made the margin alarm fire only when the pen reports that it's writing.
  • Losing the pen under the hand. The camera's view of the pen drops out at exactly the moment the student starts writing. Rather than hide that, we made the pen's own state the source of truth during those moments.
  • Integrating four people's work. We agreed on one contract file and wrote a fake for every component, driven by a single scripted timeline. A headless simulation had to pass before any merge, so hardware, vision, voice and AI came together without blocking each other.
  • Everything has to fail gracefully. The network, the camera, the sensor and the voice can each fail on stage. Every component has a fallback or a keyboard override, so one flaky part can't end the demo.

What We Learned

Designing for someone who can't see the page changes what "feedback" means. Silence is a feature: the most important thing the system does is know when not to speak. We also learned that fakes and a shared contract matter as much as the hardware when four people have to integrate quickly.

Built With

Share this project:

Updates

Submission history