Inspiration

I know that I hanged on a windy tree

nine long nights, wounded with a spear,

dedicated to Odin, myself to myself…

I took up the runes, screaming I took them.

Hávamál, stanzas 138 and 139

Odin didn't read the runes. The poem is careful about that: he took them up. The Norse didn't call them letters, either. They called them staves: marks cut into wood and stone, things you could run a finger along.

Bragi, the god of poetry, came at language from the other end. The Poetic Edda (Sigrdrífumál, 16 to 18) lists his tongue among the things the runes were carved into. For him, speaking and writing were the same act.

We kept coming back to that pair because of a gap that's hard to unsee once you notice it. Almost every piece of assistive technology replaces exactly one sense. Screen readers assume you can hear. Captions assume you can see. People who are deafblind fall between the two, and neither sense can cover for the other.

That's not a small group. About 160 million people, roughly 2% of the world, live with some degree of deafblindness, and around 16 million live with it severely, according to the World Federation of the Deafblind. The Helen Keller National Center counts 2.4 million Americans with combined vision and hearing loss. Most of what gets built for their neighbours simply doesn't apply to them. What does apply tends to cost thousands of dollars and still needs an interpreter in the room.

So we asked a narrow question: what if one small camera and one real braille cell, worn on the hand, could run in both directions?


What it does

Rune & Bragi is a hand-worn camera and a six-dot braille cell, named for the two myths. Rune is the heart of it.

Rune: the world into the hand. Point the camera outward. It finds what matters in front of you, a bottle, a person, a sign, and taps it out letter by letter on six solenoids pressed against your fingertip. You feel points, the way braille is meant to be felt, not a buzz. If someone nearby says your name, their sentence goes to the front of the line. Everything else they say is thrown away before anyone sees it.

Bragi: the hand into the world. Turn the camera toward your other hand and fingerspell. It tracks 21 points on your hand, reads all 26 letters, J and Z included, and clicks softly as each one lands. When you sign Space, it works out the word you meant and says it out loud, so someone who doesn't know fingerspelling can understand you without anyone translating.

Same camera. Same hand. The only thing that changes is which way it's pointing.


Why one cell

Braille displays have rows of cells because readers sweep across whole lines. They're priced like it: the 20-cell Orbit Reader 20 is $799, and the 40-cell Mantis Q40 is $2,682. They do far more than we do, and we aren't trying to replace them.

When we asked about the idea on Reddit, cost came up before we mentioned it:

"Love the idea but cheapest non garbage braille display is still north of 500 USD."

One cell is a different thing. It's a tactile notification: a word, an alert, a sentence meant for you right now. At about a character a second, "bottle" arrives in six beats and a paragraph takes a while. We decided that was something to design around, not something to hide. Rune says less, and chooses carefully what.


How we built it

The cell

The hardware came first. We could have faked the output with a vibration motor, or a picture of a braille cell on a screen. The point was a real cell pressing real dots into a real finger, and everything else was built to feed it.

Six 12V push-pull solenoids sit in the braille layout, one per dot. A Raspberry Pi Zero 2 W drives them, but a Pi pin puts out 3.3V at a few milliamps, and a solenoid switching off throws back a spike that would kill it outright. So every solenoid gets its own ULN2803 Darlington driver, with all eight of the chip's channels ganged together: inputs to one GPIO, outputs to the solenoid, so no single channel carries the whole coil. The chip's COMMON pin goes to 12V, which gives its built-in flyback diodes somewhere to send the spike. One 12V source, an 8×AA pack, runs the coils directly, and a buck converter steps it down to 5V for the Pi. Every ground is tied together, because without that, "high" means nothing.

Rune circuit: a 12V pack feeds six solenoids and a buck converter; the Raspberry Pi Zero 2 W drives each solenoid through its own ULN2803 with all eight channels ganged

We verified the map by firing each output one at a time:

Braille dot Position GPIO (BCM) Header pin
1 top-left 17 11
2 middle-left 25 22
3 bottom-left 22 15
4 top-right 27 13
5 middle-right 23 16
6 bottom-right 24 18

On the Pi, a small Python server built on gpiozero takes a six-bit mask and raises exactly those dots. An OV5647 camera sits on the Pi's ribbon connector. The board boots straight into the solenoid controller and the camera stream, so power is all it needs.

This thing presses against skin, so the safety rules live in the Pi, not in the page. No solenoid is ever held on for more than two seconds, whatever anyone asks for. Every pin drops on exit, on error, and on Ctrl+C. A pin lock refuses every actuation until it's released, and because it lives on the device, reloading a browser can't get around it. Stop drops every pin and clears the queue at once. We wrote those limits before the first pulse ever fired.

All of it, Pi and camera included, came to about $50, and that counts every part we bought, including the ones that failed along the way. That's a prototype bill of materials, not a price tag. But it does show that a real tactile cell doesn't have to be priced like medical equipment.

Rune: from the world to the cell

The Pi streams 720p video at 30 fps to a laptop, because a Pi Zero has nowhere near the compute for what comes next. Three kinds of input feed the cell.

Objects go through YOLO26. The model chooses freely from everything it knows, and only then do we check the winner against 18 everyday classes worth feeling. When the camera moves to a new scene, anything that has left the frame is dropped from the queue before it wastes the wearer's time.

Text is gated before it's read. PP-OCRv4's small text detector runs locally on every frame and answers one question: is there readable text here? It catches print too small for the gate we started with. Only when it says yes does a crop go to a vision-language model to be transcribed, never more than once every five seconds, and never twice for the same text. The gate is cheap, so the expensive read only happens when it counts.

Speech goes through voice-activity detection and OpenAI's gpt-transcribe, then a name gate. A sentence survives only if it contains the wearer's name or one of their aliases. The rest is dropped before it's shown, stored, or queued. An opt-in nearby-voice mode also lets through someone standing close in front of the camera, for the moments when a stranger doesn't know your name.

All three land in one queue. Speech beats text, text beats objects, and speech is never pruned. When something more urgent arrives, whatever is tapping out gets cut short. Everything is converted to Grade 1 braille, and each message starts with its own touch pattern, the full cell ⠿ for speech and a square ⠶ for text, so the wearer knows what kind of message is coming before the first letter arrives.

The web app is the clock. Every cell it shows is also sent to the Pi, so the screen and the finger never disagree. When a letter repeats, like the "ll" in "hello", every pin drops for a beat in between, so you feel two letters instead of one long one.

Bragi: fingerspelling → speech

Every frame goes through Google's MediaPipe Hand Landmarker, which returns 21 landmarks per hand: wrist, knuckles, joints, fingertips. Our camera is mounted on its side, so for Bragi the landmarks are turned upright first, the way the model learned them.

Then we do something that sounds backwards. We turn the hand back into a picture. The 21 points are drawn as a 96×96 colour-coded skeleton, and a 1.2 MB CNN reads that drawing. A skeleton carries no background, no lighting and no skin tone, only the shape of the hand, which is the one thing that should matter. J and Z are traced through the air, so they come from short motion trails, not single frames. A letter only counts once the model holds it at 65% confidence or more for three reads in a row, and never the same letter twice running, so a hand passing through one shape on its way to another doesn't stutter.

The network only knew 26 letters, so we taught it Space ourselves: an open hand, all five fingers. We added a 27th output and trained only that, on our own recordings, with every letter weight frozen. On held-out recordings it caught 15 of 15 Spaces and mistook none of 720 letters for one.

When Space arrives, the letters go to a word decoder. The classifier is wrong in predictable ways. A doubled letter comes out once, a hand moving between shapes adds a stray letter, and from the Pi's angle some letters get swapped for each other: M for E, U for R, and the closed-fist letters for one another. So the decoder is a weighted edit distance over about 10,000 common English words plus the wearer's own names and terms, where exactly those mistakes are cheap.

Camera caught Decoded What went wrong
HELLWO hello stray letter, lost double
NAEM name two letters swapped
UEBC UMBC M read as E
BATHRUUM bathroom O read as U

On simulated noisy spellings it decodes 80% of single words, in under 0.2 seconds, on the laptop, with no network. If nothing is close enough, the raw letters are kept, so names and unusual words survive. The word is then spoken aloud.

We don't call this ASL translation. ASL is a language, with grammar, motion and a face. Bragi reads the fingerspelling alphabet and turns it into words, and that's all we claim.


Challenges we ran into

3.3 volts, 12-volt coils. A Pi can't drive a solenoid, and a solenoid can kill a Pi. Getting every dot to strike firmly without cooking a chip took six driver chips, forty-eight ganged channels, flyback diodes tied to the right rail, and one shared ground. When a dot doesn't fire, the answer is almost always the ground.

Solenoids cook themselves. Hold one on too long and it overheats, and this one is touching someone's finger. So nothing is ever held for more than two seconds, and every pin is forced off on every way the program can end.

The Pi couldn't find the network. The Pi Zero 2 W only speaks 2.4 GHz Wi-Fi, and the venue network keeps devices from seeing each other. A surprising share of the hardware time went into just being able to talk to the board. We ended with three ways in: a USB cable with a fixed address, a phone hotspot, and the Pi hosting its own network.

When the Pi dropped, everything froze. A browser only opens six connections to a server at once. With the Pi off the network, every request to it hung for over a minute, those hung requests used up all six, and the live hand tracking stalled behind them. The fix was making a missing Pi fail in two seconds, and sending one camera frame at a time instead of letting them pile up.

The camera sees sideways. It's mounted on its side, so every frame gets turned 90° before Rune reads it. That fixed Rune and quietly broke Bragi, whose models had only ever seen upright hands. Fingerspelling on the Pi camera fell to 4%. Turning the landmarks upright before classifying brought it to 67%, for any signer.

A gate built for sighted users snuck into our own design. The first version of the text reader asked the user to hold the camera steady on a preview until a crop looked clean, exactly the kind of visual feedback loop a deafblind user cannot close. We rebuilt it to tolerate normal hand motion instead of demanding a held, deliberate pose.

Restricting the vocabulary backfired. Limiting the object detector to only its allowed classes meant an unfamiliar object wasn't rejected. It was confidently misidentified as the nearest allowed class instead. The fix wasn't a hardcoded exception. It was letting the model choose freely across its full vocabulary first, and only checking the winner against the allowlist afterward.

The best model on paper was the worst one in practice. We started Bragi with five pretrained CNNs, ResNet50V2 and Xception, trained on a public ASL alphabet dataset of roughly 87,000 photos. Just getting them to load took a Python downgrade and a legacy-Keras shim. Once they ran, they were close to useless on a live camera. When we fed two of them pure random noise, they answered with near-100% confidence. They had learned the dataset's backgrounds and lighting, not hands. We cut all five.

Five fists. A, S, T, N and M are all closed hands. What separates them is where the thumb goes: beside the fingers, across them, between the first two, under two, under three. Telling those apart from a single camera is still the hardest thing Bragi does, which is why the word decoder treats those letters as near-misses for each other.

A smarter decoder wasn't smarter. We tried handing the choice to a language model, first an online one and then a local one. The local rescorer added less than a point of accuracy, and the online one put a network call in front of every word. We cut both, and decoding stays on the laptop.

Approach Letters Trained on On the Pi camera
Image CNNs on raw photos (5 models) 26 ~87k public photos Cut: confident on pure noise
KNN, laptop-webcam samples 24 1,440, one signer 4% on held-out frames
KNN, Pi-camera samples (personal) 24 1,440, one signer 96% held out, for that signer
Skeleton CNN, hands turned upright (default) 26 + Space public ASL datasets 4% → 67% offline, any signer

Accomplishments that we're proud of

  • A real six-dot braille cell: six solenoids and six driver chips, wired by hand, pressing actual dots into a fingertip. No vibration motor standing in for one.
  • The whole tactile build, Pi and camera included, for about $50, counting the parts that didn't survive.
  • Safety that lives in the hardware, written before the first pulse ever fired.
  • Objects, printed text and a spoken name, all arriving on that cell, live, in priority order.
  • Throwing away every sentence that isn't meant for the wearer, instead of transcribing the room.
  • Fingerspelled words, not just letters, spoken aloud, with J, Z and a Space sign we taught the model ourselves.
  • Catching our own sighted-user assumption before it shipped, not after.
  • Building for a group almost nothing is built for, and taking that seriously enough to say clearly what the device can't do yet.

What we learned

Hardware for the body is a different discipline. A bug in software shows the wrong letter. A bug in a solenoid driver can burn someone's hand. The safety limits had to exist before the first pulse, and they had to live on the device, where no browser could talk its way past them.

Trust the hand's shape, not its pixels. Five image models trained on 87,000 photos learned backgrounds. A small model that only ever sees the hand's skeleton works for anyone, and one trained on the wearer's own hand, through the camera they actually wear, reached 96%.

For most software, generalization is the goal. For assistive technology worn on one person's body, it may be the wrong goal entirely. A device that has learned how you sign, the angle of your wrist, how far your thumb tucks, is worth more than one that's roughly right about everybody.

We learned how easily a sighted assumption hides inside an interface. "Hold the camera steady until the preview looks clean" reads as a reasonable default, right up until you remember the person using it can't see the preview at all.


What's next for Rune & Bragi

  • More, smaller cells. A short row, so a whole word is felt at once, with tighter spacing closer to real braille and less click than a solenoid.
  • Off the laptop. An enclosure, a proper battery, and recognition moved onto portable hardware, so the whole device is a camera, a cell, a speaker and a strap.
  • Fewer words, safely. "Ritesh, your ride is waiting outside at the east entrance" could arrive as RIDE, EAST ENTRANCE, but only with the wearer in control of what gets cut.
  • Sentences, not just words, using what was already said to choose between close candidates.
  • Deafblind testers. Everything so far was built for deafblind people. The next version has to be built with them.

Rune brings the world into the hand.

Bragi gives the hand a voice.

Built With

Share this project:

Updates

Submission history