Inspiration

People learn physical skills the same way they always have - a teacher grabs your hand and moves it through the motion, over and over, until your hand remembers it without your brain having to think. That doesn't scale: one doctor, one patient at a time; one piano teacher, one student. We wanted to know if that hand-over-hand teaching could be captured once and replayed forever - and whether an AI could be the one doing the guiding. Usually, humans write skills for Claude. We built a platform for Claude to write skills for humans.

What it does

Reflex is a robotic exoskeleton glove you can talk to, or that can learn from you.

Type into Claude - "make a fist," "point," "open your hand" - and it drives the glove's servos live over MCP.

Teleoperate the glove with your own hand, and it records the motion. Feed a handful of those recordings into a fine-tuned VLA (SmolVLA) and the glove starts reproducing the motion on its own.

A live web console shows the whole thing happening - 3D URDF viewport, depth camera feed, per-finger telemetry, commanded vs. measured joint state, preset gestures.

How we built it

Hardware: ST3215 bus servos, one per finger, 3D-printed RCM (remote center of motion) linkages, dowel-pinned joints. Intel RealSense D435i for depth/RGB/IMU, iPhone + Record3D for an egocentric overhead view. A HAL layer so the same control interface hits real servos or a simulated hand.

Stack:

  • ROS 2 Humble + ros2_control + Gazebo Fortress, bridged out via rosbridge/Foxglove
  • FastAPI backend: roslibpy over WebSockets, OpenCV, aiortc, MediaPipe
  • Next.js/React frontend: Three.js + urdf-loader for the live 3D hand, Zustand, shadcn
  • Claude connects straight into the same command path as the console, over MCP
  • Glove packaged as a LeRobot robot; demos teleoperated + pose-tracked, logged, used to fine-tune SmolVLA's action head; async inference over rosbridge with a threshold-controller fallback so the demo doesn't die if the model doesn't
    • Piano copilot built on Qwen OMNI, combining voice input, camera vision, and movement proposals into one multimodal loop

Challenges we ran into

We wanted the glove to be backdrivable - torque on, but soft, so pushing a finger moves it like a natural in hand-guide mode. The plan was right (weak-spring PID, position error as force), the tuning wasn't done in time. Still designed, not demoed.

Fusion 360 → URDF export kept producing joint geometry that looked fine in Fusion and broke on load, for reasons we never fully nailed down. Ended up diffing a hand-fixed version against one reconstructed from scratch.

Fifty demonstrations isn't a lot of data for a VLA, but that was the point; Reflex is designed for one-shot-style skill transfer, not general manipulation - but it meant every recording had to count. For 50 high-quality episodes, we went through nearly 100 attempts.

Accomplishments that we're proud of

Claude driving real hardware over MCP, live, synced to a 3D console. A HAL clean enough that sim and real share one code path. A usable teleop dataset off nothing but MediaPipe and a laptop camera.

What we learned

Admittance control is easy on paper and hard on a servo bus at 2am. A small, narrow dataset is a real research choice, not something to apologize for - as long as you're straight about what it does and doesn't prove.

What's next for Reflex

Finish the compliance tuning - that's what unlocks hand-over-hand motor relearning for Parkinson's/stroke rehab, and passive assist for repetitive work (learn a task from a few reps, start anticipating it). Bigger demonstration set. Recording sessions triggered straight from the console.

Built With

ros2 · ros2-control · gazebo · python · fastapi · roslibpy · opencv · mediapipe · nextjs · react · typescript · tailwindcss · threejs · react-three-fiber · urdf-loader · lerobot · smolvla · pytorch · anthropic-api · mcp · feetech-servo · platformio · esp32 · stm32 · kicad · docker · foxglove · rosbridge

Built With

Share this project:

Updates

Submission history