Inspiration
People learn physical skills the same way they always have - a teacher grabs your hand and moves it through the motion, over and over, until your hand remembers it without your brain having to think. That doesn't scale: one doctor, one patient at a time; one piano teacher, one student. We wanted to know if that hand-over-hand teaching could be captured once and replayed forever - and whether an AI could be the one doing the guiding. Usually, humans write skills for Claude. We built a platform for Claude to write skills for humans.
What it does
Reflex is a robotic exoskeleton glove you can talk to, or that can learn from you.
Type into Claude - "make a fist," "point," "open your hand" - and it drives the glove's servos live over MCP.
Teleoperate the glove with your own hand, and it records the motion. Feed a handful of those recordings into a fine-tuned VLA (SmolVLA) and the glove starts reproducing the motion on its own.
A live web console shows the whole thing happening - 3D URDF viewport, depth camera feed, per-finger telemetry, commanded vs. measured joint state, preset gestures.
How we built it
Hardware: ST3215 bus servos, one per finger, 3D-printed RCM (remote center of motion) linkages, dowel-pinned joints. Intel RealSense D435i for depth/RGB/IMU, iPhone + Record3D for an egocentric overhead view. A HAL layer so the same control interface hits real servos or a simulated hand.
Stack:
- ROS 2 Humble +
ros2_control+ Gazebo Fortress, bridged out via rosbridge/Foxglove - FastAPI backend:
roslibpyover WebSockets, OpenCV,aiortc, MediaPipe - Next.js/React frontend: Three.js +
urdf-loaderfor the live 3D hand, Zustand, shadcn - Claude connects straight into the same command path as the console, over MCP
- Glove packaged as a LeRobot robot; demos teleoperated + pose-tracked, logged, used to fine-tune SmolVLA's action head; async inference over rosbridge with a threshold-controller fallback so the demo doesn't die if the model doesn't
- Piano copilot built on Qwen OMNI, combining voice input, camera vision, and movement proposals into one multimodal loop
Challenges we ran into
We wanted the glove to be backdrivable - torque on, but soft, so pushing a finger moves it like a natural in hand-guide mode. The plan was right (weak-spring PID, position error as force), the tuning wasn't done in time. Still designed, not demoed.
Fusion 360 → URDF export kept producing joint geometry that looked fine in Fusion and broke on load, for reasons we never fully nailed down. Ended up diffing a hand-fixed version against one reconstructed from scratch.
Fifty demonstrations isn't a lot of data for a VLA, but that was the point; Reflex is designed for one-shot-style skill transfer, not general manipulation - but it meant every recording had to count. For 50 high-quality episodes, we went through nearly 100 attempts.
Accomplishments that we're proud of
Claude driving real hardware over MCP, live, synced to a 3D console. A HAL clean enough that sim and real share one code path. A usable teleop dataset off nothing but MediaPipe and a laptop camera.
What we learned
Admittance control is easy on paper and hard on a servo bus at 2am. A small, narrow dataset is a real research choice, not something to apologize for - as long as you're straight about what it does and doesn't prove.
What's next for Reflex
Finish the compliance tuning - that's what unlocks hand-over-hand motor relearning for Parkinson's/stroke rehab, and passive assist for repetitive work (learn a task from a few reps, start anticipating it). Bigger demonstration set. Recording sessions triggered straight from the console.
Built With
ros2 · ros2-control · gazebo · python · fastapi · roslibpy · opencv · mediapipe · nextjs · react · typescript · tailwindcss · threejs · react-three-fiber · urdf-loader · lerobot · smolvla · pytorch · anthropic-api · mcp · feetech-servo · platformio · esp32 · stm32 · kicad · docker · foxglove · rosbridge





Log in or sign up for Devpost to join the conversation.