Inspiration
Mimesis started with a goal: make it cheap for robots to learn from human demonstrations. Robots pick up new skills from recordings of an operator guiding them through a task, but the tools for capturing those recordings are expensive and hard to use. Affordable research arms like the open-source SO-ARM101 are normally driven with a second "leader" arm, which doubles the hardware cost and takes practice to operate.
We wanted any student, hobbyist or researcher to be able to control a real robot arm with nothing more than a laptop, two webcams and their own hands.
What it does
Mimesis is an end-to-end, vision-based teleoperation system for the SO-ARM101. It works as follows:
- Hand Tracking: Two consumer webcams and the WiLoR neural network find 21 keypoints on each hand. We triangulate them into a 3D hand pose.
- Two-Hand Control: The right hand moves the gripper through 3D space. Twisting the left hand rotates the wrist, and making a fist closes the gripper.
- Motion Control: A custom inverse-kinematics solver turns the hand pose into joint angles. A 50 Hz driver moves the arm with safety checks on each joint.
- Research Ready: Sessions can be recorded and replayed exactly. A ROS 2 interface lets researchers plug in their own trackers, controllers or AI policies.
How we built it
Mimesis combines stereo computer vision, Kalman filtering, inverse kinematics and real-time servo control, connected through ROS 2.
The Hardware: An SO-ARM101 arm with six Feetech STS3215 servos, plus two Elgato Facecam MK.2 webcams at 1080p and 60 fps.
The Vision Pipeline: A YOLO hand detector and WiLoR (a vision-transformer hand model) run in PyTorch. An IMM Kalman filter then smooths the triangulated hand position.
Core Optimizations: Fixing how the driver handled messages cut command-to-arm latency from 700 ms to ~220 ms. A PyTorch upgrade and a multithreaded pipeline doubled tracking to 20–21 updates/s. The new filter cut smoothing lag from ~90 ms to 20–40 ms.
| Metric | Before | After |
|---|---|---|
| Command-to-arm latency | ~700 ms | ~220 ms |
| Hand-tracking rate (one hand) | ~10 updates/s | 20–21 updates/s |
| Preview frame rate | 30 fps | 60 fps |
| Tracking filter lag | 85–90 ms | 20–40 ms |
| Jitter in the IK solver's output | 35 % | 6 % |
| Usable wrist rotation | ~205° | 333° |
| Base rotation reachable by hand | ±30° | full −64° to +80° |
AMD Cloud Server: For laptops without a GPU, WiLoR runs as a FastAPI service on an AMD GPU with ROCm. We sped it up by batching both camera views into one FP16 model call, running the detector only every third frame, auto-tuning GPU kernels with TunableOp and keeping several requests in flight.
Laptop (2 webcams → JPEG) ──POST /detect over SSH tunnel──► AMD GPU server (FastAPI → YOLO + WiLoR)
▲ │
└──────────────── JSON hand keypoints (both views) ◄────────────────┘
| Optimization | Result |
|---|---|
| AMD cloud deployment | Runs without a local GPU; adds ~300 ms of network delay |
| WiLoR instead of MediaPipe | More accurate hand detection |
| Tuned GPU kernels | ~12 % less delay, ~7 % higher frame rate |
| BF16 → FP16 | ~400 ms less delay; ~8× more accurate |
| Detector skipping | ~20 ms less delay |
| Stereo batching | ~80 ms less delay |
Challenges we ran into
Network latency was our biggest challenge. The cloud server needs only ~20 ms per frame, but network transit adds about 300 ms per round trip, which we can't reduce from our side. We therefore support two modes: local GPU (~50–60 ms) for the most responsive control, and the AMD cloud server for machines without a GPU. On the hardware side, we had to correct a misfocused camera that caused wrist glitches, joint limits narrower than the robot model's, gravity sag in the shoulder, and a wrist encoder that couldn't rotate past a fixed point.
Accomplishments
We're proud of building a full-stack system that drives a physical robot arm through pick-and-place tasks using only two webcams and natural hand motion. We also cut its latency by more than half and ran hand tracking on NVIDIA, Apple and AMD GPUs.
What we learned
- High-Performance Computing: GPU inference with PyTorch on ROCm, CUDA and Apple MPS, including kernel tuning and FP16 precision.
- Robotics: Inverse kinematics, Kalman filtering and servo tuning on real hardware.
- System Integration: Combining computer vision, a cloud server, ROS 2 and a physical robot into one application.
What's next
- Record demonstrations in the LeRobot format and train AI policies that do tasks on their own.
- Move the cloud server to a region with lower network delay.
Special Thanks to AMD
We thank AMD for access to the AMD Developer Cloud GPU that powers Mimesis's cloud mode.
Log in or sign up for Devpost to join the conversation.