Inspiration

Mimesis started with a goal: make it cheap for robots to learn from human demonstrations. Robots pick up new skills from recordings of an operator guiding them through a task, but the tools for capturing those recordings are expensive and hard to use. Affordable research arms like the open-source SO-ARM101 are normally driven with a second "leader" arm, which doubles the hardware cost and takes practice to operate.

We wanted any student, hobbyist or researcher to be able to control a real robot arm with nothing more than a laptop, two webcams and their own hands.

What it does

Mimesis is an end-to-end, vision-based teleoperation system for the SO-ARM101. It works as follows:

  • Hand Tracking: Two consumer webcams and the WiLoR neural network find 21 keypoints on each hand. We triangulate them into a 3D hand pose.
  • Two-Hand Control: The right hand moves the gripper through 3D space. Twisting the left hand rotates the wrist, and making a fist closes the gripper.
  • Motion Control: A custom inverse-kinematics solver turns the hand pose into joint angles. A 50 Hz driver moves the arm with safety checks on each joint.
  • Research Ready: Sessions can be recorded and replayed exactly. A ROS 2 interface lets researchers plug in their own trackers, controllers or AI policies.

How we built it

Mimesis combines stereo computer vision, Kalman filtering, inverse kinematics and real-time servo control, connected through ROS 2.

The Hardware: An SO-ARM101 arm with six Feetech STS3215 servos, plus two Elgato Facecam MK.2 webcams at 1080p and 60 fps.

The Vision Pipeline: A YOLO hand detector and WiLoR (a vision-transformer hand model) run in PyTorch. An IMM Kalman filter then smooths the triangulated hand position.

Core Optimizations: Fixing how the driver handled messages cut command-to-arm latency from 700 ms to ~220 ms. A PyTorch upgrade and a multithreaded pipeline doubled tracking to 20–21 updates/s. The new filter cut smoothing lag from ~90 ms to 20–40 ms.

Metric Before After
Command-to-arm latency ~700 ms ~220 ms
Hand-tracking rate (one hand) ~10 updates/s 20–21 updates/s
Preview frame rate 30 fps 60 fps
Tracking filter lag 85–90 ms 20–40 ms
Jitter in the IK solver's output 35 % 6 %
Usable wrist rotation ~205° 333°
Base rotation reachable by hand ±30° full −64° to +80°

AMD Cloud Server: For laptops without a GPU, WiLoR runs as a FastAPI service on an AMD GPU with ROCm. We sped it up by batching both camera views into one FP16 model call, running the detector only every third frame, auto-tuning GPU kernels with TunableOp and keeping several requests in flight.

Laptop (2 webcams → JPEG) ──POST /detect over SSH tunnel──► AMD GPU server (FastAPI → YOLO + WiLoR)
          ▲                                                                   │
          └──────────────── JSON hand keypoints (both views) ◄────────────────┘
Optimization Result
AMD cloud deployment Runs without a local GPU; adds ~300 ms of network delay
WiLoR instead of MediaPipe More accurate hand detection
Tuned GPU kernels ~12 % less delay, ~7 % higher frame rate
BF16 → FP16 ~400 ms less delay; ~8× more accurate
Detector skipping ~20 ms less delay
Stereo batching ~80 ms less delay

Challenges we ran into

Network latency was our biggest challenge. The cloud server needs only ~20 ms per frame, but network transit adds about 300 ms per round trip, which we can't reduce from our side. We therefore support two modes: local GPU (~50–60 ms) for the most responsive control, and the AMD cloud server for machines without a GPU. On the hardware side, we had to correct a misfocused camera that caused wrist glitches, joint limits narrower than the robot model's, gravity sag in the shoulder, and a wrist encoder that couldn't rotate past a fixed point.

Accomplishments

We're proud of building a full-stack system that drives a physical robot arm through pick-and-place tasks using only two webcams and natural hand motion. We also cut its latency by more than half and ran hand tracking on NVIDIA, Apple and AMD GPUs.

What we learned

  • High-Performance Computing: GPU inference with PyTorch on ROCm, CUDA and Apple MPS, including kernel tuning and FP16 precision.
  • Robotics: Inverse kinematics, Kalman filtering and servo tuning on real hardware.
  • System Integration: Combining computer vision, a cloud server, ROS 2 and a physical robot into one application.

What's next

  • Record demonstrations in the LeRobot format and train AI policies that do tasks on their own.
  • Move the cloud server to a region with lower network delay.

Special Thanks to AMD

We thank AMD for access to the AMD Developer Cloud GPU that powers Mimesis's cloud mode.

Built With

Share this project:

Updates

Submission history