What it is

AutoGrasp is a self-supervised robot-grasping system I built in CU Boulder's Correll Lab. A UR5e arm with a MAGPIE gripper detects an object, grasps it, and then grades its own attempt, so it collects clean training data with zero human labels.

How it works

A scripted expert detects the object (Gemini + SAM3), plans a grasp (PCA and NVIDIA GraspGenX in parallel), grasps with force-adaptive control (DeliGrasp), and a vision-language judge scores each attempt. Only episodes where the object measurably stayed held and the judge scored the grasp at 0.6 or higher enter the dataset. That runs at about 67 clean episodes per hour, fully unattended.

The honest failure

My first policy (60 random-placement episodes, an ACT model trained in 62 minutes on an NSF GH200) could replay its own data to 1.9 mm accuracy and then froze on rotated blocks in the real world. It turned out 71% of the training grasps were at one angle, and imitation learning averages conflicting demonstrations into inaction. The model was never the problem. The data was.

The fix and the result

I rebuilt collection as a 5x5 position grid crossed with 7 angles, with a continuous aperture action space, and audited every episode before training. The new policy, scored by a fully unattended graded evaluator I built, reached 97/100 picks (95% CI 91.5-99.0), including a perfect 75/75 on the rotated blocks that broke the first one.

Where it stops

I also tested the edge on purpose: one grid step (3 cm) outside the trained zone, success collapses into single digits. The policy interpolates cleanly within its data and does not extrapolate, which tells me exactly where the next episodes need to go.

Built with ROS2, SAM3, GraspGenX, DeliGrasp, ACT, LeRobot, Gemini, and NSF ACCESS (DeltaAI). SPUR final project, mentored by William Xie, PI Prof. Nikolaus Correll. Full case study: https://acwa-portfolio.netlify.app/#/projects/autograsp

Built With

Share this project:

Updates