Inspiration

I had wanted to build a robotic arm project for a long time, and one day while I was using Codex I thought: what is actually stopping me from giving an AI tools to control a robot arm?

It can have tools to move the joints, and it can also have a tool to look through a camera. So I decided to build the physical arm and experiment with letting an LLM see, move, and interact with real objects.

I think AI controlling physical things through vision and tools is still really underexplored, and I wanted to make it easier for other people to experiment with this idea too.

For this challenge, I made a lighter web version of the project where an agent can control a simulated robot using WebMCP tools.

What RAI does

RAI lets users watch AI trials and experiment with models controlling a robot.

The agent can look through the robot camera, read information about the arm, move its joints, and try to complete tasks.

It does not get hidden object positions or shortcuts. It has to look at what happened, make a move, see the result, and adjust if it fails.

Users can also try different prompts and compare how the agent behaves.

How I built it

I mainly used Codex to build the project.

I already had a lot of the control logic and ideas from my physical robot arm and my experiments with Isaac Sim, so I adapted those ideas into a lightweight web simulation that could run in the browser and be controlled through WebMCP.

Codex helped me build most of the interface, simulation, and WebMCP integration.

Challenges

At first I wanted RAI to be much bigger. I wanted agents to control multiple robots and experiment in more complicated environments.

That became too much for the time I had, and I also had some personal things happening during the submission period, so I came pretty close to missing the deadline.

Because of that, I focused on making one smaller environment that clearly shows the main idea instead of trying to build everything.

What I learned

One of the most interesting things I noticed is that sometimes the agent is too careful.

It can hesitate to move the arm far enough or avoid moves that seem risky, even when those moves are needed to complete the task.

My favorite experiment is tipping over a can. With the right prompt, Sol can usually figure it out in around 2 minutes using only the camera and the arm controls.

It is really interesting to watch the model try something, see that it did not work, and then change its next move.

I think models could become much better at this if training and evaluation included more examples of vision + tool use for controlling physical systems.

What's next

I want to turn RAI into a better environment for testing how different AI models control robots.

One thing I am especially interested in is letting an AI look at its own failed experiments and then change its prompt or strategy to improve the next attempt.

Longer term, I want to do more research on this and look into collecting large amounts of robot interaction data in simulation.

Because the same experiments can be repeated many times in a simulator, I think this could be useful for evaluating or training future models that are better at interacting with the physical world.

Built With

Share this project:

Updates

Submission history