Inspiration

Training a robot to perform a task is only part of making it useful. An operator still needs to launch the right software, select the trained policy, and manage the execution process.

I wanted to make that interaction more natural: could someone request a trained robot skill in ordinary language, review what will happen, and approve it?

That idea became RoboStrands, a voice-guided robot assistant built with Strands Agents SDK.

What it does

RoboStrands connects spoken or typed requests to a trained physical robot skill, with human approval before movement.

The current prototype supports one task: pick up the orange cube, place it in the box, and return home.

The operator:

  1. Speaks a request through a ReSpeaker Lite microphone.
  2. Reviews and accepts the transcription.
  3. Reviews the task prepared by the agent.
  4. Types RUN ROBOT to authorize execution.
  5. Supervises the robot and confirms the physical result.

The intended users are laboratory technicians and small-workcell operators who need to run an already-trained robot routine. The intended benefit is less routine launch work through a guided interface that keeps the operator in control.

How I built it

The application is written in Python.

Strands Agents SDK connects a local Qwen model, served through Ollama, to a fixed task-preparation tool. The tool describes the supported action without moving the robot.

For voice interaction, ReSpeaker Lite captures audio, faster-whisper transcribes English locally, and eSpeak provides spoken replies.

A separate Python approval step controls execution. After approval, the application requests that Ollama unload Qwen, then launches a trained ACT policy through LeRobot. Camera observations feed the policy that controls the SO101 follower arm during the configured 35-second rollout.

The language model handles conversation and task preparation. The trained robot policy controls movement.

Challenges I faced

One challenge was sharing limited GPU memory between the language model and robot inference. I added an explicit model-unloading step before launching the robot policy.

Voice input also required iteration. Short recordings sometimes captured incomplete requests, so I tested a longer recording window and added transcript review before sending requests to the agent.

Another challenge was keeping conversation separate from execution. A model response alone cannot authorize movement: the preparation tool must be called, and the operator must provide keyboard approval.

Finally, a completed software process does not prove that a physical task succeeded. The application asks the operator to check the result instead of assuming success or retrying automatically.

Accomplishments

I connected voice input, local AI, Strands tool selection, human approval, and a physical robot rollout into one workflow.

The voice-initiated demonstration completed with operator-confirmed pick-and-place and return home. The project also includes 21 passing offline tests covering approval gates, dry-run behavior, transcript handling, and failure paths.

What I learned

Building an agent for physical hardware means coordinating more than an AI response. Resource management, clear interaction steps, bounded recording, process handling, and physical outcome checks all matter.

I also learned the value of separating responsibilities: the agent prepares a task, the operator authorizes it, and the trained policy executes it.

What’s next

I plan to make the robot configuration more portable, document and package the custom LeRobot runtime and checkpoint access, improve voice retry handling, and measure repeatability.

The current physical demonstration depends on my configured hardware and trained policy. It is a supervised prototype, not a safety-certified industrial system. Additional robot skills will require their own training and testing.

Built With

  • act
  • agents
  • faster-whisper
  • lerobot
  • ollama
  • python
  • qwen
  • respeaker
  • sdk
  • so101
  • strands
Share this project:

Submission history