Inspiration
Most AI coding tools can talk about your software.
The harder problem is letting an AI actually work on a real machine without turning the entire computer into one giant uncontrolled tool call.
ROSIE started as the local execution layer for HACKASS.
HACKASS needed more than a model that could generate code. It needed a way for an agent to inspect a workspace, understand what is already there, modify files, run commands, test its work, recover from failures, and keep going.
That is ROSIE.
What it does
ROSIE gives HACKASS a controlled way to work on a real local software project.
HACKASS is the user-facing agent.
ROSIE is the machinery underneath it that can actually interact with the development environment.
Through ROSIE, the agent can work with the project workspace, read and write files, execute shell commands, run tests, inspect results, and continue an engineering task instead of stopping after generating a block of code.
The goal is not "AI writes code."
The goal is:
AI can participate in the software-development lifecycle.
ROSIE provides the bridge between model reasoning and actual machine execution.
How we built it
ROSIE is built around a runtime-owned execution model.
Instead of scattering machine access, state, approvals, and tools across the application as globals, the runtime owns the active session and the resources used during that session.
That includes the workspace, execution tools, approval boundaries, task state, and agent interaction.
The architecture deliberately separates responsibilities.
The model decides what work needs to happen.
ROSIE provides the bounded tools for doing that work.
HACKASS provides the interaction layer that lets the user direct the agent.
This separation also means the local execution machinery is not permanently tied to one interface. The same runtime can be driven by different clients while preserving the underlying execution model.
A major design principle throughout the project has been:
Design software so intelligence can reason locally and the system can verify globally.
Challenges we ran into
The difficult part was not getting an LLM to generate code.
That part is easy.
The difficult part was everything that comes after it.
Where is the project?
What files already exist?
What is the agent allowed to change?
How does it run the software?
How does it know whether a command succeeded?
How does it recover when something fails?
Who owns the state of an active build?
What happens when multiple parts of the system need access to the same runtime?
Those questions pushed ROSIE away from being a collection of convenient tools and toward becoming an actual runtime with explicit ownership and boundaries.
Another challenge was avoiding architecture that only works for the current demo.
We deliberately moved execution responsibilities into session-owned components so ROSIE could evolve without every interface, tool, and subsystem becoming permanently coupled together.
Accomplishments that we're proud of
The biggest accomplishment is that HACKASS can move beyond conversation and actually perform software-development work through ROSIE.
The agent can operate against a real workspace, use local execution tools, run commands, modify software, and verify the results.
We also moved the architecture toward explicit runtime ownership instead of relying on global state.
That matters because an autonomous builder needs more than intelligence.
It needs a controlled place to act.
ROSIE provides that place.
The result is an early version of something much larger than a coding chatbot:
a software system where an AI agent can be given a job, operate inside a bounded engineering environment, and work toward completing it.
What we learned
Giving an AI more intelligence does not solve the execution problem.
A model can understand exactly what should happen and still be useless if it has no reliable way to act on the machine.
We learned that the bridge between reasoning and execution needs to be treated as first-class infrastructure.
State ownership matters.
Tool ownership matters.
Workspace boundaries matter.
Approvals matter.
Verification matters.
And the architecture has to assume the agent will eventually need to inspect, build, test, repair, and replace parts of the system itself.
That changed the question from:
How do we make an AI coding assistant?
to:
How should software be constructed when machines are expected to participate in its lifecycle?
What's next for ROSIE
The hackathon version of ROSIE proves the local execution path.
The next step is continuing to harden that runtime: stronger task isolation, deeper observation, better verification, clearer authority boundaries, and more complete lifecycle ownership.
The larger engineering system will continue evolving beyond the hackathon.
ROSIE herself will also eventually move into a different specialized role, while the local autonomous software-construction work continues under the broader system architecture.
For this submission, though, ROSIE's job is simple:
She gives HACKASS hands.
Log in or sign up for Devpost to join the conversation.