Inspiration

We signed up for TikTok TechJam because we wanted to experience what it is like to code for change and do something meaningful to test our skils. We wanted to build beyond what felt familiar. The scale of the event made that feeling even stronger. Seeing nearly 3,000 participants join a 72-hour challenge was slightly intimidating, but it was also deeply motivating. Thousands of students were looking at the same broad set of problems and would still produce entirely different interpretations of them. That sense of collective creativity was one of the most exciting parts of the experience. It like an invitation to test how far an idea could develop under pressure.

The idea for RunVault did not arrive as a complete solution. It emerged while we were exploring the original Agent Launchpad and observing how an Agent Run moved through the system. The Agent could receive a task, perform work, and return an output, but we repeatedly found ourselves asking the same questions:

  • What actually happened during this Run?
  • Which files changed?
  • Were those changes safe?
  • Did the tests run?
  • If something went wrong, what state was the workspace left in?
  • What information would a person need before trusting the result?

We realised that the problem was not simply whether an Agent could complete a task. The larger issue was that we had very little visibility into the space between sending a prompt and receiving an answer. A Run could appear successful while still leaving behind changes that deserved closer attention. Conversely, a useful proposal could be rejected without giving the user enough information to understand or improve it.

That gap became the foundation of RunVault. We wanted to make every Run observable, reviewable, and reversible before it became trusted work. More than blocking risky actions, we wanted to give the person using the Agent enough context to make a meaningful decision.

How we built it

We began by mapping the complete journey of a Run, from the moment a prompt is submitted to the moment its result becomes part of the workspace. Looking at the whole journey first helped us identify where information disappeared and where a decision boundary was missing.

From there, we broke the system into smaller questions. Where should the Agent work? What evidence should be collected? What should happen when tests fail? Which changes should require attention? How should a user approve, discard, or revise a proposal without losing context?

RunVault now allows an Agent to work in a separate staged workspace. Once the Run is complete, the proposed changes are inspected, available tests are evaluated, and the result is classified before it can affect trusted state. Safe work can move forward automatically, while uncertain work remains available for human review.

We also built a review experience around this process. Instead of presenting only a final Agent response, RunVault shows the workspace outcome, verification status, findings, changed files, revision history, and operational evidence. Users can approve the proposal, discard it, or request a revised version. A separate history view makes it possible to return to earlier decisions and understand how the workspace evolved over time.

Technically, we extended the Volc Agent Launchpad using React, TypeScript, Fastify, Codex CLI, Volcengine Ark, and container-based execution. However, our main focus was not simply connecting these technologies. It was ensuring that the entire experience from the interface to the final workspace decision felt coherent and understandable.

Challenges we faced

The most obvious challenge was the 72-hour time limit, but the harder part was deciding what deserved that limited time.

At several points, we had multiple directions open at once. We could improve the interface, add another policy rule, strengthen recovery, expand the evidence, write more tests, or refine the demo. All of them felt important. We had to stop, step back, and ask ourselves what would most clearly support the central idea.

We also underestimated how difficult it would be to communicate the result simply. The underlying behaviour can become complicated very quickly, but the user should not need to understand every internal detail to know whether their workspace changed and what they can do next. Writing the interface copy, README, architecture explanation, and demo forced us to confront places where our own understanding was still unclear.

What we learned

The most important thing we learned was how to approach a large, unfamiliar problem without being overwhelmed by it. We learned to decode it bit by bit, validate each smaller assumption, and gradually reconnect those pieces into a complete picture.

At the same time, working piece by piece only helps if the top-level view is preserved. We became much more conscious of asking whether each task contributed to the actual user journey or simply felt interesting to build.

We also learned that information itself can be a meaningful product feature. Before RunVault, the Run produced an answer, but much of the reasoning needed to trust that answer remained invisible. By exposing outcomes, findings, history, and possible next actions, we were helping the user understand what the Agent had done.

Finally, this experience changed how we think about building with AI agents. Capability is valuable, but capability without visibility can leave the human outside the decision-making process. RunVault reflects what we came to care about during the hackathon: allowing an Agent to move quickly while ensuring that the person using it never loses context, control, or the ability to reconsider what happens next.

Built With

+ 18 more
Share this project:

Updates

Submission history