Inspiration

Challenge 10 of the hackathon asked us to solve something that developers regularly find annoying. We chose bugs because they are an unavoidable part of the development lifecycle. Even a small runtime error can interrupt someone’s work and force them to stop what they are doing, reproduce the problem, study the error, locate the correct file, write a fix, test it, and prepare a pull request.

We were also interested in how quickly BoxLang could be used to build agents and automation. That led us to Pipeline Doctor: an agentic workflow that can investigate a reproducible failure, attempt a repair, verify the result, and prepare a pull request automatically. The developer does not have to repeatedly copy errors into an AI chatbot or write prompts manually. Instead, they receive a proposed fix that is already organized for human review, allowing them to focus on other work.

What it does

Pipeline Doctor runs a project using a command defined in that repository’s configuration file. Our current MVP focuses on a small Java project containing an intentional NullPointerException.

When the command fails, Pipeline Doctor:

  1. Captures the exit code, standard output, error output, and stack trace.
  2. Uses the failure information to identify the relevant Java source file.
  3. Creates a separate repair branch so the original main branch remains untouched.
  4. Sends the error and source code to a Gemini-powered repair agent.
  5. Applies the proposed repair to the identified file.
  6. Runs the original command again to verify that the program now succeeds.
  7. Checks that only the expected file changed and rejects unsafe paths or unexpected files.
  8. Commits and pushes the verified repair.
  9. Opens a GitHub pull request containing the fix and a description of the problem.

Pipeline Doctor never merges the pull request automatically. The final decision always belongs to a human reviewer.

How we built it

We built the main application in BoxLang and separated the workflow into several focused components.

CommandRunner uses BoxLang’s process execution features to run the target project and capture its output, exit code, duration, and timeout information. FailureContext examines the stack trace and loads the relevant source code.

RepairAgent receives that structured failure context, creates a prompt containing the error and source file, and communicates with the Gemini API. It returns a diagnosis, summary, and list of files that it changed.

SafetyGuard checks whether the repair passed verification and compares the files reported by the agent with the files Git actually detected. It rejects sensitive files such as .env, absolute paths, path traversal attempts, failed verification, and unexpected changes.

Finally, GitHubPublisher creates the repair branch, stages only approved files, creates a commit, pushes the branch, and uses GitHub CLI to open the pull request.

We also created a separate pipeline-doctor-demo repository containing the intentionally broken Java application. Keeping the test project separate helped us prove that Pipeline Doctor could inspect and modify another repository instead of only working on its own source code.

Challenges we ran into

Our biggest challenge was that this was our first time building a complete project with BoxLang. We had to learn its syntax, process execution behavior, file handling, and runtime functions while developing the project.

We also divided the system into three separately developed parts. Each part worked on its own, but integration exposed small differences in the data passed between them. For example, one component used success while another expected succeeded, and one returned safeToVerify while another expected repaired. We also had to preserve the full path scr1/App.java instead of passing only App.java. These issues taught us that agreeing on shared data contracts is just as important as making each individual component work.

Git and GitHub automation also introduced challenges. We needed to authenticate GitHub CLI, handle SSH authentication, prevent existing branches from being overwritten, and make the demo repeatable. We spent additional time ensuring the system would never accidentally stage every changed file or publish secrets.

Another challenge was making AI behavior reliable. API errors, rate limits, or an unexpected response format can cause an AI repair to fail, so we added clear error handling and a limited fallback for our specific demonstration.

Accomplishments that we're proud of

We are proud that Pipeline Doctor performs real development actions instead of only displaying a simulated result. Our BoxLang code can execute another project, inspect a genuine runtime exception, create a real Git branch and commit, push it to GitHub, and open a real pull request.

We are especially proud of the safety checks. The goal was not simply to let AI modify code, but to make the result reviewable and prevent it from quietly publishing unrelated or unsafe changes.

We also successfully combined work from three teammates who focused on different stages of the pipeline. This gave each of us ownership of a major part of the project while still requiring us to understand how the complete system works.

What we learned

We learned how to build applications and automation tools with BoxLang, execute external commands, process error output, work with files safely, and connect an AI model to a larger development workflow.

We also gained a better understanding of Git branches, commits, GitHub authentication, and automated pull-request creation. Most importantly, we learned how important integration testing is. Three components can all pass their own tests and still fail when connected if they disagree about a field name, file path, or expected result.

What's next for Pipeline Doctor

Our current MVP is intentionally limited to one reproducible Java runtime failure and one changed source file. In the future, we would like Pipeline Doctor to support additional programming languages, compiler errors, failing unit tests, build tools, and repairs involving multiple related files.

We would also like to run it automatically through GitHub events or CI pipelines, provide clearer diffs and confidence information, support additional AI providers, and place repairs inside a stronger sandbox.

Even as the project grows, we would preserve its most important rule: Pipeline Doctor can prepare and verify a repair, but a human must review and approve it.

Built With

Share this project:

Updates

Submission history