Orville — one report, one verified handoff
Inspiration
Support work doesn't fail loudly. A customer reports a bug, someone opens the right engineering issue, someone else makes a follow-up card, someone tells the team in chat — and afterwards nobody can prove all three happened, or that they point at the same report. The work is spread across three tools, and the coordination is the part that breaks.
We wanted to build the version of that handoff an agent could actually be trusted with. "An LLM that calls an API" is an easy demo; the hard question is what would make you let it write to your real GitHub, Trello and Discord. Our answer became the project: the agent picks the steps, code enforces the limits, and nothing is called complete until the records have been read back.
The theme shaped it too. Agents for Humans — the human keeps the judgement, the agent does the coordination, and the system should be honest about which is which.
What it does
One messy bug report goes in. Orville:
- Inspects the live GitHub issue candidates through a typed tool.
- Selects the matching issue — or stops for a person if the match is ambiguous.
- Attaches the report to that issue as a GitHub comment.
- Creates one linked Trello follow-up card for the customer.
- Posts one Discord team update linking to both records.
- Reads all three records back, and only then reports
complete.
That is three writes across three apps on the existing-issue path. If the report also asks for unrelated destructive work, the request is recorded as refused.
How we built it
Python, with the Strands Agents SDK as the orchestration engine. The model is openai/gpt-oss-120b served by Groq through its OpenAI-compatible endpoint, wired in with Strands' OpenAIModel and a SequentialToolExecutor so side effects cannot run in parallel.
The tool surface is deliberately small and typed: inspect_report_context, select_issue, record_engineering_handoff, create_customer_followup, publish_team_update, request_human_review. Tool arguments cannot carry credentials, destinations, repository names or report bodies — those are injected from trusted run context the model never sees.
Under the tools sits a guarded operations layer. Each app stage is a separate gated operation: GitHub must be verified before Trello is attempted, and both before Discord. Every write persists its returned ID immediately, then performs an independent read that checks the marker, the destination and the cross-links (the Trello card must contain the verified GitHub link, the Discord message must contain both). Later steps are deferred until the records they link to are verified, so a status message cannot claim an unverified step succeeded.
The original single-pass CLI runner is still in the repository and still passes its tests. It is the regression baseline the new runtime was measured against, not something we quietly deleted.
The guard is the product
Four rules did most of the work.
The model proposes, code decides. The selected issue ID is validated against the candidate list fetched from GitHub at that moment. A fabricated ID, a confidence below the floor, or a flagged ambiguous match never reaches a write.
Reports are data, not instructions. Report text is scanned for delete, close, archive and unrelated-modification requests, which are recorded as refused. More importantly, the execution layer has no such operation to call — the model cannot invoke what does not exist.
Uncertainty pauses. If a Discord send outcome is unknown, the run marks it uncertain and stops for manual reconciliation. There is deliberately no approve-and-resend control, because resending is how you get duplicate customer-visible messages.
Completion is computed. complete comes from the read-back evidence in code. The model's own summary cannot set status.
Challenges we ran into
Uncertainty is harder than failure. A failed write is easy — you see the error. A write whose response was lost is dangerous: retrying may duplicate it, and not retrying may drop it. We handled it by persisting intent and IDs around each call, marking uncertain outcomes explicitly, and refusing to resolve them automatically.
Our own prompt was wrong about ambiguity. The system prompt told the model to call request_human_review for an ambiguous match, but that path persisted the pause as manual reconciliation, which the resume command could not resume. The model was doing exactly what we asked and hitting a dead end. We found it while building the offline human-review example, and fixed it so a pre-selection review is resumable while post-selection uncertainty still needs a person.
Provider friction is real. Groq's OpenAI-compatible endpoint emits repeated reasoningContent is not supported in multi-turn conversations warnings during the Strands tool loop. We confirmed the loops and read-backs still complete and verify, disclosed the limitation instead of hiding it, and did not swap providers late in the build.
Environment archaeology. A broken Python 3.11 virtual environment, temp directories on a synced drive causing file-replace failures, and a disk that filled up and killed an asset render mid-run. None of it was interesting; all of it was in the way.
What we learned
- Tool-calling agents deserve the same suspicion as any untrusted input. Validating the model's arguments in code is not paranoia, it is the design.
- Read-back verification is the line between a demo and something you would let near real systems. It also changed how we wrote tests: the interesting cases are the ones where the model sounds finished but the evidence says otherwise.
- Binding an immutable report and destination scope to a report ID makes retries safe and cheap. Running the same report twice reused all three records with zero duplicates.
- Honest disclosure costs nothing and buys credibility. The provider warning sits in the README next to the proof, not buried at the bottom.
Evidence and honest limits
The demo report is a sample support ticket. The app records are real.
The current connected proof is HARBOR-STRANDS-05: a fresh run that selected GitHub issue #1 and produced comment 5667869897, Trello card 6aa82c5d3a1ff6da7c9129b6 and Discord message 1549107085609013442, each read back and verified. Running the identical command again reused those exact three records and created nothing new, with marker counts staying at one across the repository and the list. The terminal output of both runs is saved in the repository. An earlier connected proof (HARBOR-STRANDS-04) and its retry are recorded as well.
42 local tests cover the failure paths with in-memory fakes: out-of-order tool calls writing nothing, false model completion staying incomplete, fabricated IDs rejected, changed input rejected, human-choice resume, and the connected-run scope. They are not a model-accuracy benchmark, and we do not quote an accuracy number.
It is a prototype. State is a local JSON file, GitHub searches are bounded, and there is no authenticated public execution service — you run it with your own credentials. Ambiguous matches go to a person by design.
What's next
An authenticated service so a team can run it without a local checkout, real reconciliation for uncertain sends instead of manual handling, and a small review interface for the human decisions that currently happen on the command line.
Credits
Built for Agents for Humans, September 2026. Strands Agents SDK with openai/gpt-oss-120b on Groq. GitHub, Trello and Discord via their REST and webhook APIs. The landing page and demo video were produced with local tooling. AI coding assistants were used for implementation and review; the architecture, guard design and tests are ours. MIT licensed.
Log in or sign up for Devpost to join the conversation.