Inspiration
If you ship LLM features, the bugs that page you are not the ones a linter finds. A model id that gets deprecated in the middle of a function. No output limit. JSON parsed straight from a model. User text spliced into a prompt. A prompt nobody has ever measured.
Coding agents can write fixes for these. But the first time I let an agent work unattended, the problem was not capability. It was trust. It told me it had "updated the system prompt to separate instructions from data". The diff changed one line and never touched a system prompt. Another time it made a failing test pass by deleting behaviour the docstring promised.
An agent will tell you it did something, and the only defence is checking. So I built an AI engineer whose main job is to check itself, and to refuse to hand over work it cannot prove.
What it does
Journeyman is an AI engineering agent that works beside you and keeps working when you leave.
review: finds what a senior AI engineer would stop in a pull request, in Python and TypeScript: hardcoded model ids, nomax_tokens,json.loadson model output, prompt injection, no timeout, live-model tests, prompts with no eval, and Bedrock-specific failure modes. Each finding says why it matters and how to fix it. No model is involved, so it runs in CI and fails a pull request only on problems that pull request introduced.inventory: every model call in the codebase, which model or environment variable it uses, whether output is bounded, and whether the text leaves the machine.eval: turns a prompt with no eval into test cases with a held-out split, recorded responses, and a test that fails when quality drops.pair: while you work, runs only the tests your save affects, and stays silent unless something goes red.shift: while you are away, it takes the most important broken thing, fixes it in a throwaway git worktree, verifies the result itself, and either commits to a new branch or refuses and says why.proposewrites the pull request description from the evidence. It never pushes or merges.- On AWS: the reviewer is deployed on Amazon Bedrock AgentCore Runtime, where a Strands agent on Claude explains the findings, so a team can use it from CI or chat without installing anything.
How we built it
Strands Agents is the core:
- The worker is a Strands
Agentwith@toolfunctions bound to an isolated git worktree. It can read, edit, run tests and see its diff. It cannot push, merge, install packages or leave the worktree, and those refusals are tests. - Budgets (model turns, commands, paid calls) are enforced by a Strands
HookProvider. A watchdog timer callsAgent.cancel()so the wall-clock limit holds even in the middle of a generation. - Verification is code, not a prompt. When the worker says DONE, Journeyman runs the test suite before and after, checks the target test now passes or the finding no longer fires, compares what the agent claimed against the actual diff, and blocks changes that delete behaviour a docstring promises. If something is wrong, the worker is sent back with exactly what failed.
- The blind spec check is two more Strands agents that never see the change. They write tests from the task and the pre-change docstrings, and the change must pass them.
- Containment: every test the agent causes to run executes in the macOS sandbox, with no network, no writes outside the worktree, and no access to
~/.ssh,~/.awsor credential-looking environment variables. - Models through Strands model providers: Qwen3-Coder 30B locally via Ollama (free, private, used for every benchmark run) and Claude on Amazon Bedrock (the live demo and the AgentCore service).
- AgentCore: a
BedrockAgentCoreAppentrypoint wraps the deterministic review. A Strands agent on Claude explains the top findings through one read-only tool confined to the repository, and nothing from the repository is ever executed. It was deployed with direct code deploy in us-east-1.
Python, about 10,000 lines of code and 4,000 lines of tests (358 passing). The public git history starts on 13 September 2026.
Challenges we ran into
Honesty was harder than capability. Early shifts reported changes they had not made, deleted documented behaviour to satisfy a test, and ran with budgets that were declared but never enforced. Each failure became a check in code.
Silent failures everywhere. Ollama's default 4,096-token context silently discarded the task description. A test suite that could not even start printed no failures, so it read as green. Sandboxed tests ran under the wrong Python. A scheduled job was registered and ran zero times. None of these raised an error; each was found only by running the thing for real.
A checker that cries wolf is worse than none. The first blind checker raised false alarms on four of five correct fixes. Tuning it (float-safe assertions, module constants in the brief, ignoring failures inside third-party libraries, and requiring two independent checkers to agree) brought that down to one false alarm on 17 correct solutions.
AWS surprises. The first Claude model I tried was marked Legacy for my account. One newer model was listed as active but still returned "not available for this account". And a report line reading "running in us-west-2" revealed my agent was calling the wrong region.
Accomplishments that we're proud of
- Live on AWS. On a demo support-ticket app, the AgentCore reviewer returned six findings in 4.7 seconds, and with Claude's explanation in about 20 seconds.
- The same bug, two outcomes, both honest. On a local open model, the agent made the visible test pass, but the blind checker proved its fix threw away the system prompt, so it was not committed. On Claude in Bedrock, the fix passed the repo's tests and all ten blind tests in under a minute and was committed to a branch.
- Measured against hidden oracles the agent never sees (local Qwen3-Coder 30B):
- 12 routine bugs, three runs: 35 of 36 correct, no wrong fix delivered.
- 5 hard bugs, three runs: without the spec check, 3 of 15 correct and 10 wrong fixes delivered; with it, 6 of 15 correct and 4 wrong.
- We publish the bad numbers too. Hard bugs are still its weakness, and five cases repeated three times is a small sample.
What we learned
- Put the rules in code the model cannot talk its way past, and test the refusals.
- The agent's summary is a claim, not evidence. Verify the outcome yourself.
- When the tests that verify an agent's work were written by the same context that wrote the code, nothing has been verified.
- For work nobody is watching, "I did not commit this, and here is why" is a feature.
What's next
- A Linux sandbox so shifts can run on a team server, not only a Mac.
- Running the reviewer across a larger set of real open-source LLM applications and publishing its precision.
- Using the benchmark to decide when the blind spec check should be on by default.
What's next for REPLICA OF ME
Built With
- agents
- amazon-bedrock
- amazon-bedrock-agentcore
- amazon-cloudwatch
- amazon-web-services
- aws-cli
- aws-iam
- boto3
- claude
- git
- github-actions
- macos
- model-context-protocol
- node.js
- ollama
- pytest
- python
- qwen3-coder
- sarif
- strands-agents
- typescript
- vitest
Log in or sign up for Devpost to join the conversation.