Inspiration
A cron job that fails loudly gets fixed by lunchtime. Someone gets paged, someone reads a stack trace, someone ships a patch.
The dangerous one exits 0. It turns the dashboard green, writes nothing, and does that every night for four months. Nobody finds out until somebody asks for the export and it turns out to be an empty file.
Every monitoring tool I have used watches the job. Almost nothing watches the evidence. The check that would have caught this, "is there actually a non-empty CSV in there from the last six hours", is different for every single job, which is exactly why nobody writes it. Deadman exists because that check is the thing you already didn't write.
What it does
You register a job with two sentences of plain English: what evidence of success looks like, and how you would re-run it by hand.
Job(
name="nightly-export",
container="deadman-fleet",
workdir="/out/export",
evidence_claim="a non-empty CSV under /out/export, modified in the last 6 hours",
rerun="/bin/nightly-export",
)
No schema. No cron expression to re-declare. No assertion DSL to learn.
Deadman works out for itself which observations would settle that claim, makes them inside a sandbox, and reports only the jobs that are green and dead. A healthy job gets one quiet line, or none at all.
$ deadman
DEAD nightly-export
export-2026-09-10.csv is 0 bytes
fix: /bin/nightly-export (needs approval)
DEAD db-backup
backup.tar.gz is real, and four days stale
4 checked | 2 green and dead | 2 alive
It exits non-zero when anything needs attention. A watcher that always exited 0 would be the very thing it exists to catch.
Ask it to repair something and a human decides:
$ deadman nightly-export --fix
Deadman wants to run remediation for 'nightly-export': /bin/nightly-export
Approve? [y/N] y
remediation approved and executed - re-checking from scratch
alive nightly-export
Answer n and the file is untouched, same zero bytes and same timestamp, and the audit trail records REFUSED run_remediation - CONFIRMATION_FAILED. Answer y and it runs the command, then throws away its own opinion and investigates again from scratch. The agent does not get to grade its own repair.
How we built it
Strands Agents on Amazon Bedrock, with the safety properties pushed underneath the model instead of written into the prompt.
There are two intervention handlers, and the contrast between them is the design. Authorization sets on_error = "deny", so if the guard itself raises, the tool call is refused. AuditTrail sets on_error = "proceed", because a broken logger must never stop an investigation. Those two lines encode the whole posture, and the framework enforces them rather than the prompt asking nicely.
Remediation is unreachable until somebody asks for it. During an investigation run_remediation is not merely denied, it is never registered, so the model is not told it exists. --fix registers the tool and arms the handler, which then returns Confirm and stops the agent with stop_reason: interrupt. Nothing proceeds until a person types. A latch keyed on the tool call means a refused request cannot come back around as a second prompt.
Generated shell lands in a container, never on the host. DockerSandbox subclasses PosixShellSandbox with one primitive, execute_streaming, implemented as docker exec into the job's own container: no network, no capabilities, no host mounts, no-new-privileges.
Every observation is scope-checked. list_directory, stat_path and read_head each take a path, and anything outside the registered workdir is denied. Traversal, sibling-prefix collisions like /out/export-secrets, and relative paths, which are ambiguous and so refused rather than guessed at. The rule is a pure function, tested without a model, a container, or a network.
Verdicts come back through a per-call structured_output_model, with the text contract as a fallback. It is deliberately not set on the fix turn, because forcing a tool choice there would let the model satisfy the request by describing the repair instead of calling run_remediation, and the approval gate would never appear.
The model is amazon.nova-lite-v1:0, under a customer-managed IAM policy granting InvokeModel and little else, rather than AmazonBedrockFullAccess, which would permit provisioned throughput.
Challenges we ran into
The agent had no clock, and I blamed the wrong thing first. Deadman called a healthy log-rotate job dead. The obvious suspect was a stale test fixture. Wrong. Every evidence claim is relative to a moment in time, "in the last 6 hours", "from today", and nothing was supplying now: not the tools, not the system prompt, not the SDK. The model was guessing what today was, and two files with identical timestamps got split alive and dead. The fix is to read the clock inside the container, the same clock the file timestamps come from, and hand it to the agent. My host runs UTC+3 and the container runs UTC, so a fixture seeded at ten past midnight local was stamping files with the previous calendar day. That is a class of bug an agent cannot reason its way out of without a reference.
The approval prompt fired twice. The model reached for run_remediation during the investigation turn, before anyone asked it to. Prompt wording would have been a wish. The fix was structural: unregistered tool, armed flag, per-call latch.
The audit trail was missing the only event worth auditing. It logged tools completing, but never the remediation request or the human's refusal, because the intervention dispatch stops at the first handler that decides, so the audit's before-hook never ran for exactly that call. Refusals are now recorded from the after-hook, which carries the cancellation reason.
Raw chain-of-thought reached the terminal, and once a paragraph of prose was offered as a "fix". Now nothing reaches the terminal unparsed, and a command that is really an English sentence, or that contains a placeholder like /path/to, is refused rather than printed.
Accomplishments that we're proud of
Both halves of the safety claim are demonstrated rather than asserted. Deny, and the file is byte-for-byte unchanged. Approve, and /bin/nightly-export actually runs, the empty file fills with real rows, and a fresh investigation reports it alive. Not the agent's own say-so, a new look at the evidence.
A verdict that cannot be parsed is reported unreadable, never alive. A watcher that fails open is worse than no watcher.
The security boundary is tested as a pure function, with no model, no container and no network in the loop, which means the part that matters most is also the part that cannot flake.
What we learned
Exit code 0 is not proof of work, and that turns out to apply to the agent as well. Anything reporting on its own success needs an independent check, which is why an approved repair triggers a re-investigation instead of a success message.
Rules written in a prompt are requests. Rules written as interventions are rules. Every time I tried to fix behaviour with wording, the model eventually did the other thing.
And one about debugging: the obvious cause of a wrong answer was the wrong cause. Reading the SDK's own source settled in ten minutes what a plausible theory would have had me chasing for a day.
What's next for Deadman
Deploying on Bedrock AgentCore for a hosted demo. An eval suite over synthetic green-but-dead cases, so verdict accuracy becomes a number instead of an anecdote. Evidence backends beyond the filesystem: a row in a table dated today, an object in a bucket, a queue that drained. And a fleet-wide digest, because the moment this becomes useful is the moment you have ninety jobs and no idea which three are lying.
Built With
- amazon-bedrock
- amazon-nova
- amazon-web-services
- aws-iam
- docker
- pydantic
- pytest
- python
- strands-agents
Log in or sign up for Devpost to join the conversation.