Inspiration

Every Python developer knows the quiet dread of a dependency upgrade. A bot opens a PR bumping pydantic from 1.x to 2.x, and suddenly you're reading a migration guide, rewriting class Config blocks, and hoping your tests catch what you missed. The tedious part isn't the typing, it's the judgment: which upgrades are mechanical and safe to automate, and which ones genuinely need a human to think?

Most automated dependency bots ignore that distinction. They either blindly bump versions (and break your build) or flood you with PRs you don't have time to review. We wanted an agent that respects the developer's time and attention, one that quietly does the boring, mechanical work on its own, and interrupts a human only when there's a real decision to make. That's the heart of "Agents for Humans," and it became LusiScan.

What it does

LusiScan is an autonomous DevSecOps AI agent that takes the repetitive, judgment-heavy work out of Python dependency upgrades. On each run it:

  1. Monitors a target repository, parses its pyproject.toml, queries PyPI, and finds outdated packages (without installing anything).
  2. Plans each upgrade fetches the changelog and asks Amazon Nova for a structured migration plan: a confidence level, a strategy (auto_fix / guided_pr / human_required), an estimated risk, and the specific breaking changes it found.
  3. Executes safe changes, applies AST-aware refactors with libcst (never blind string replacement), bumps the manifest, and opens a pull request.
  4. Validates with real tests, triggers the repo's GitHub Actions workflow and reads the actual pass/fail result via the Actions API.
  5. Surfaces a decision to a human: persists the migration to DynamoDB and notifies the developer in their Slack channel (or as a PR comment). A Streamlit control panel shows the diff, confidence, risk, reasoning, test result, and PR link, with Approve / Review / Ignore buttons.

The two demo scenarios show both sides of its judgment:

  • requests 2.31.0 → 2.34.2: a safe bump. LusiScan applies it, opens a PR, and CI goes green.
  • pydantic 1.10.13 → 2.x: a breaking major upgrade. LusiScan classifies it human_required, leaves the code completely untouched, flags the breaking changes, and pings a human.

Very important, LusiScan never merges a PR without an explicit human approval. The agent proposes; the human decides. When you click Approve, the agent merges it on its next cycle, and never before.

How we built it

LusiScan is built entirely on the AWS-native agent stack:

  • Strands Agents SDK: the four stages (Monitor → Planner → Executor → Validator) are plain @tool functions that return structured dicts and are composed by ordinary Python calls. No opaque planner, no hidden control flow order and determinism matter for a code-changing pipeline.
  • Amazon Nova (via Bedrock) Nova Pro produces the migration plan (the one call where reasoning quality matters), and Nova Lite summarizes changelogs cheaply and at high throughput. Low temperature keeps the JSON output deterministic enough to parse reliably.
  • Amazon Bedrock AgentCore Runtime: hosts the agent loop. We wrote the orchestrator and wrapped it in a BedrockAgentCoreApp; deployment is agentcore configure + agentcore launch. No API Gateway, no servers to babysit.
  • DynamoDB: a single-table state store for migrations, decisions, and run logs, with a pending_review → approved/ignored → merged/closed state machine.
  • AWS Secrets Manager: the GitHub token and Slack webhook are read at runtime, never committed or baked into the image.
  • EventBridge Scheduler + Terraform: provisions the state table and a schedule that invokes the runtime on a cadence, all in one root module that tears down with a single terraform destroy.
  • Streamlit: the human control panel, a viewer that reads state and records decisions but runs no agent logic itself.

The safety core is a confidence gate: after the planner returns its strategy, human_required migrations skip the executor entirely (code untouched), while everything else runs the full refactor → PR → validate → notify path. And every AST transform is re-parsed and validated before it's ever committed if a transform produces invalid Python, we throw it away and fall back to a guided PR. We never emit broken code.

Challenges we ran into

  • The @tool decorator crashed the container on startup. Locally, Strands wasn't installed, so our @tool fell back to an identity decorator and everything imported fine. In the deployed container, the real decorator ran at import time and tried to build a Pydantic schema from every function signature and choked on our dependency-injection parameters typed as Protocol classes (PydanticSchemaGenerationError). The container never started. The fix: keep injected models as plain Any in tool signatures, since they were never meant to be LLM-facing inputs.
  • Decimal is not JSON serializable. This only appeared after we added DynamoDB persistence. DynamoDB returns every number as a Decimal, and the moment the agent read a stored migration (with a pr_number) back into its JSON response, the runtime threw a 500. We added a recursive deserializer at the store's read boundary to normalize Decimal → int/float.
  • A proxy that truncated Docker layers. agentcore launch's ECR push kept dying with write: broken pipe on one large layer through Docker Desktop's proxy. After two failed retries we switched approaches entirely to docker buildx build --push, whose BuildKit uploader handled the proxy cleanly.
  • Making AST refactoring genuinely safe. Confidence isn't correctness. Teaching the agent to validate its own output parse, transform, re-parse, and reject anything that doesn't round-trip — was the difference between a flashy demo and something you'd actually trust on your codebase.

Accomplishments that we're proud of

  • Both demo scenarios run end-to-end on a live AgentCore runtime — requests auto-fixed with green CI, pydantic correctly deferred to a human with zero code changes.
  • The human-in-the-loop loop actually closes. We proved the full cycle live: a decision recorded in Streamlit → the agent merges the PR on its next cycle → the bump lands on main. And we verified the guardrail holds: nothing merges without an explicit approved.
  • It genuinely runs on its own: secrets from Secrets Manager, a scheduler that fires it on a cadence, and idempotent re-invocation that reuses existing PRs instead of spamming new ones.
  • 342 passing tests covering the state machine, AST transforms, the confidence gate, the Actions poller, and the notifier: the safety-critical paths are locked down.
  • Least-privilege everywhere: scoped IAM roles for the runtime and the scheduler, no wildcards, and a single-command Terraform teardown.

What we learned

  • The environment where you test and the environment where you deploy can disagree about what your code does. Our @tool guard behaved differently with and without the SDK installed a lesson we only learned in production.
  • Restraint is a feature. The best decisions in LusiScan were all about doing less on purpose: return the original source rather than a broken one, downgrade to guided instead of guessing, hand off architectural calls, never merge without a yes. An agent for humans isn't the one that needs you least it's the one that knows exactly when it needs you.
  • Structured outputs beat free-form text for pipelines. Having every stage return a typed dict (instead of parsing model prose) made the whole system testable and deterministic.
  • AgentCore removes an enormous amount of deployment friction we spent our time on agent logic and safety, not on wiring up infrastructure.

What's next for LusiScan

  • More packages and ecosystems. The refactor pattern registry is currently scoped to the two demo packages; expanding it (and eventually supporting Node/Cargo) would make LusiScan broadly useful.
  • Richer changelog understanding. Generalized changelog discovery and diffing, beyond the hardcoded demo mapping, so the planner can reason about any dependency.
  • A hosted, multi-repo dashboard. A public Streamlit link watching many repositories at once, with per-repo policies for what auto-fixes are allowed.
  • Learning from decisions. Feeding a team's Approve/Ignore history back into the planner so LusiScan's confidence calibration improves over time for your codebase.
  • Interactive Slack actions. Approve or ignore a migration directly from the Slack message, closing the loop without leaving the channel.

Built With

  • amazon-bedrock
  • amazon-cloudwatch
  • amazon-dynamodb
  • amazon-ecr
  • amazon-eventbridge
  • amazon-nova
  • aws-codebuild
  • aws-iam
  • aws-secrets-manager
  • bedrock-agentcore
  • boto3
  • docker
  • github-actions
  • libcst
  • pygithub
  • pytest
  • python
  • strands-agents
  • streamlit
  • terraform
Share this project:

Updates

Submission history