Inspiration

Every home has appliances that quietly need maintenance — HVAC filters, water heaters, smoke detectors — but almost nobody tracks it until something breaks. By then it's an emergency repair, a rushed replacement decision, or living with a failure that could've been avoided. When I saw the Agents for Humans hackathon's Everyday Agents track, this felt like exactly the kind of repetitive, judgment-heavy busywork an agent should quietly handle in the background — not another dashboard to check, but something that only speaks up when there's an actual decision to make.

What it does

Maintain-AI tracks a household's appliances — brand, model, install date — and runs a daily check against known maintenance intervals. When something is due or overdue, it doesn't stop at a reminder: a Cost Estimator sub-agent weighs repair cost against replacement cost and recommends which makes more sense. Critically, the agent can inform freely, but it can never act — submitting a repair/replace decision is gated behind explicit, per-appliance human approval. Nothing is recorded as decided until a person says so.

Interval lookups check a structured table first; anything not covered falls back to a RAG lookup over appliance manuals, so the agent is never just guessing.

How we built it

The core is two agents built on the Strands Agents SDK, composed with the Agent-as-Tool pattern: an orchestrator that handles tracking, checking, and notifications, and a Cost Estimator it invokes as a tool for the repair-vs-replace decision. OpenAI models power the reasoning and tool-use for both agents.'

From the start, I designed every external dependency — model, storage, vector store, notifier, trigger — behind a clean interface, deliberately decoupled from the agent and tool logic. That let me choose a fast, lightweight deployment stack for the hackathon timeline: Railway for the agent runtime and daily cron, Neon Postgres for appliance state, Chroma Cloud for vector search on the RAG fallback, and Resend for email notifications. An AWS/Bedrock-based architecture is documented as an alternative deployment path, since the interfaces make swapping infrastructure a config change, not a rewrite.

For visibility, Strands' OpenTelemetry instrumentation exports every chat turn and tool call to Honeycomb, and a Next.js dashboard on Vercel streams the agent's tool-trace live over WebSocket, alongside a pending-approvals panel for the human-in-the-loop gate.

I used Claude as an AI pair-programmer throughout the build — for scaffolding the agent and tool code, working through the deployment architecture, and reviewing the human-in-the-loop design — while making the product and architecture decisions myself.

Challenges we ran into

The real challenge was getting the human-in-the-loop gate right. My first pass only gated the cost recommendation, but on review that wasn't enough — the agent could still inform freely without gating the action that actually mattered. I rebuilt the gate around submit_maintenance_request() itself, using Strands' HumanInTheLoop intervention handler, and verified it against real multi-appliance batches — approving some, denying others, in a single round trip.

The RAG fallback took similar care: getting the structured-table-first, RAG-second lookup order right, and adding a cache-back step so a successful fallback lookup doesn't repeat the same expensive vector search for the next request, meant thinking carefully about the failure modes rather than just wiring up a vector store and calling it done.

Accomplishments that we're proud of

  • Both agents working end-to-end, with structured lookup, RAG fallback, and cache-back on a successful fallback hit
  • A genuine human-in-the-loop approval gate, not just a UI checkbox — verified against real approve/deny scenarios
  • Full observability: OTel tracing to Honeycomb, plus a live tool-trace dashboard for real-time visibility into agent reasoning
  • 47 backend tests passing, including coverage guarded against a live database
  • A deployment architecture where every service sits behind a swappable interface, with a documented AWS alternative that never required touching agent code

What we learned

Designing behind clean interfaces isn't just good practice — it's what made the deployment stack a genuinely open decision rather than a locked-in one. I also learned that "human-in-the-loop" is easy to say and easy to under-build; it took a real second pass to find the actual action that needed gating, not just the recommendation that preceded it.

What's next for maintain-ai

  • Replace the mock manual excerpts with real manufacturer documentation
  • Appliance-specific fulfillment: contractor and vendor integrations so an approved decision can turn into an actual booked service
  • SMS notifications alongside email
  • Multi-household support for property managers tracking several homes at once

Built With

Share this project:

Updates

Submission history