Household Chief of Staff
Submitted to the All Things Agentic Hackathon - Taskmaster track.
Inspiration
Running a household is a constant background list of small decisions that never stop: has anyone checked whether the kids still fit last season's clothes? Is a gift sorted for the anniversary next week? Did the grocery reorder go out? None of it is individually hard - but doing it well means remembering past purchases, checking prices, and staying inside a budget, and that mental overhead never really goes away. The idea was an agent that carries that background load: it watches, decides what needs to happen, justifies its reasoning, and only acts once a human says yes.
What it does
A daily "sweep" agent that checks a family's data across four categories and proposes what needs to happen next - grounded in real data, not generic suggestions:
- Wardrobe - flags a kid whose sizing data is stale (90+ days) or who's crossed a growth threshold, proposes a capsule purchase sized to their latest measurements.
- Gifts - flags a birthday/anniversary within 14 days, proposes a gift idea within budget that explicitly avoids repeating a past year's gift for that person (checked against real gift history, not assumed).
- Groceries - flags items overdue for their usual reorder cadence, proposes the standard reorder.
- Travel - fetches a live flight price per watched trip, proposes booking when the current price is at or near its historical low.
Every proposal passes through a deterministic budget guardrail - plain Python, not the LLM - before anything is written or sent. The model proposes; code enforces the budget. A proposal is never silently dropped for being over budget; the user is still emailed, flagged explicitly as needing an override. On approval (one click from the email, no login required), a simulated transaction is logged to Firestore and the category budget updates
- nothing here moves real money or makes a real purchase/booking.
A single-page admin dashboard (added after the original scope, on request) gives a live view of everything that would otherwise mean digging through the Firestore console or individual emails: what needs approval by category, today's activity, items trending toward qualifying soon, recent history, and a budgets-this-period view with progress bars that shade in what's pending approval - so it's visible at a glance whether approving everything queued would push any category over budget.
How I built it
- Gemini (via Vertex AI) for each category's reasoning and justification text.
- Google ADK for multi-agent orchestration - one
LlmAgentsub-agent per category (Wardrobe, Gifts, Groceries, Travel), each given only the already-qualified candidates for its category. Trigger rules (does this even need looking at?) run first in plain Python - the LLM never decides whether to look at something, only what to say once it's qualified. - Cloud Run hosts the HTTP service (
/sweep,/approve,/reject,/admin) - deployed public (--allow-unauthenticated) because the Approve/Reject links are clicked from a plain browser tab opened out of Gmail, which can't carry a Cloud Run IAM identity token; the real security boundary is a random per-transaction approval token instead./sweepis separately gated by a shared-secret header, and/adminby HTTP Basic Auth, since it's a broader attack surface (full family data, approve/ reject without a per-transaction token) that a human loads in a browser. - Firestore holds family data, budgets, transaction/approval log, and price history - no dedicated DB cluster needed.
- Cloud Scheduler fires the daily sweep.
- Gmail API sends the approval-request email and is genuinely wired up (not mocked) - the one real external action in an otherwise fully simulated execution model.
- Secret Manager for every credential (SerpAPI key, Gmail refresh token, sweep secret, admin password) - nothing committed to the repo.
- Lyria 2 (
lyria-002, via Vertex AI on the same project) generates the demo video's background score - a second Google AI model alongside Gemini.scripts/build_music_lyria.pycalls the Lyria predict endpoint with three distinct ambient-underscore prompts, crossfades the ~30s clips into a loopable bed, and the video build sidechain-ducks it under the narration. Used in video production only - not wired into the agent runtime.
Other data sources: SerpAPI for live current flight/shopping prices. Historical price trends (used to judge "is this near the yearly low?") are synthetic seed data standing in for a production time-series pipeline - no retailer scraping, stated explicitly rather than left ambiguous.
Challenges I ran into
- Distinguishing a real automatic sweep from a manual test one. With Cloud Scheduler live and unpaused during ongoing development, a genuine daily sweep fired mid-session and created real proposals/emails - a same-day cleanup pass meant to clear manual test data wiped those out too, since nothing distinguishes "test data from this session" from "a real automatic run that happened to land the same day." Concrete lesson for any agent with a live scheduled trigger: test-data cleanup needs to be aware of what's actually real, or the scheduler needs pausing before iterating.
- Auth boundary design for the public endpoints. The public service can't use Cloud Run IAM for the email-clicked Approve/Reject links, so security had to move to a per-transaction random token instead - and the admin dashboard needed its own separate gate (HTTP Basic Auth, chosen over a query-string token so the browser handles the login natively without leaking the credential into browser history or server logs), since it's a strictly larger attack surface than a single-transaction link.
- Keeping the LLM out of the budget decision. Early design risk was the agent "talking itself into" skipping a budget check via its own reasoning. Fixed by making budget enforcement a hard deterministic gate in plain Python that runs after the LLM proposes and before anything is written or sent - the model never gets a vote on whether its own proposal is affordable.
Accomplishments I am proud of
All four categories built end-to-end and verified against the live deployed service - not just locally: a real sweep run, real Gmail approval emails with justifications grounded in actual Firestore data (verified by reading them, not just diffing a template), a real approve/reject round-trip with Firestore state changes and budget updates on screen, and the "avoid repeating last year's gift" rule genuinely checked against seeded gift history.
What I learned
Deterministic code around an LLM's edges - trigger rules before it, a budget guardrail after it - turned out to matter more for trustworthiness than the LLM's reasoning quality itself. The agent is only as reliable as the boundaries it's not allowed to reason its way out of.
Log in or sign up for Devpost to join the conversation.