Watch the final public demo — 3:24 · Try Society Relay · Judge guide
Inspiration
“Any update?” is the unofficial operating system of many apartment societies. A-304 needs water restored. A-502 cannot leave work for a revised visit. The committee has a shared budget to protect. These fictional households make the coordination problem concrete: getting a repair done must not mean taking people's decisions away.
Society Relay follows a shared problem through to a verified outcome. Its defining rule is that a vendor's completion report cannot close an issue. The residents affected must verify restoration.
What it does
The Good Neighbor demonstration connects two households reporting a water-supply problem in Tower A. It preserves their original reports, proposes a fixed INR 1,800 repair, and finds an access window compatible with both households, the vendor, and society quiet hours.
Then the demo makes life difficult:
- A resident changes availability. The common window disappears and the old proposal cannot be approved. Relay asks for a one-visit exception. The resident declines. The agent respects that answer and asks a different household, without rewriting either household's general availability.
- The technician reports a delay. Sentinel checks the saved milestone and brings a decision back to the committee without increasing the quote.
- The missed visit needs a new date, fresh one-time consent and new committee approval. An old yes cannot authorize a different visit.
- The facility team reports completion. One household confirms restoration; the issue remains open. Only the second household's confirmation closes it. Residents can also dispute completion and reopen attention.
A committee member can separate a mistaken report match before authorization; that correction prevents automatic remerging. Approved plans produce downloadable demonstration work orders and calendar invitations.
Who does what
Residents retain access consent and restoration checks. The committee retains spending approval. Relay interprets reports, coordinates feasible alternatives, tracks missed milestones, and preserves evidence behind each decision. The interface labels whose turn it is and compares the shared responsibilities with manual coordination. We do not claim measured time savings.
How we built it
A Strands Graph runs Sensemaker, Coordinator and Sentinel. An evidence specialist uses the model to classify original reports before validated groupings are applied. The optional gateway path additionally performs an independent model review of each proposed link. Follow-ups with no new reports can enter directly at Sentinel.
The model interprets reports and invokes tools. Deterministic Python rules enforce spending authority, current proposal identifiers, access constraints, human corrections, and resident verification. Versioned workspace writes recheck rules after concurrent changes. One-time consent is bound to the exact date, household group, vendor, quote and constraints, and can be withdrawn before authorization.
A public AWS Lambda Function URL serves our Python/FastAPI application and original HTML/CSS/JavaScript interface. Cloudflare provides the public domain. A private Lambda dispatcher invokes an IAM-protected Bedrock AgentCore runtime using Amazon Nova Pro. DynamoDB stores workspaces, conditional version writes, durable job leases and usage counters. EventBridge Scheduler checks for due work every minute, even with no browser open.
The shared demo is limited to 100 workspaces, 100 dispatched jobs and 150 model calls. These are workload controls, not a hard dollar spending cap. Optional local SQLite and Free AI gateway adapters remain available for development.
Try it
Choose Start the difficult repair on the live community desk. The guide reads actual saved state, makes human decisions explicit, and stops for unexpected model decisions instead of fabricating progress. No login is required; use fictional information. Failed runs preserve completed work and require a manual retry.
For an immediate walkthrough, replay a complete verified AWS run. Fifteen checkpoints show a declined request, an alternative involving another household, a missed visit, fresh consent, and separate confirmations. Five actual Nova Pro runs took 13.15, 6.80, 15.97, 3.73 and 14.59 seconds. The INR 1,800 quote and general household availability stayed unchanged. This clearly labeled recorded evidence exposes saved state and tool calls without consuming AI capacity.
The final 3:24 video combines edited September 10 live-application captures, continuous inspection of saved AWS evidence paced to narration, and explanatory graphics. The replay is not a new model invocation. Narration uses the stock Kokoro af_heart synthetic voice, not an imitation of a real person.
What we verified
The AWS acceptance test verified actual Nova Pro tool calls, duplicate linking, exact-proposal approval, provisional vendor completion and closure only after both households confirmed. The initial agent run took 10.33 seconds. A separate workspace completed through the external scheduler without a manual agent request; the test waited 49.11 seconds including the next tick. States and checks are in docs/aws-live-acceptance.json. A separate DynamoDB proof verified stale-write rejection, conflict retry and committed state read by another process.
We froze eight difficult synthetic interpretation cases before execution and compared the deployed evidence prompt against three simple rules baselines. Nova Pro passed 6/8; every baseline also passed 6/8, with different mistakes. Nova returned an invalid empty group alongside a detected hazard and omitted a private report. Both failed component cases passed separately through the full AWS workflow; this does not replace the first-pass failures or establish model superiority. Cases, raw outputs, model request IDs and limitations are public.
The current build passes 55 Python tests and six JavaScript checks covering lifecycle authority, stale approvals, concurrency, report separation, missed milestones, false completion, retries, gateway compatibility and workspace isolation. Fault-injection tests include a provider timeout after a saved tool action, accept/decline races, simultaneous confirmations and stale consent on a new date. Saved work survives and private diagnostics are redacted. The manual recovery test explicitly uses a fixture. GitHub Actions runs the checks.
Six earlier targeted live-model cases passed, including unrelated reports, an embedded instruction attempt, no common access window, delay and negotiation after refusal. Their calls, timings, observed models and states are in docs/evaluations/. A separate hosted GPT-OSS-120B run completed the basic lifecycle in 71.28 seconds; an earlier provider-boundary failure is also recorded in docs/hosted-verification.json. These are bounded synthetic checks, not general reliability or real-user feedback.
Challenges and lessons
The gateway initially stripped native tool-call history fields. Our adapter preserves the conversation as text while leaving tool execution to Strands. Early model runs showed why confident prose is not enough: success must be checked against stored outcomes.
Report similarity is not proof of a common physical cause. Original evidence and durable human corrections matter. A changing schedule must invalidate stale approval, and a vendor's statement must remain provisional until affected households agree.
What's next
Authenticated resident and staff roles, scoped vendor integrations, reliable external notifications, and a consenting society pilot. We have not measured adoption or resident time savings. The judge-facing role switch is a demonstration tool, not production identity verification. No real communications, bookings, payments or emergency dispatch occur.
Build disclosure
Application code and original interface assets were created September 9–11, 2026 with assistance from OpenAI Codex. The existing inference gateway is an external service, not new hackathon work. No Fleet or client application code was incorporated. Source is MIT-licensed; dependencies retain their own licenses.
Built With
- amazon-bedrock
- amazon-bedrock-agentcore
- amazon-dynamodb
- amazon-eventbridge
- amazon-nova
- aws-lambda
- fastapi
- python
- strands
Log in or sign up for Devpost to join the conversation.