-
-
The maintenance burden — Different issues compete for the landlord’s attention.
-
It starts with the tenant — Everyday problems reported in everyday language.
-
Autopilot at work — Understand the issue, apply the rules, determine what happens next.
-
Different problems, different decisions — Routine, unclear, costly and critical issues follow different paths.
-
What needs me? — A simple landlord view of what progressed, what needs attention and what is waiting.
-
Bounded autonomy by design — AI interprets the problem; deterministic policy controls the authority.
-
Giving the landlord back their attention — Routine work progresses while important decisions come to the surface.
-
A better tenant experience — Faster attention, clearer progress and less back-and-forth.
Maintenance Autopilot
Why I built Maintenance Autopilot
Maintenance is full of small decisions — and too many of them still need a human just to decide whether they need a human.
For a landlord, one message might say the kitchen tap is dripping. Another says a bedroom window will not latch. A technician sends a repair quote. Then a tenant reports a smell of gas.
Every message needs to go somewhere — but not every message needs the landlord.
That was the problem I wanted to solve:
What if a maintenance agent could handle more of the routine decision-making, while knowing exactly when the landlord still needs to be involved?
I came into this hackathon from Trinidad and Tobago with years of experience around maintenance and operations, but limited hands-on experience building AI agents.
That experience shaped the idea. In maintenance, good decision-making is not about automating everything. It is about keeping routine work moving while making sure risk, uncertainty and important decisions reach the right person.
Maintenance Autopilot became my way of exploring what that could look like for an ordinary landlord.
And as I built it, the problem became more interesting.
The important question was no longer just:
Can AI understand the maintenance problem?
It was:
Is AI actually authorized to make the decision that follows?
Giving the landlord back their attention
Maintenance Autopilot accepts a maintenance report in everyday language, interprets what is happening, and determines what should happen next.
Depending on the situation, it can progress routine work, ask for missing information, wait for a response, return a decision to the landlord, take the protective escalation path, or close something that no longer requires action.
Internally, those governed workflow decisions are represented as:
ACT · ASK · AWAITING · ESCALATE · ACT+ESCALATE · CLOSE
In practical terms:
- Routine problem? Progress it.
- Something important is missing? Ask for it.
- Already asked and waiting for information? Hold it.
- Needs the landlord's approval or judgment? Bring it back to them.
- Critical hazard? Take the protective path and escalate.
- Problem already resolved? Close it.
The Landlord View then brings those individual decisions together.
Instead of working through every maintenance conversation to determine what needs attention, the landlord gets a simple operational picture:
What came in? What progressed? What needs me? What are we waiting on?
The aim is simple:
Less time administering routine maintenance. More attention available for the decisions where human judgment actually matters.
The idea behind the agent
The difficult part was deciding how much authority AI should have.
I did not want a language model to read a maintenance report and simply choose whatever action sounded reasonable.
So Maintenance Autopilot deliberately separates interpretation from authority:
Use AI to interpret what is happening. Use deterministic policy to control what is allowed to happen next.
AI handles the messy language:
"The AC isn't cooling."
"The second bedroom window doesn't latch."
"There's a strong smell of gas and it seems to be getting stronger."
Using the Strands Agents SDK and Amazon Bedrock, the semantic pipeline interprets those reports through:
Issue decomposition → Hazard assessment → Case assessment → Repeat assessment → Validation
This produces structured, decision-relevant concepts such as hazards, conditions, information sufficiency, urgency, trade and repeat-failure signals.
Then AI stops being the authority.
Those concepts pass into a deterministic policy engine that applies explicit rules for safety, security exposure, repair authority, missing information, repeat failure, replacement decisions, property damage and resolved work.
AI handles ambiguity. Policy handles authority.
That separation is the core of Maintenance Autopilot.
What makes this different
There are already good reasons to apply AI to maintenance: understanding resident messages, categorizing problems, troubleshooting issues and automating workflows.
Maintenance Autopilot focuses on a different question:
Who decides how far the agent is allowed to go?
Many AI workflows focus on what an agent can do.
Maintenance Autopilot focuses equally on what it is allowed to do.
The language model is used where its strengths matter — interpreting ambiguous, everyday maintenance reports. But it does not determine its own authority.
A separate deterministic policy controls when work can progress, when more information is required, when a decision must return to the landlord, and when a potential hazard requires the protective path.
This means the probabilistic model does not both interpret the situation and invent its own permission to act.
The model interprets.
The policy governs.
The human retains the decisions outside the delegated boundary.
The goal is not maximum automation.
It is controlled autonomy: let routine work move without unnecessary human involvement, while keeping judgment, risk and authority boundaries explicit.
Decision rights with clear limits
The policy layer turns interpreted facts into explicit decision boundaries.
For example, the demo property gives Maintenance Autopilot delegated repair authority up to $200.
That is not the property's maintenance budget.
It is the limit of what the agent has been authorized to approve without involving the landlord.
A $195 repair may be entirely appropriate and within the agent's delegated authority, so it can progress.
A $205 repair may be equally appropriate — but the decision returns to the landlord because the agent's authority has ended.
The same principle applies beyond cost.
An unclear security issue can trigger a request for the information needed to make the decision.
A strong gas smell takes the protective path and returns the issue to human attention rather than waiting for an ordinary maintenance decision.
The point is not simply that the agent can make decisions.
It is that the agent knows the limits of the decisions it is allowed to make.
From decision engine to maintenance workflow
Maintenance Autopilot is deployed as a live end-to-end application on AWS:
Browser → AWS Amplify → Amazon API Gateway → AWS Lambda → Amazon Bedrock AgentCore → Strands-based V2.5 agent → Amazon Bedrock
Today, the process begins when a user submits a maintenance report through the live application.
That report can be written in everyday language. Autopilot interprets the issue, applies the governed decision policy, and returns a clear next step — including what it understood, what should happen next, and why.
But the web interface is only the current trigger.
In a fuller implementation, the same process could begin from an existing tenant portal, messaging channel, property-management platform or work-order system.
Maintenance Autopilot does not need to become another system of record to provide the decision layer.
And an individual decision is only part of the maintenance problem.
The Landlord View brings those decisions together into a simple operational picture:
What came in? What progressed? What needs my attention? What are we waiting on?
This is the experience I was aiming for: routine maintenance can move without demanding the landlord's attention, while the issues that need judgment, approval or intervention are brought clearly to the surface.
Building trust into the workflow
The same principle of bounded authority extends beyond the decision engine.
The public application includes:
- server-controlled property context
- strict request validation and input limits
- API throttling and restrictive CORS
- bounded retry and timeout handling
- response-contract validation
- privacy-conscious logging
- least-privilege AgentCore invocation for the public Lambda
- browser double-submit protection
There is also an important distinction between a decision and a real-world action.
If Maintenance Autopilot returns ACT, that means the work is authorized to progress.
It does not claim that a technician has been dispatched when no contractor integration exists.
That distinction becomes increasingly important as agents move from answering questions toward participating in real workflows.
The architecture evolved through testing
The final architecture was not where I started.
An early sealed 24-scenario evaluation showed that simply putting a governor around an AI decision was not enough. It corrected some decisions but also introduced regressions.
That result changed the design.
Instead of asking the model to reason directly toward an action and trying to constrain it afterward, V2.5 separates semantic interpretation from deterministic authority.
That became one of the biggest lessons from the project:
The authority boundary works better as part of the architecture than as another instruction in the prompt.
Tested beyond the demo
The four examples in the live application were selected because they make the main decision boundaries easy to see.
They are not hard-coded flows.
The same free-text system was exercised across a much broader range of residential maintenance situations during development — including plumbing, HVAC, electrical, appliances, access and security, water ingress, progressive damage, repeat problems, missing information, repair-authority boundaries and safety-related conditions.
Reports included everyday descriptions such as:
- "washing machine making a loud grinding sound when it spins"
- "I think there's a leak somewhere, water bill doubled but I can't see anything"
- "power's out in half the apartment, other half fine, no storm"
- "dishwasher's been leaking and now the laminate floor is swelling up"
- "toilet won't flush at all and it's the only bathroom"
- "there's mold in the bathroom ceiling corner, spreading"
- "garage door won't close, stuck halfway, car's stuck inside"
- "keep smelling sewage in the downstairs bathroom, comes and goes"
- "ceiling light fixture fell and is hanging by the wires"
- "the balcony railing feels loose when I lean on it"
The public application accepts the same kind of free-text input.
Try reports like these — or enter your own maintenance issue in ordinary language.
Testing whether it knows its limits
A handful of successful examples would not tell me whether the design was actually working, so I tested Maintenance Autopilot at three different levels.
Each test answers a different question, so I keep the results separate rather than combining them into a single accuracy number.
Development evaluation — 111 cases
Across the developed V2.5 evaluation set:
111/111 governed outcomes were correct.
These cases were used during development and regression testing. They demonstrate policy correctness across the developed test envelope, but I do not present them as an unseen final holdout.
Live deployment hardening — 40 evaluations
I selected four representative scenarios spanning the main decision boundaries and ran each one ten times against the deployed Amazon Bedrock AgentCore system.
40/40 returned the expected governed outcome and policy rule.
This tested whether the deployed system could consistently preserve the same governed decisions across repeated live evaluations.
Post-freeze challenge — 80 cases
For the final evaluation, I made the test harder.
The application was frozen first.
An additional 80-case holdout dataset and its expected outcomes were then locked before official execution, with the answer-key fingerprint and Git history preserving that sequence.
The official evaluation was completed once, with no post-result tuning or selective reruns.
Result: 63/80 — 78.8%
The system was not perfect.
But one result was particularly important for the bounded-autonomy objective:
All 15/15 holdout cases expected to require protective ACT+ESCALATE behavior were correctly governed.
That is not a claim that the prototype is universally safe.
It is a specific result within the locked evaluation: all 15 cases preregistered as requiring the protective path received that governed outcome.
The challenge also identified where the next version needs to improve, particularly clarification and non-critical escalation decisions where repeat failure, progression or missing information must first be recognized by the semantic layer.
I kept those misses rather than tuning them away.
That made the final evaluation more useful: it showed not only what Maintenance Autopilot could do, but where its current boundaries really are.
What the failures taught me
Building Maintenance Autopilot changed how I think about agent autonomy.
Deterministic policy does not make AI uncertainty disappear.
The policy can enforce a boundary only when the interpretation layer correctly understands the situation and surfaces the decision-critical facts that activate it.
The post-freeze challenge exposed this clearly.
Several misses occurred because the semantic layer did not correctly recognize a relevant signal such as repeat failure, progression or missing information.
Once incomplete or incorrect facts reach the deterministic policy, the policy can faithfully apply a rule to the wrong interpretation.
That exposed the next engineering challenge:
Deterministic policy can control an agent's authority. It cannot compensate for a material fact the interpretation layer failed to recognize.
The next iteration therefore needs to strengthen semantic recognition and validation without weakening the deterministic authority boundary.
More broadly, the project reinforced another lesson:
Evaluating an agent should not only ask whether it produced a plausible answer. It should ask whether the agent acted when it could — and stopped when it should.
What I accomplished
I started this hackathon with a maintenance problem I understood and limited hands-on experience building AI agents.
I finished with:
- a live public end-to-end application
- a deployed Amazon Bedrock AgentCore runtime
- semantic agent components built with the Strands Agents SDK
- a deterministic authority and safety policy
- clarification and human-escalation logic
- a landlord-focused operational view
- an evaluation and evidence layer
- a secured public API boundary
- repeated testing against the live deployment
- a post-freeze 80-case challenge
- a reproducible evaluation trail in Git
But the accomplishment I value most is the operating idea that emerged from all of that work:
The most useful agent is not necessarily the one that does the most.
For a landlord, the better agent may be the one quietly moving routine maintenance forward in the background — while bringing the human back in exactly when their judgment, authority or attention is needed.
What's next
The immediate next step is to strengthen the semantic interpretation areas identified through the post-freeze challenge, including a stronger check that a report represents an actionable maintenance condition before routine work is allowed to progress.
Beyond that, Maintenance Autopilot could extend from governed decision-making into the wider maintenance workflow:
Tenant report → governed triage → approved work → resource selection → contractor/work-order integration → completion → landlord oversight
Once work is authorized to progress, a future resource layer could select from an approved contractor or maintenance-resource library using factors such as:
- required trade or capability
- location
- availability
- workload and capacity
- previous performance
- ratings or service history
- configured commercial rules
Those capabilities are not part of the current prototype.
The important architectural principle would remain the same:
Resource selection and downstream execution should only occur after the system has established that it has authority to progress the work.
A production implementation would also require secure identity, durable workflow state, notifications, audit history, deeper integrations and much broader real-world validation.
A decision layer, not necessarily another system
One of the more interesting possibilities is that Maintenance Autopilot does not need to become another complete property-management platform.
It could operate as a decision-control layer within systems landlords and property managers already use.
Conceptually:
Tenant / resident report
↓
Existing intake channel
↓
Maintenance Autopilot
Interpret → Apply policy → Govern authority
↓
Existing work-order / resource system
↓
Human intervention where required
Existing platforms could continue managing properties, suppliers, work orders, payments and records.
Maintenance Autopilot would answer a narrower but important question:
Given what has been reported and the authority I have been given, what am I allowed to do next?
Residential maintenance is the prototype domain.
More broadly, the architecture explores a pattern that may be useful wherever AI must interpret ambiguous human input without being given unlimited decision authority.
Why this matters
Maintenance Autopilot is currently a prototype using synthetic residential-maintenance scenarios.
It does not dispatch contractors, execute repairs, process payments or replace emergency services.
But the problem it explores is larger than any one maintenance workflow.
As AI agents become capable of doing more, the question cannot only be whether they are intelligent enough to act.
We also need to decide where their authority begins, where it ends, and when a person needs to come back in.
For maintenance, that could mean less administration for the landlord without sacrificing visibility or control.
For agentic systems more broadly, it suggests a simple principle:
Use AI for understanding. Keep authority explicit.
Maintenance Autopilot — an agent designed not just to act, but to know when not to.
Built With
- amazon-api-gateway
- amazon-bedrock
- amazon-bedrock-agentcore
- amazon-cloudwatch
- amazon-web-services
- aws-amplify
- aws-iam
- aws-lamda
- javascript
- pydantic
- python
- strands
Log in or sign up for Devpost to join the conversation.