Inspiration

Construction projects generate a huge amount of administrative work: checking insurance and compliance records, tracking submissions, following up with contractors, comparing drawings and schedules, and preparing RFIs when documents disagree. Much of this work is repetitive, but the consequences of missing something can be significant.

My background in quantity surveying and construction project administration made this problem especially familiar. I wanted to build an agent that could take responsibility for the routine work without pretending that every construction decision should be automated.

That became BuildGuard: an autonomous construction project administration agent that handles low-risk administrative work and escalates decisions that require professional judgment.

What it does

BuildGuard works in the background over project documents, requirements, compliance records, submissions, issues, and actions.

Instead of behaving like a chatbot, it follows an administrative workflow:

inspect project facts → reason → use tools → update project state → verify → automate safely or escalate

The prototype demonstrates three core workflows:

  • Compliance monitoring: BuildGuard detects expiring compliance items and overdue submissions, creates evidence-backed issues, and prepares appropriate follow-up actions.
  • Routine follow-up: low-risk reminders can be executed only when a deterministic policy allows them.
  • RFI / document discrepancy escalation: BuildGuard compares project documents, identifies inconsistencies, gathers evidence, prepares an RFI-related action, and requests human approval rather than deciding the issue itself.

For example, in the synthetic Riverside Office Development dataset, BuildGuard independently identified that Door D07 is shown as 1200 mm on Architectural Drawing A-102 but 1000 mm on Door Schedule DS-01. It created a discrepancy issue, prepared an approval-gated RFI draft, and requested human review. It did not decide which dimension was correct.

A second live scan detected a Contractor's All Risks insurance policy approaching expiry and an overdue aluminium-window material submission. It created two issues and two proposed reminder actions while correctly leaving unrelated compliant records alone.

How I built it

The agent is written in TypeScript using the Strands Agents SDK. BuildGuard currently exposes 13 narrowly scoped tools for reading project information, creating project state, requesting approvals, and safely executing permitted actions.

The safety model is deliberately separated from the language model. A deterministic execution policy controls whether an external action is allowed. At present, send_reminder is the only automatically executable action type. RFI-related actions, contractual decisions, approvals, and other consequential actions remain outside the agent's authority.

The backend includes:

  • Strands Agents SDK for agent orchestration
  • Amazon Bedrock as the default model-provider path
  • Amazon Bedrock AgentCore Runtime for the deployed runtime
  • Amazon SES for explicitly enabled notification sending
  • Amazon CloudWatch for runtime observability
  • A synthetic construction-project dataset with drawings, requirements, compliance records, and submissions
  • A local-outbox mode so development and tests cannot accidentally send email

The frontend is a Next.js 16 / React 19 dashboard showing project health, attention items, agent activity, and a human approval screen for RFI decisions. The current dashboard uses clearly labelled synthetic preview data and is intentionally not presented as a live connection to the agent.

BuildGuard also has a configurable model-provider layer. Bedrock remains the default, while Google Gemini was added as a fallback for local live demonstrations after an AWS account-level inference quota prevented successful Bedrock inference near the submission deadline.

The project currently has 129 automated tests across 20 suites, covering tool behavior, safety policy, approvals, model-provider configuration, runtime behavior, output filtering, dataset integrity, and failure handling.

Challenges I ran into

One of the biggest challenges was making autonomy safe rather than simply impressive. A construction agent should not be able to approve its own RFI, issue contractual correspondence, or claim an email was delivered merely because a provider accepted it. I therefore designed explicit authority boundaries and kept consequential decisions outside the agent's tool surface.

Deployment also presented an unexpected challenge. The default AgentCore Node CodeZip packaging flow did not include the project's non-code synthetic dataset. I created a custom build artifact that bundles the runtime together with the required project data and deployed that artifact successfully. The AgentCore Runtime is now deployed and reports READY.

A separate AWS account-level inference quota reduced the available Bedrock model capacity to zero. Rather than fabricate a successful run, I kept that limitation visible, opened an AWS Support case, and added a provider abstraction so the same Strands agent, tools, prompts, and safety policies could run against another model for the live demonstration.

This also exposed other engineering details worth addressing, including provider rate limits, secret redaction, and reasoning metadata. BuildGuard filters model reasoning and provider thought signatures from all product and CLI output surfaces.

Accomplishments that I'm proud of

I’m especially proud that BuildGuard became more than a proof-of-concept chatbot. It is a working agent system with clear safety boundaries, real tool use, deployment infrastructure, and verified workflows.

Some of the accomplishments I’m most proud of are:

  • Building a 13-tool Strands agent that can inspect project facts, create issues and actions, request approvals, and execute only explicitly permitted low-risk work.
  • Designing the approval flow so that the agent cannot approve its own consequential actions. Human approval is outside the Strands tool surface by design.
  • Successfully running two live autonomous workflows:
    • a compliance scan that made 12 tool calls, detected an expiring CAR insurance policy and an overdue aluminium-window submission, and created two evidence-backed issues with two proposed reminder actions;
    • an RFI scan that made 12 tool calls, independently identified the 1200 mm vs 1000 mm Door D07 discrepancy, created a document-discrepancy issue, prepared an approval-gated RFI draft, and requested human review.
  • Deploying BuildGuard to Amazon Bedrock AgentCore Runtime, including solving a packaging issue where non-code project data was omitted from the default Node CodeZip flow.
  • Building a deterministic execution-policy layer so safety does not depend solely on model instructions.
  • Adding Amazon SES with explicit opt-in sending and a local-outbox mode for safe development.
  • Building a professional Next.js dashboard and approval-review interface that makes the agent’s work understandable to a construction professional.
  • Reaching 129 passing tests across 20 suites, covering tools, approvals, execution policy, provider selection, runtime behavior, output filtering, and failure handling.
  • Handling provider failures honestly. When AWS inference quotas blocked live Bedrock runs, I added a configurable model-provider layer without changing the core Strands agent, tools, or safety policies, rather than fabricating successful output.

What I learned

Building BuildGuard reinforced that useful professional agents need more than an LLM and a set of tools. They need authority boundaries, deterministic policies, evidence, verification, observability, and graceful failure handling.

I also learned that human-in-the-loop design works best when it is structural rather than just an instruction in a system prompt. In BuildGuard, the agent literally does not have the tool required to approve its own consequential actions.

Most importantly, I learned that autonomy does not have to mean removing people from the workflow. For professional work, the more useful goal is to automate the repetitive work surrounding expert judgment so that professionals can spend their time on the decisions that actually require them.

BuildGuard handles the administrative work around construction projects so professionals can focus on the decisions that require professional judgment.

What's next for BuildGuard

The next step for BuildGuard is to move from a strong hackathon prototype into a persistent, production-ready construction administration platform.

The immediate roadmap includes:

  • Persistent project state so issues, actions, approvals, and workflow history survive between invocations.
  • A live dashboard API so the frontend displays real agent state instead of preview data.
  • Scheduled autonomous scans for compliance expiries, overdue submissions, and document changes.
  • Project document ingestion, allowing teams to upload drawings, schedules, specifications, certificates, and registers rather than relying on a prebuilt dataset.
  • Construction-platform integrations so BuildGuard can work with the systems teams already use instead of creating another isolated workspace.
  • A more complete audit history showing exactly what the agent read, changed, escalated, and verified over time.
  • Human-authorised external actions, such as transmitting an approved RFI only after the appropriate reviewer has explicitly authorised it.
  • Broader construction workflows including variations, payment applications, procurement follow-up, progress reporting, and additional compliance categories.

Longer term, I see BuildGuard becoming an always-on project-administration layer that continuously watches the project, handles repetitive coordination work in the background, and brings people in only when professional judgment, contractual responsibility, or commercial decision-making is required.

Built With

Share this project:

Updates

Submission history