Inspiration

Libraries already have software for the predictable work.

A room-booking system can detect a conflict. An interlibrary-loan system can route a routine request. An overdue system can send a reminder.

But when the case becomes ambiguous, the automation usually stops.

A room conflict may have no obvious winner. An interlibrary-loan request may match multiple editions. An overdue case may require staff to consider the patron's circumstances rather than send another generic reminder.

That exception queue is where automation stops and staff judgment begins.

And that is exactly where library staff have the least capacity to absorb more work.

Stacks was built for that gap.

The central question was:

What if an agent could take a library exception from detection to resolution, while knowing exactly when it must stop and hand the decision back to a person?

That became the foundation of Stacks.

During this hackathon period, I was also a mentor for AWS User Group Madurai's Agents for Humans Builder Circle, helping other builders work with the Strands Agents SDK, Amazon Bedrock AgentCore, and Kiro.

That experience reinforced an important distinction.

It is relatively easy to demonstrate an agent that can reason and call tools.

It is much harder to build one that can safely act inside an operational workflow.

So I deliberately built Stacks around that harder problem.

Not "How autonomous can the model be?"

"How much useful work can the agent safely complete before a human needs to take over?"


What Stacks does

Stacks is a task-completion agent for library operations, focused on exception-heavy workflows that traditional automation can identify but cannot completely resolve.

It currently handles three workflows.

Room-booking conflicts

Stacks examines conflicting bookings, retrieves the applicable library priority policy, evaluates the situation, and resolves the conflict when policy permits.

Ambiguous interlibrary-loan requests

When a request matches multiple possible editions, Stacks invokes an ILL Disambiguation Specialist to narrow the candidates using the request and catalog context instead of leaving staff to reconstruct the ambiguity manually.

Overdue-item escalation

Stacks runs a durable, multi-day escalation sequence that maintains relevant context rather than repeatedly sending the same generic reminder.

The important distinction is that Stacks does not stop at generating a recommendation.

It moves a case through:

Case → Context → Policy → Decision → Action → Outcome → Audit

That is why Stacks is designed as a task-completion agent, not a library chatbot.


The problem

Existing library automation is very good at predictable work.

It can block an obvious double-booking, automatically route a routine loan, or send a templated overdue notification.

The difficult cases are different.

They require someone to:

  • understand the context
  • interpret the library's policy
  • resolve ambiguity
  • choose between competing options
  • decide whether an action is safe
  • sometimes decide that the system should not act at all

Those cases become a human review queue.

Stacks is designed specifically for that queue.

The goal is not to replace the systems libraries already depend on.

The goal is to work on the remainder they hand back to people.


Why this is an agent, not a chatbot

A chatbot primarily produces an answer.

Stacks receives an operational case and can:

inspect → reason → retrieve context → call tools → apply policy → take action → request approval → record the outcome

A chatbot:

Question → Answer

Stacks:

Case → Decision → Action → Outcome

That distinction matters.

The purpose of Stacks is not to have a conversation about library operations.

Its purpose is to complete library work.

The agent can act when policy permits it.

The agent can prepare an action when staff confirmation is required.

And the agent can stop when a human must make the decision.


Human judgment is part of the architecture

The most important design decision in Stacks is that the model does not decide whether it is safe to act.

Every action is classified into one of three safety tiers.

GREEN

The action is safe to execute automatically.

YELLOW

Stacks prepares the action, but a staff member must confirm it.

RED

Stacks does not make the decision. An authorized human must review it.

The critical rule is:

The model does not decide its own safety tier.

The GREEN, YELLOW, and RED classification is determined by application code and policy conditions.

A confident model response cannot simply persuade the system that an action is safe.

Every meaningful mutation is also recorded in the audit trail, including the case, action, safety tier, result, and timestamp.

This creates a simple operating contract:

Automate execution where policy allows it. Escalate judgment where policy requires it.


What makes Stacks different

Traditional automation generally follows:

Detect → Rule → Route → Human queue

Stacks follows:

Case → Context → Policy → Specialist reasoning → Safety gate → Resolve or escalate

That difference is the core of the project.

Stacks is not intended to be a replacement for an ILS, LibCal, or interlibrary-loan platform.

It is the intelligence layer for the cases those systems cannot confidently finish.

The agent works on ambiguity instead of only executing predetermined rules.

The safety layer prevents that flexibility from becoming uncontrolled autonomy.

And the audit trail makes the resulting action explainable and reviewable.


Memory with a purpose

Stacks uses Amazon Bedrock AgentCore Memory for operational context across sessions.

Memory is not being used simply to make the conversation feel personalized.

It exists because historical context can change an operational decision.

For example, Stacks can recall relevant prior information such as:

  • a requester's historical substitution pattern
  • a documented patron hardship flag

That means a decision made today can incorporate relevant context from previous interactions rather than starting from a blank state every time.

The design principle is:

Memory should improve an operational decision, not simply remember a conversation.


A deliberately focused agent architecture

Stacks does not use a large multi-agent swarm simply to make the architecture look more sophisticated.

The system uses a single primary Strands Agent for orchestration, with narrowly scoped capabilities where specialized reasoning provides real value.

Stacks Agent

The primary agent coordinates the overall workflow, retrieves context, selects tools, applies policy, and determines the next operational step.

ILL Disambiguation Specialist

A dedicated specialist handles ambiguous edition matching.

It is implemented as a second Agent instance invoked through a tool boundary and focused specifically on the ILL disambiguation problem.

Overdue Escalation Sequencer

The overdue workflow uses durable session management to maintain a multi-day sequence rather than relying on an in-memory loop.

The result is a compact architecture where each component has a clear responsibility.


How we built it

Stacks is deployed as a real AWS-backed application rather than a notebook or isolated model demo.

Agent layer

  • Strands Agents SDK
  • Amazon Bedrock AgentCore Runtime
  • Amazon Nova Lite
  • Amazon Bedrock Guardrails
  • Amazon Bedrock AgentCore Memory

Identity and application layer

  • Amazon Cognito
  • AWS Amplify Hosting
  • Next.js
  • FastAPI
  • AWS Lambda

Data and durability

  • Amazon DynamoDB
  • Amazon S3
  • Amazon EventBridge
  • AWS Lambda

Infrastructure

  • Terraform
  • IAM
  • Bedrock Guardrail configuration
  • Reproducible AWS infrastructure

The staff application communicates with a FastAPI backend.

The backend verifies identity and authorization before invoking the AgentCore Runtime.

The agent then operates against the library's operational tools and data.

The result is returned to the staff application, while meaningful actions are recorded in the audit trail.


Safety is layered

Stacks uses multiple independent controls because no single safety mechanism should be responsible for the entire decision boundary.

Identity boundary

Amazon Cognito establishes who is making the request.

Authorization boundary

The backend enforces tenant and role permissions server-side.

Model safety boundary

Amazon Bedrock Guardrails provide independent content safety controls.

Action safety boundary

Application code determines whether an action is GREEN, YELLOW, or RED.

Human boundary

Sensitive actions stop for staff approval.

Audit boundary

Meaningful mutations are recorded for later inspection.

The model is therefore one component of the decision process, not the authority over the entire system.


Challenges we ran into

The hardest part was not getting the model to call the correct tool.

It was making sure the system could not cross boundaries that the agent itself was not supposed to control.

During development, adversarial testing exposed two significant issues before deployment.

Cross-tenant data isolation

An interlibrary-loan catalog search path could expose information across library boundaries.

The issue was traced and fixed by enforcing tenant isolation at the application layer rather than assuming that the model would behave correctly.

Human approval was not actually pausing execution

The HITL mechanism appeared connected at the application level, but execution was not actually stopping where approval was required.

Tracing the Strands SDK behavior directly exposed the problem.

The approval boundary was redesigned so the code path itself enforced the pause.

These failures became some of the most valuable lessons from the build.

A safety mechanism is only real when you can demonstrate that the agent cannot bypass it.


Building a public deployment

Getting the agent working locally was only part of the challenge.

Making the entire application work as a public system introduced another class of engineering problems.

The deployment required:

  • real Cognito authentication
  • tenant-scoped data access
  • server-side authorization
  • AgentCore Runtime invocation
  • AgentCore Memory integration
  • durable DynamoDB audit records
  • scheduled EventBridge processing
  • Lambda Function URL configuration
  • browser-compatible CORS behavior
  • reproducible Terraform infrastructure

The final system is deployed as a real public application with a live frontend and real AWS-backed agent execution.

The demo uses synthetic library operational data so that the system can be demonstrated without exposing real patron information.


An architectural adjustment we made

The original development environment was intended to use a Claude model through Amazon Bedrock.

During deployment, the required model-use-case agreement was not available on the AWS account being used for the project.

Instead of treating that as a blocker, the architecture was adapted to use Amazon Nova Lite.

This was a useful engineering lesson in itself.

A production-shaped architecture has to adapt to real account constraints, service availability, permissions, and deployment requirements.

The final deployed system therefore reflects what is actually running rather than what was originally planned.


Accomplishments we are proud of

Stacks is a production-shaped system rather than a notebook prototype.

It has:

  • a real staff-facing web application
  • real Cognito authentication
  • role-scoped access
  • tenant-aware DynamoDB data
  • a deployed AgentCore Runtime
  • AgentCore Memory
  • Bedrock Guardrails
  • specialized agent reasoning
  • scheduled workflows
  • human approval
  • durable auditability
  • reproducible infrastructure
  • a public live demo

More importantly, the safety design was not only documented.

It was adversarially tested.

The testing process found real weaknesses in the implementation and those weaknesses were fixed before the system was treated as ready for demonstration.


Why this fits Good Neighbor Agents

The value of Stacks is not limited to the staff member using the application.

It flows outward to the community.

When Stacks resolves a room conflict, a community group can use the space.

When it helps disambiguate an interlibrary-loan request, a patron can receive the material they actually need.

When it reduces repetitive overdue work, library staff can spend more attention on cases where human judgment matters.

The beneficiary is therefore not simply the person operating Stacks.

It is the community served by the library.

That is why the project belongs in the Good Neighbor Agents track.

The project is designed around a simple idea:

Give community-serving organizations more capacity by letting agents handle the exception work that consumes human attention, without taking away human authority.


Potential impact

The three workflows implemented in Stacks are only the starting point.

The same pattern could extend to other library exception queues, including:

  • policy interpretation
  • membership exceptions
  • inter-branch coordination
  • special collections workflows
  • event conflicts
  • service escalations

The architectural pattern remains the same:

Let automation own predictable work.

Let agents work through ambiguity.

Let humans retain authority over consequential decisions.

For libraries, that means converting staff time spent moving cases through queues into staff time spent serving people.


What we learned

The hardest part of agentic automation is not making the agent act.

It is deciding exactly where the agent is allowed to act on its own.

Building Stacks reinforced five principles:

  1. Agents need boundaries, not just capabilities.
  2. Safety decisions should be enforced in code, not delegated to model confidence.
  3. Memory should exist because context changes decisions.
  4. Human approval should be an actual execution boundary, not a UI decoration.
  5. A useful agent should complete work and leave an auditable outcome.

The most important lesson was simple:

The best operational agent is not the one that acts the most. It is the one that knows when to act, when to ask, and when to stop.


What's next

The next stage would be a real library pilot using institution-specific policies and live integrations.

Planned directions include:

  • OCLC/WorldShare integration
  • LibCal integration
  • institution-specific policy configuration
  • expanded operational workflows
  • first-class approval events
  • structured audit querying
  • pilot validation with practicing librarians

The current implementation intentionally uses synthetic data and does not claim live production integration with these external library systems.

The objective is to demonstrate a credible architecture that can become a real operational system without overstating what the hackathon prototype already does.


Built for Agents for Humans

Stacks was built for the Agents for Humans Hackathon under the Good Neighbor Agents track.

The project uses the Strands Agents SDK, Amazon Bedrock AgentCore, Amazon Nova, and supporting AWS services.

The project was deliberately designed around the dimensions of the hackathon judging criteria:

Technological Implementation

A real Strands-based agent deployed on Amazon Bedrock AgentCore with tools, memory, guardrails, scheduled execution, and human oversight.

Design

A complete staff-facing application designed around operational workflows rather than a conversational AI interface.

Potential Impact

A concrete problem affecting libraries and the communities they serve.

Creativity & Originality

An agent focused on the exception queue left behind by traditional library automation rather than another generic library assistant.

Presentation

A complete journey from an operational case through context, policy, reasoning, action, human approval where necessary, and audit.


Built through AWS User Group Madurai

Stacks is an individual hackathon submission.

During the Agents for Humans preparation period, I served as a mentor for AWS User Group Madurai's Agents for Humans Builder Circle, helping other builders explore the Strands Agents SDK, Amazon Bedrock AgentCore, and Kiro.

That experience directly influenced the engineering direction of Stacks, particularly its focus on:

  • agent memory
  • identity
  • human oversight
  • safe tool execution
  • production-shaped deployment

The community experience helped turn the project from a simple agent demonstration into an exploration of what it takes to let an agent operate responsibly on behalf of people.


Try Stacks

Live Demo

https://main.d1f4dnxaaugevn.amplifyapp.com/

GitHub

https://github.com/yogeshselvarajan/stacks

Judge quick start

  1. Open the live demo.
  2. Click Sign In.
  3. Select Continue as Hackathon Judge.
  4. Open Approval Inbox.
  5. Review a pending case and the policy clause supporting the decision.
  6. Approve the action.
  7. Open Audit Trail and inspect the recorded outcome.
  8. Try the ILL Queue to experience ambiguous request routing.
  9. Explore the Overdue Queue and Calendar workflows.

The hosted application is the primary way to experience Stacks.

No AWS CLI, curl command, or local development environment is required for the judge experience.

The deployed infrastructure is real, while the library records used in the demonstration are synthetic.


Final thought

Libraries do not need another system that tells staff what they already know.

They need more capacity for the cases that require attention.

Stacks is built for that space between automation and judgment.

Your systems handle the routine. Stacks works what they leave behind.

Built With

  • ai-agents
  • amazon-bedrock
  • amazon-bedrock-agentcore
  • amazon-nova
  • bedrock-guardrails
  • community-technology
  • fastapi
  • generative-ai
  • human-in-the-loop
  • library-technology
  • multi-agent-systems
  • nextjs
  • python
  • strands-agents-sdk
  • terraform
  • typescript
Share this project:

Updates

Submission history