Inspiration

The hackathon's Everyday Agents track pushed me to build something that genuinely helps someone during a hard moment, not another productivity gimmick. Before writing a single prompt, I researched the actual state of purpose and engagement in Latin America, specifically Peru.

The numbers were stark: 40% of Peruvians aged 18-24 report clinical-level mental health difficulties (Sapien Labs Global Mind Project), only 25% of Peruvian workers are engaged with their jobs (Gallup, State of the Global Workplace), and just 41% consider themselves "thriving" in life overall. At the same time, 8 out of 10 Peruvians with a mental health condition never receive treatment, with fewer than 1 psychiatrist per 100,000 people against a WHO recommendation of 5. That gap, real need, almost no professional access, is where I saw room for a self-guided tool to genuinely help, without pretending to replace therapy.

I also looked at what already exists: AI-powered Ikigai finders and habit-tracking apps like Atoms. Both attack half the problem. Ikigai tools help you discover a direction but stop there. Habit apps help you execute but often rely on streaks and guilt-based mechanics that I found actively counterproductive. Nobody connects discovery, validation, systemized execution, and honest, judgment-free follow-up into one continuous experience. That gap became TelOS.

What it does

TelOS is a team of five orchestrated AI agents, built on the Strands Agents SDK, that walks a person from "I don't know what I want" to a sustainable daily system built around what actually matters to them:

Explorer conducts an adaptive interview covering flow states, recognized strengths, current tension, and the four dimensions of Ikigai, never suggesting a purpose yet. Synthesizer cross-references the person's own words against three frameworks, Ikigai, Viktor Frankl's logotherapy, and Stanford's Designing Your Life, to draft two or three grounded purpose hypotheses, presented as structured cards with phrase, explanation, and concrete example—never a single "correct" answer. The Synthesizer is engineered to always deliver well-formed data: the orchestrator detects empty candidates and forces a retry with structured_output_model=ListaCandidatosProposito, guaranteeing structured data reaches the frontend. Each card respects strict length limits (phrase ≤15 words, explanation ≤100 chars, example ≤120 chars), which forced the synthesizer's own reasoning to stay grounded instead of drifting into generic motivational language. The constraint is architectural, not aesthetic. Validation Coach helps the person commit to one, socratically, anchoring the discussion in real past evidence they provide rather than letting them dodge the choice with hypotheticals. Systems Strategist converts that purpose into a daily operating system: positioning, leverage, a discipline standard checked against James Clear's four laws of behavior change, and an adaptation protocol for when life gets in the way, deliberately rejecting rigid goals in favor of sustainable structure. Follow-up Agent checks in whenever the person reopens the conversation, reflects their own standard back to them instead of judging it, tracks context changes as a versioned history of the system (not a progress bar toward the purpose, which would misrepresent what purpose even is), and rotates between an open, unscored question about compliance, self-perception ("does this still feel like yours?"), and whether the system itself needs adjusting. A habit streak is shown warmly in the UI when there is one, but the language never turns punitive if it breaks or if compliance was low, and the purpose itself is never treated as a "goal" that gets completed.

How I built it

I started with the interview and synthesis design before touching code: mapping which parts of the flow needed genuine model reasoning versus which were better as deterministic, code-controlled transitions, since an agentic system that invokes the model for everything is both slower and more expensive than it needs to be. Phase transitions, calendar creation, and the crisis guardrail's detection logic are fixed code; question ordering, framework synthesis, and the interpretation of what someone actually means during validation are left to the agents' judgment.

The five agents run through a deterministic orchestrator on the Strands Agents SDK, invoking Amazon Bedrock (Claude Sonnet 4.5), with Amazon Bedrock AgentCore Memory handling persistence—both the versioned "ficha" (purpose + system, never overwritten, so a person can see their own evolution) and full turn-by-turn conversational history. AgentCore Memory's event API became the data layer, eliminating the need for a custom schema or resume-logic; it handles versioning and eventual consistency out of the box.

The interface started as Streamlit, kept deliberately simple so the conversation carried the product. It later moved to a FastAPI backend plus a Next.js/TypeScript frontend for one concrete technical reason, not aesthetics: sustaining real follow-up outside the app needed genuine browser push notifications, which require registering a Service Worker, and Streamlit has no way to do that. The migration was additive and tested end to end (Cognito login, SSE streaming, the same five agents underneath) before Streamlit was retired. The app now runs on Amazon EC2 behind Amazon CloudFront for HTTPS, with Amazon Cognito handling login (no public self-signup), Amazon EventBridge driving the scheduled push-reminder job, and the whole stack, IAM included, defined as code with AWS CDK.

Calendar integration (crear_evento_calendario) is designed to run through AgentCore Gateway, wrapping Google Calendar as an MCP-compatible tool—a pattern that avoids hand-writing OAuth connectors. For this submission it returns a simulated confirmation; real calendar binding is a post-launch integration. AgentCore Runtime, the optional hosting layer that would run agents directly instead of inside EC2, is architecturally prepared but not yet deployed—the AgentCore Memory stack is independent of the runtime choice, making migration straightforward once coverage requirements justify it.

Challenges I ran into

The hardest problem wasn't technical, it was philosophical, and it had to become architectural. My systems-thinking framework explicitly rejects tracking "closeness to a goal," since a purpose isn't a finish line. That meant designing an entire follow-up experience, and a final summary view, that shows real progress and consistency without ever collapsing into a percentage or a score, because that would have quietly contradicted the philosophy the product is built on. I solved it by separating two ideas I'd initially conflated: reflecting a person's own standard back to them is legitimate; the system generating a judgment about them is not. That same line is why, when the product owner later asked to bring back a visible habit streak, it came back as a warm counter, never as a punitive one.

A close second was tone. The same product needs to sound warm and Socratic during validation and direct and executive during systems design, and a single prompt blending both registers diluted each one. Splitting that into two separate agents, rather than one agent with two "parts," fixed it. The same instinct showed up again once real users tried the flow: agents were narrating their own internal handoffs ("now I'll pass you to the next agent"), which broke the feeling of one continuous conversation even though the routing underneath was already 100% deterministic code. The fix was a shared prompt rule, not a code change, since the mechanism was already right.

Persistence introduced its own honest surprises. AgentCore Memory's writes are eventually consistent: an agent would close a phase correctly, but the orchestrator's immediate re-read to decide whether to advance could land before that write was visible yet, silently stalling the conversation with no error thrown. A bounded retry on the re-read fixed it, verified against a fake backend that deliberately returns stale data for the first couple of reads. Related, I couldn't fully trust "the model said it saved" as proof anything was saved, since text describing an action isn't the same as executing it. The fix was to stamp a real save from inside the tool's own code, and force a retry whenever generated text sounds like a closing confirmation without that stamp being set.

Infrastructure had a few of its own lessons that no amount of local testing could have surfaced. The first real deploy failed creating the App Runner service outright, the AWS account was still on Free Tier credits without a verified payment method, and App Runner has no free tier at all, which is what pushed the EC2 + CloudFront architecture described above. Later, building the scheduled push-reminder job, cdk synth produced perfectly valid CloudFormation that AWS rejected at creation time anyway, EventBridge Scheduler's documented "universal targets" don't actually accept an API destination directly, despite reading as if they do; the fix was the older, plainer EventBridge Rule + ApiDestination pattern instead.

I also had to design against a very real risk: this product handles someone's most personal reflections. That pushed me to decouple verifiable identity from the psychological profile at the data layer from day one, and to explicitly ban any scoring or "type of person" classification, not as an afterthought, but as a non-negotiable constraint on every agent's system prompt. The crisis guardrail follows the same principle in the other direction: it's a small deterministic function the orchestrator runs on every single turn before any agent sees the message, checked in both languages the app supports regardless of the selected UI language, because that's a safety property, not a preference.

Accomplishments that we're proud of

I'm proud of how deliberately I drew the line between what needed genuine model reasoning and what didn't. It would have been easy to route every step of TelOS through an LLM call, phase transitions, check-in messages, silence handling, and call it "agentic." Instead, I treated model invocation as a resource to spend only where judgment is actually required.

Phase transitions between the five agents, calendar event creation, scheduled follow-up triggers, and the check-in message itself (a template pulling the person's own stored standard, not a freshly generated sentence) are all deterministic code, not model calls. The model only runs where the task genuinely can't be reduced to a rule: adaptive interview questions, cross-referencing frameworks during synthesis, socratic validation, systems design, and the crisis guardrail's interpretation of what someone actually means. Even the "user didn't respond" case in follow-up is a state check, not an agent decision to "choose not to insist."

The result is a system that's cheaper to run and more predictable to debug than a version where the model decides everything turn by turn, without giving up any of the reasoning that actually makes the product useful. That discipline, model power reserved for what strictly requires it, structure for everything else, held up under a real cdk deploy, a real person testing the deployed app end to end, and real AWS spend protections, not just in the pitch. It's the architectural decision I'd point to first if asked what makes TelOS a genuine agentic system rather than a chatbot wearing an orchestration diagram.

What I learned

That "systems over goals" isn't just a productivity slogan, it's a real architectural constraint once you take it seriously enough to build with it. That the most important design decisions in an agentic product are often about what the system must never do (never judge, never score, never nag), not just what it should do. And, from the infrastructure side, that a real cdk deploy against a live account will always know things cdk synth can't tell you, App Runner's Marketplace subscription requirement and EventBridge Scheduler's target restrictions were both semantic failures no amount of local template validation would have caught, so I stopped treating a clean synth as proof anything would actually deploy.

What's next

Real-time crisis-resource lookup against verified, current sources rather than a static list; wiring the calendar tool to a real Google Calendar connector through AgentCore Gateway instead of the current mocked confirmation; a scheduler-driven follow-up cadence instead of the simplified session-triggered version I shipped for the hackathon; and full user-facing delete controls, export already works today, but deleting your own data is still a developer-only script, not a button in the product yet.

Built With

Share this project:

Updates

posted an update —

Arrancó Telos

Definí la arquitectura completa para el hackathon AWS Agents for Humans (track Everyday Agents): un agente conversacional que acompaña a una persona a través de 4 fases fijas (explorar, sintetizar, validar, convertir en sistema) hasta un propósito de vida concreto y un sistema de 4 preguntas accionable. Un quinto agente hace seguimiento breve, sin rachas ni gamificación.

Log in or sign up for Devpost to join the conversation.

Submission history