Inspiration
Most AI assistants are good at telling people what to do, but they often stop after producing an answer. We wanted to explore what is required for an agent to continue beyond that first response: create a structured plan, process its tasks, track progress, detect failures, request approval, and produce a final outcome.
That idea inspired GoalForge — a Strands-powered prototype for turning a professional goal into a structured and observable execution workflow.
What We Built
GoalForge accepts a natural-language professional goal such as:
“Help me prepare for a software engineering interview next Thursday.”
It then:
- Interprets the user's goal.
- Breaks the goal into actionable tasks.
- Creates a structured plan with an expected outcome for every task.
- Processes tasks through a safe prototype execution layer.
- Checks whether each execution produced a usable result.
- Retries failed or unverified tasks up to two times.
- Detects potentially consequential actions and pauses for human approval.
- Tracks task status and agent activity.
- Synthesizes the task information into a final outcome for the user.
The goal of this prototype is to explore how AI systems can move from generating one response toward managing a complete goal-oriented workflow.
How We Built It
GoalForge uses the Strands Agents SDK as the foundation of its agent architecture.
The system includes:
- Strands Agent Orchestrator — connects the Qwen model with GoalForge's structured planning and tool layer.
- Structured Planner — converts a natural-language goal into an objective and a validated list of tasks.
- Prototype Execution Tool — simulates safe task-completion points where real professional integrations can be connected in future versions.
- Baseline Verification Layer — checks whether the execution layer produced a usable, non-empty result.
- Recovery Controller — retries failed or unverified tasks up to two times and stops safely if recovery is unsuccessful.
- Approval Detection — identifies potentially consequential tasks and surfaces an approval request.
- Runtime State — tracks the plan, tasks, activity events, approval status, and final result.
- FastAPI Backend — exposes the GoalForge workflow through an HTTP API.
- React Frontend — provides an interactive dashboard for goals, plans, tasks, activity, approvals, progress, and results.
- Ollama and Qwen 2.5 7B — provide the local model runtime used by the Strands agent.
The frontend is deployed on Vercel. During the hackathon demonstration, the FastAPI backend and Ollama model runtime were operated through a GitHub Codespace.
Strands Agents Implementation
GoalForge creates its core agent using strands.Agent and the Strands OllamaModel integration.
The agent has access to three registered Strands tools:
research_topic— supports prototype research operations.execute_task— processes a task through the safe execution layer.verify_task— performs baseline validation of the execution result.
The planning response follows a strict JSON structure containing:
- The overall objective
- Task title
- Task description
- Expected outcome
The response is parsed and validated before it is accepted as a GoalForge plan. This prevents unstructured model output from being passed directly into the execution workflow.
The model handles goal interpretation, plan generation, and final-outcome synthesis. Deterministic application logic manages task status, retry limits, approval detection, activity records, and workflow stopping conditions.
Agent Workflow
GoalForge follows this lifecycle:
GOAL → PLAN → PROCESS TASKS → VERIFY → RETRY OR PAUSE → SYNTHESIZE OUTCOME
For each generated task, GoalForge records:
- Title
- Description
- Expected outcome
- Current status
- Execution result
- Verification state
- Related activity events
If execution fails or produces an unusable result, GoalForge retries the task within a fixed two-attempt limit. If recovery is unsuccessful, the task is marked as failed and the workflow stops safely instead of silently claiming success.
If a task appears to involve a consequential external action—such as sending, purchasing, publishing, booking, deleting, or submitting—the system creates an approval request and pauses the workflow.
Human-in-the-Loop
GoalForge includes an approval state for potentially consequential tasks.
The dashboard can display:
- Why approval is required
- Which task triggered the request
- Approval and decline controls
- The resulting workflow activity
The current prototype demonstrates approval detection and the user-facing approval state. Durable execution of an approved external action remains future work.
Example Workflow
A user can enter:
“Create a seven-day Python preparation plan for my upcoming software engineering interview.”
GoalForge then:
- Interprets the requested outcome.
- Generates a structured objective.
- Breaks the goal into practical preparation tasks.
- Assigns an expected outcome to each task.
- Processes the tasks through the prototype execution lifecycle.
- Records task and activity states.
- Performs baseline verification.
- Retries failed or empty executions when necessary.
- Produces a final synthesized preparation outcome.
The dashboard makes this process visible instead of hiding the workflow behind a single chat response.
Challenges We Ran Into
One of our biggest challenges was making the workflow structured and observable rather than simply returning a model-generated answer.
We had to design:
- A strict planning format
- Pydantic validation for generated plans
- Task and activity states
- Bounded retry behavior
- Verification status
- Approval requests
- Final-outcome synthesis
- Communication between the React frontend and FastAPI backend
We initially experimented with Amazon Bedrock, but authorization and operation restrictions prevented us from using it reliably during the hackathon. With the deadline approaching, we adapted by using Ollama with Qwen 2.5 7B through the Strands OllamaModel integration.
We also had to connect a deployed Vercel frontend to a backend and model runtime operating in a GitHub Codespace. This required handling CORS, public access, request timeouts, and the additional latency of local model inference.
What We Learned
Building GoalForge taught us that an AI agent requires more than connecting a model to a chat interface.
A goal-oriented system needs to:
- Represent the intended outcome
- Create a structured plan
- Validate model output
- Maintain workflow state
- Record execution results
- Decide whether to continue, retry, pause, or stop
- Surface consequential decisions to the user
- Produce an outcome that reflects the workflow
We also learned that a generated answer is not the same as a verified result. Our current prototype implements baseline result checking, while task-specific outcome verification remains an important next step.
Finally, we learned how to integrate the Strands Agents SDK with a local model, FastAPI backend, and interactive React dashboard.
Accomplishments That We're Proud Of
We are proud that GoalForge demonstrates a complete agent-workflow prototype rather than only a prompt and response.
The project includes:
- Genuine Strands Agents SDK integration
- Configurable Ollama model support
- Structured JSON planning
- Pydantic validation
- Registered Strands tools
- Task and activity tracking
- Baseline verification
- Bounded retry control
- Approval detection
- Final-outcome synthesis
- FastAPI backend
- Interactive React dashboard
- Public source code and architecture documentation
Current Prototype Scope
GoalForge currently demonstrates Strands-powered goal planning, structured task generation, runtime state, task-status tracking, bounded retry control, approval requests, and final-outcome synthesis.
The current execution tool uses a safe simulated executor rather than performing external actions. Verification currently checks that execution produced a usable, non-empty result.
Runtime state is stored in memory and is not yet persisted to a database. The live backend also depends on the availability of the GitHub Codespace and local Ollama model runtime.
Real professional integrations, task-specific verification, persistent state, and durable approved-action execution are planned next steps.
Why GoalForge?
GoalForge is built around a simple idea:
AI should not stop at generating a plan. It should manage the workflow needed to move that plan toward an outcome.
By combining Strands-powered planning, a prototype execution lifecycle, baseline verification, bounded retries, approval states, progress tracking, and final-outcome synthesis, GoalForge explores the architecture required for more practical autonomous agents.
What's Next
Future improvements include:
- Connecting real professional tools and external services
- Adding task-specific deterministic verification
- Correctly recording and resuming approved actions
- Persisting goals and workflow state in a database
- Deploying the backend on stable cloud infrastructure
- Exploring Amazon Bedrock AgentCore
- Adding calendar, document, research, and communication integrations
- Creating specialized workflows for interview preparation and other professional goals
- Adding automated tests for retries, approvals, verification, and failure recovery
Built With
- codespace
- fastapi
- github
- ollama
- python
- qwen-2.5
- react
- strands-agents-sdk
- typescript
- vercel
Log in or sign up for Devpost to join the conversation.