Berry

One task tracker for people and AI agents, where a person always makes the final call.

Berry is a self-hosted workspace where people and AI agents share the same projects, goals, tasks, discussions, and reviews.

A task can be assigned to a person or an agent. When an agent receives work, it can inspect the project, write code, run checks, collaborate with other agents, and open a pull request. The result then waits for a person to review.

Berry is built around one principle:

Agents work for people, not around them.

Agents may plan, implement, review, delegate, and suggest work, but they cannot release their own work. Organization agents cannot merge code, mark work Done, or cancel tasks. Risky operations require human approval, and delivered work passes through a human review gate.

Berry itself acts as the control plane. Agent execution happens in a separate isolated runtime, normally Amazon Bedrock AgentCore Runtime. The Berry server does not call AI models directly.


Core Concepts

Task

The main unit of work.

A task contains a title, description, status, priority, assignee, project, due date, labels, and optional custom fields.

The codebase may call these issues; the product calls them tasks.

Agent

An AI worker in the workspace.

Agents have:

  • a name
  • instructions
  • a model
  • skills
  • permissions
  • optional MCP tools
  • an optional role in Berry's built-in organization

Agents can be assigned, mentioned, and chatted with much like human teammates.

Run

One attempt by an agent to perform a task.

A run is queued, starts, reports progress, and finishes, fails, or is cancelled. Its messages, tool calls, usage, and results are recorded.

Runtime

The isolated environment where agents work.

Berry's standard runtime is Amazon Bedrock AgentCore Runtime, although the same runtime image may also be hosted over ordinary HTTP.

Autonomy Level

Berry's built-in organization gives each role an autonomy level from 1 to 5.

The level limits which tools the role may use.

No autonomy level allows an organization agent to:

  • merge code
  • mark a task Done
  • cancel a task

Approval

A human decision required before a sensitive action may proceed.

Examples include risky planned work, escalations, and proposed work with material impact.

Review Gate

When an agent finishes work, the task moves to In review.

Agent reviewers may inspect the work first, but the final decision belongs to a person.

Proposal

Structured work discovered by an agent.

A proposal includes evidence, impact, severity, suggested action, responsible role, and required reviewers.


A Typical Day

Maya wants customers to export invoices as CSV.

She opens Plan something…, chooses the Invoices project, and enters:

Let customers export their invoices as CSV.

Berry analyzes the request.

If something important is ambiguous, Berry asks before creating work:

Should export include every invoice or only the current filtered view?

Maya answers.

Berry generates a plan containing milestones, tasks, dependencies, capabilities, and risk information.

For example:

Goal: Export Endpoint

  • Implement CSV endpoint
  • Validate permissions
  • Add tests

Goal: Export UI

  • Add export button
  • Send current filters
  • Handle download errors

Maya starts the plan.

The Orchestrator assigns each task to the best-fitting agent.

The Backend Engineer starts the API work. During implementation, it discovers a schema question and creates a linked sub-task for the Database Engineer.

When coding finishes, the runtime pushes a branch and Berry opens a pull request containing:

  • task reference
  • implementation summary
  • checks performed
  • check results

The task moves to In review.

If AutoGate or role-based reviews are enabled, other agents may review first.

Their approval never closes the task.

Maya opens Reviews, sees the pull request, diff, checks, and agent verdicts, then chooses:

  • Approve → task becomes Done
  • Send back → task returns to Todo with her feedback

Later, the Security Engineer's weekly discovery finds a missing authorization check and submits a proposal with evidence.

Because it has security impact, Berry waits for a person before starting the work.


1. Work Tracking

Berry is a complete task tracker regardless of whether work belongs to a human or an agent.

Tasks

Tasks support:

  • title and description
  • status
  • priority
  • human or agent assignee
  • project
  • due date
  • labels
  • custom fields
  • attachments
  • comments
  • reactions
  • followers
  • dependencies
  • sub-tasks
  • activity history

Tasks can be created manually, through planning, by agents, or through the API.

Statuses

Tasks move through:

  • Backlog
  • Todo
  • In progress
  • In review
  • Done
  • Blocked
  • Cancelled

Agents may move tasks through operational states such as Todo, In progress, In review, and Blocked.

Only people may mark work Done or Cancelled.

Assigning Work

Human and agent assignees appear in the same picker.

For agents, Berry supports:

  • Assign without starting
  • Assign and start

Assigning a Todo task to an available agent can start a run automatically.

Sub-Tasks and Delegation

Tasks may have nested sub-tasks.

Agents may delegate work by creating linked sub-tasks for approved roles.

Examples:

  • Backend Engineer → Database Engineer
  • Engineering Manager → Backend Engineer
  • Engineering Manager → Frontend Engineer

Dependencies

Tasks may depend on other tasks.

Berry rejects circular dependencies.

Tasks created by a plan remain Blocked until their prerequisites are complete.

Comments and Mentions

Every task has a threaded conversation.

Mentioning an agent can start a run for that agent on the task.

A note mode allows discussion without waking agents.

Activity History

Berry records changes to:

  • status
  • priority
  • assignee
  • fields
  • labels
  • parent/child relationships
  • comments
  • reactions

Each event identifies whether a person or agent performed it.

Views

Tasks can be displayed as:

  • List
  • Board
  • Table
  • Swimlanes
  • Gantt

Views support filtering, sorting, grouping, and configurable fields.

Filters

Tasks can be filtered by:

  • status
  • priority
  • assignee
  • creator
  • project
  • label
  • date
  • custom fields
  • personal scope

Saved Views

Users can save task layouts and filters as personal or shared views.

Inbox

The Inbox collects activity requiring attention, including:

  • mentions
  • comments
  • assignments
  • review requests
  • completed runs
  • failed runs
  • blocked agents
  • approvals
  • proposals

It is the main way a person follows agent activity without watching every run.


2. Planning and Automation

AI Planning

Berry can turn a natural-language objective into structured work.

A plan may contain:

  • milestones
  • tasks
  • dependencies
  • required capabilities
  • approvals
  • assumptions
  • blocking questions
  • risk assessment

Nothing enters the real task board until a person starts the plan.

Clarifying Questions

Berry should ask instead of inventing important requirements.

Blocking questions stop task generation until answered.

Once answered, they become settled context for replanning.

Plan Validation

Before a plan starts, deterministic checks verify that:

  • task IDs are unique
  • dependencies reference real tasks
  • dependency graphs contain no loops
  • approvals point to real tasks
  • milestones contain tasks

Risk is assessed separately from the model.

Goals

Each plan milestone becomes a Goal.

A goal tracks:

  • related tasks
  • progress
  • pending approvals
  • blocked work
  • status

Goal states include:

  • Draft
  • Planned
  • Active
  • Blocked
  • Completed
  • Cancelled

Projects

Projects group related work.

A project may contain:

  • tasks
  • goals
  • activity
  • status
  • health
  • target date
  • linked repository

Agents working on code use the project's repository.

Automatic Routing

After a plan starts, the Orchestrator evaluates:

  • tasks
  • available agents
  • roles
  • capabilities

It assigns work to appropriate agents and may select a predefined workflow.

If routing fails, tasks remain available for manual assignment.

AutoGate

AutoGate adds agent review before human review.

Agent implements
      ↓
Agent reviewer checks
      ↓
Changes may be requested
      ↓
Required agent reviews pass
      ↓
Human decides

Agent approval never marks the task Done.

Autopilots

Autopilots are recurring instructions assigned to an agent.

Triggers may include:

  • schedules
  • signed webhooks
  • manual Run now

Examples:

  • daily standup
  • stale-task sweep
  • release notes
  • CI failure monitoring
  • security discovery
  • weekly digest

Each firing becomes an ordinary agent run or task.

Quick Actions

Quick actions are reusable workspace prompts.

Examples:

  • Summarize this task
  • Review this discussion
  • Draft release notes
  • Explain this failure

They run through a chosen agent.


3. Agents and Execution

Agent Roster

The Agents page shows:

  • availability
  • workload
  • runtime health
  • model
  • owner
  • access
  • recent activity
  • success/failure statistics
  • usage

Creating Agents

Agents may be:

  • configured manually
  • drafted through an AI-assisted builder
  • created automatically as part of Berry's organization

Agent Capabilities

An agent may have:

  • instructions
  • conversation starters
  • skills
  • MCP servers
  • model configuration
  • repository permissions
  • runtime binding
  • environment variables

Skills

Skills are reusable knowledge and instruction packages.

They may be:

  • written inside Berry
  • imported from GitHub
  • imported as archives
  • assigned to multiple agents

Imported skills can be refreshed from their original source.

MCP Servers

Agents may access external MCP servers.

MCP servers may belong to:

  • one agent
  • the whole workspace

They may connect directly or through AgentCore Gateway.

Runs

A run includes:

  • task context
  • conversation context
  • agent instructions
  • skills
  • tools
  • temporary credentials

During a run, the agent may:

  • inspect project files
  • write files
  • run commands
  • comment
  • create sub-tasks
  • escalate
  • open pull requests
  • propose work

Every step is recorded.

Chat

Users may chat directly with agents.

Agent conversations use the same instructions, model, and runtime as task work.

Tasks and projects can be brought into chat as context.

Usage and Cost

Berry records model usage per run.

Usage can be summarized by:

  • agent
  • model
  • day
  • week
  • project
  • workspace

Where pricing is known, Berry calculates cost.


4. Berry's Built-In Agent Organization

Every workspace receives a predefined AI organization.

There are 19 agents across 7 departments.

Department Role Level
Operations Orchestrator 2
Product Product Lead 5
Product Business Analyst 2
Product UX Researcher 2
Product Product Designer 2
Engineering Software Architect 5
Engineering Engineering Manager 2
Engineering Backend Engineer 4
Engineering Frontend Engineer 4
Engineering Database Engineer 3
Engineering Integration Engineer 3
Quality & Security QA Engineer 5
Quality & Security Security Engineer 5
Platform DevOps Engineer 3
Platform Site Reliability Engineer 4
Growth & Insight Data & Analytics Engineer 3
Growth & Insight Technical Writer 3
Growth & Insight Growth Engineer 3
Leadership CTO 5

Each role defines:

  • mission
  • responsibilities
  • expected inputs
  • expected outputs
  • autonomy level
  • allowed tools
  • delegation relationships
  • escalation rules
  • reviewers
  • discovery responsibilities

Autonomy Levels

Level 1 — Advisory

May:

  • read workspace context
  • inspect files
  • comment
  • escalate

Level 2 — Contributor

Adds:

  • create tasks
  • create projects
  • write files
  • attach files
  • propose work
  • delegate
  • change permitted task states

Level 3 — Executor

Adds:

  • run commands
  • modify code
  • collect artifacts
  • prepare pull requests

Level 4 — Autonomous

Uses the same core execution tools as Level 3 but may have routine, low-risk proposals accepted automatically.

Level 5 — Authority

Adds the ability to submit review verdicts within the role's domain.

Restrictions at Every Level

No level can:

  • merge code
  • mark Done
  • cancel work

Required Reviews

Code-writing roles normally receive:

QA Engineer

  • blocking

Security Engineer

  • blocking for security-sensitive changes

Software Architect

  • blocking for architecture-sensitive changes

Additional advisory reviews may come from:

  • Database Engineer
  • Product Designer

A reviewer cannot approve its own work.

Delegation

Roles may hand work to approved roles.

Delegation creates linked sub-tasks with acceptance criteria.

Escalation

Agents can escalate decisions they do not own.

Typical categories:

  • product
  • technical
  • security
  • operational

Escalations may become:

  • tasks for senior agents
  • human approval requests

The original work may remain Blocked until answered.

Work Discovery

Professional agents may periodically inspect the workspace for problems or opportunities.

Examples:

QA Engineer

  • missing tests
  • flaky tests
  • skipped tests

Security Engineer

  • vulnerable dependencies
  • exposed secrets
  • missing authorization

Technical Writer

  • outdated documentation
  • undocumented behaviour

Engineering Manager

  • blocked work
  • ownerless work
  • conflicting tasks

Discoveries become proposals backed by evidence.

Agents are instructed not to implement discovered work during the discovery run.

Proposal Policy

Agents may propose work with:

  • evidence
  • severity
  • impact
  • suggested action
  • effort
  • dependencies
  • owner
  • reviewers

Routine proposals may start automatically only when:

  • proposer is Level 4 or 5
  • severity is low or medium
  • impact is routine

Everything else waits for a person.


5. Human Control

Human control is structural, not just prompt-based.

Review Gate

Agent delivers work
        ↓
    In review
        ↓
 Human decision
     ↙       ↘
Approve    Send back
   ↓           ↓
 Done        Todo
           + feedback

Approvals

Approvals happen before sensitive work.

They may gate:

  • risky planned tasks
  • work proposals
  • escalations

Approval decisions are recorded in workspace history.

Agent Reviewers

Agent reviews may:

  • approve
  • request changes
  • provide advisory findings

They cannot release the work.

GitHub Delivery

For code tasks, the runtime may:

  1. create a branch
  2. commit changes
  3. run project checks
  4. push the branch
  5. open a pull request

The pull request contains:

  • task key
  • summary
  • verification results

Berry does not merge the pull request.

A person merges it through GitHub.


6. GitHub

GitHub is Berry's primary repository integration.

People sign in using GitHub.

Their encrypted GitHub authorization allows Berry to:

  • list repositories
  • link repositories to projects
  • clone repositories for runs
  • push branches
  • open pull requests

Agents never receive unrestricted GitHub credentials.

Repository actions are individually checked against the agent's permissions.

Projects link to existing repositories; Berry does not create repositories.


7. Workspaces and Administration

Berry supports multiple workspaces.

Each workspace contains its own:

  • members
  • agents
  • organization
  • projects
  • goals
  • tasks
  • runtimes
  • integrations
  • plugins
  • skills
  • settings

Roles

Owner

Full control.

Admin

Manages workspace configuration and normal members.

Member

Creates and updates work and can start agents.

Viewer

Read-only.

Permissions are checked on every mutation.

Invitations

Workspaces support:

  • email-targeted invitation links
  • reusable join links

Invitation secrets are shown once and stored hashed.

Personal Access Tokens

Users may create API tokens with:

  • scopes
  • expiry
  • revocation

Tokens act as the user and use the same permission checks.


8. Integrations and Extensions

Integrations

Berry exposes tools to agents from several sources.

Berry-Native Tools

Examples:

  • read task
  • comment
  • create task
  • update permitted status
  • read project
  • save files

GitHub Tools

Examples:

  • read repository
  • create branch
  • push changes
  • open pull request

MCP Tools

External MCP servers may provide arbitrary additional capabilities.

AgentCore Gateway

MCP servers and external tools may optionally be routed through Amazon Bedrock AgentCore Gateway.

Plugins

Berry supports external plugins.

Plugin installation shows requested permissions before installation.

Plugins may contribute:

  • tools
  • pages
  • scheduled actions
  • event handlers
  • integrations

Individual plugin tools remain disabled until an admin enables them.

Plugin access uses short-lived scoped credentials.

Public API

Berry provides an API for scripts and plugins.

It supports operations such as:

  • identify current user
  • list accessible workspaces
  • read tasks
  • edit task fields
  • read comments
  • create comments

All operations use normal Berry permission checks.


9. Runtime Architecture

Berry is a control plane, not the model runtime.

Human
  │
  ▼
Berry
Projects / Goals / Tasks / Policies
  │
  ▼
Run package
  │
  ▼
Amazon Bedrock AgentCore Runtime
  │
  ▼
Agent loop
  │
  ├── Model
  ├── Berry tools
  ├── MCP tools
  ├── Repository
  └── Commands

The Berry server never directly calls an AI model.

A runtime receives the work package, performs the run, and streams events back.

Agent Loop

The standard runtime uses the Strands Agents SDK with models available through Amazon Bedrock.

The loop follows:

Think
  ↓
Call tool
  ↓
Inspect result
  ↓
Continue
  ↓
Deliver

Run-Scoped Credentials

Every run receives a short-lived token.

It is:

  • tied to one run
  • tied to one workspace
  • tied to the relevant task
  • time limited
  • revoked when the run ends

Agents cannot access Berry's database.

They only interact through named tools.

Tool Authorization

Before a tool executes, Berry checks:

  • workspace
  • run
  • task
  • agent
  • autonomy level
  • role contract
  • explicit permissions

The model cannot override these checks.

Workspace Isolation

All resources are scoped to the current workspace.

The run token determines workspace and task context server-side.

An agent cannot choose another workspace through model output.


10. Security

Human Release Authority

Organization agents cannot:

  • merge code
  • mark work Done
  • cancel tasks

These restrictions exist in tool capabilities, not only instructions.

Secret Encryption

Sensitive values are encrypted using AES-256-GCM under the deployment encryption key.

Examples include:

  • plugin secrets
  • agent environment variables
  • MCP headers
  • connected provider credentials

If encryption isn't configured, features requiring protected secret storage are disabled.

GitHub Credentials

GitHub tokens come from user sign-in and remain encrypted.

They are decrypted only when Berry needs GitHub access.

Agents receive operations, not raw user credentials.

Auditability

Every agent-visible action becomes a named tool call.

Berry records:

  • who acted
  • which task
  • which run
  • what tool was called
  • what changed
  • when it happened

11. Events, Observability, and Files

Live Updates

Workspace and board updates are streamed to open clients.

Events are first written durably to PostgreSQL, then streamed.

Clients can reconnect and replay recent events.

Run History

Each run has an append-only ordered event log containing:

  • lifecycle events
  • messages
  • tool calls
  • usage
  • results
  • failures

Run status is derived from this record.

Metrics

Berry exposes:

  • health endpoints
  • readiness checks
  • Prometheus-format metrics
  • run counts
  • failure information
  • model usage
  • estimated cost

Artifacts

Agent-generated files can be stored in S3-compatible storage.

Uploads are associated with the run that created them.

Downloads use temporary signed URLs.


12. Deployment

Berry is self-hosted.

The main application consists primarily of:

  • Node.js server
  • web application
  • PostgreSQL

The agent runtime is separate and normally deployed to Amazon Bedrock AgentCore.

Multiple Berry application servers may run together.

Database locking prevents duplicate run claims and duplicate scheduled executions.


13. Capability Reporting

Berry reports what the current deployment can actually do.

Examples:

  • agent runtime available
  • model access configured
  • storage configured
  • GitHub configured
  • planner available
  • live updates available

Features that depend on missing infrastructure display Needs setup instead of pretending to work.


What Berry Is Fundamentally

Berry combines four systems:

Task Tracker

People and agents share the same work records.

AI Organization

Specialized agents have explicit roles, capabilities, and supervision.

Control Plane

Berry determines what work may happen and under which permissions.

Human Review System

Agents may perform and review work, but a person remains responsible for the final release decision.

The resulting loop is:

Human defines direction
        ↓
Berry plans work
        ↓
Berry routes tasks
        ↓
Agents execute
        ↓
Agents may review
        ↓
Human decides
        ↓
Berry records outcome
        ↓
Agents discover further work
        ↺

The core promise remains:

Agents can do the work. People stay in control.

+ 1 more
Share this project:

Updates

Submission history