AdPilot Memory Arena
Inspiration
AdTech platforms make millions of decisions under strict time and budget constraints. An optimization that looks successful immediately can become harmful once delayed conversions, CPA, or fraud signals arrive. We wanted to demonstrate that autonomous agents should not only react quickly—they should remember what happened after previous decisions and avoid repeating failures.
What it does
AdPilot Memory Arena is a live AdTech simulation where two autonomous AI agents optimize the same advertising campaign:
- A baseline agent reacts only to current campaign metrics.
- A memory agent uses time-aware Elasticsearch memory to retrieve previous experiments, recognize superseded strategies, and avoid known failures.
Both agents receive identical auction traffic and make bounded optimization decisions. Initially, the baseline may appear to perform better by increasing bids and expanding supply. As delayed conversion and fraud data arrive, the hidden cost of that strategy becomes visible. The memory agent instead applies a safer strategy grounded in previous outcomes.
A live dashboard displays campaign performance, agent decisions, memory evidence, guardrail results, rollbacks, and the final winner.
How we built it
We used:
- Mastra to build and trace the autonomous optimization agents.
- Elasticsearch as episodic, time-aware memory.
- ES|QL
FORK, weightedFUSE, andDECAYfor hybrid retrieval that combines keyword relevance, semantic similarity, and recency. - A seeded TypeScript auction simulator so both agents face identical, reproducible traffic.
- Deterministic guardrails to validate bid changes, pacing, spend, supply quality, CPA, ROAS, and invalid traffic.
- Next.js for the live control-room dashboard.
The agents propose actions, but deterministic code owns policy enforcement, application, rollback, and winner calculation. This keeps the system autonomous while preventing unsafe model output from directly controlling the simulation.
Challenges
The hardest challenge was making delayed outcomes meaningful within a short demo. Advertising decisions often look successful before conversions and fraud labels mature, so we compressed that feedback loop into a reproducible 60–90 second simulation.
We also needed a fair comparison. Both agents therefore receive the same opportunities and potential outcomes using common seeded randomness. Their results differ only because of their policy decisions.
Finally, memory retrieval had to favor current evidence without ignoring relevant history. We combined exact operational matching, semantic search, time decay, and explicit supersession relationships so newer findings can override older recommendations.
What we learned
We learned that memory is most valuable when an agent must reason about consequences that arrive later. Current metrics explain what is happening now, but episodic memory provides evidence about what happened after similar actions in the past.
We also learned that reliable autonomy requires clear boundaries: the model can interpret evidence and propose a strategy, while deterministic guardrails protect budgets, quality, and safety.
AdPilot Memory Arena demonstrates how an optimizer can progress from reactive automation to a continuously learning system that makes safer decisions with every experiment.
Built With
- codex
- elasticsearch
- github
- mastra
- mateo'sgreatbrain
- next.js
Log in or sign up for Devpost to join the conversation.