-
The agent calls listSubjects, listTopics, then findStatements for each topic — querying Elasticsearch 11 times to build the full timeline.
-
Sam Altman — 50% Consistency Index. Five topics held steady, five shifted. Real quotes from primary sources with dates."
-
Shifted topics show THEN/NOW quotes. e.g. Open Source: 'share with the world' (2015) to 'we were wrong about openness' (2023).
Inspiration
Public figures say things on the record — then sometimes say the opposite a year later. Someone on the internet digs up the old clip, and that person just got "receipted." But it's always manual: one person, one topic, one viral moment. Nobody tracks consistency systematically across topics and time. We wanted to build the tool that does.
What it does
Receipts ingests timestamped public statements into Elasticsearch, organizes them by person and topic, and then an AI agent reads the chronological timeline for each topic and classifies it as consistent or shifted. The result is a Consistency Index: what percentage of a person's stated positions held steady over time.
The demo compares Sam Altman (OpenAI) and Dario Amodei (Anthropic) across 12 topics — AI safety, open source, jobs, regulation, military contracts, corporate structure, and more. Altman scored 40%. Amodei scored 92%.
Then someone in the audience names a third person, and the agent runs a Rapid Receipt — extracting known positions, indexing them into Elasticsearch, and scoring consistency in about 15 seconds.
Two tiers of truth: Gold Standard (67 curated quotes from 29 primary sources) and Rapid Receipt (LLM-extracted, lower confidence, real-time). The system is honest about which one you're getting.
How we built it
- Elasticsearch stores all statements and serves as the agent's memory layer. We use term-based filtering (subject + topic), date sorting for chronological timelines, aggregations for topic discovery, and bulk indexing with immediate refresh for real-time onboarding of new subjects.
- Mastra orchestrates the agent with four custom tools (Zod-validated schemas): list_subjects, list_topics, find_statements, and rapid_receipt. The agent decides which tools to call and in what order based on the user's question.
- Gemini 2.5 Flash via OpenRouter does the consistency judgment. We deliberately chose a model not made by either company being evaluated — neither OpenAI nor Anthropic grades its own CEO's homework.
- The corpus (67 statements across 12 topics) was researched from 29 primary sources before the event. The agent, tools, and Elasticsearch integration were built during the hackathon.
Challenges we ran into
Getting the Rapid Receipt to return structured, indexable data reliably — LLM output parsing across different response formats required defensive handling. Also, tuning the agent's instructions so it consistently calls tools in the right order (list subjects first, then topics, then statements) rather than trying to answer from its own knowledge.
Accomplishments that we're proud of
The two-tier confidence model. Most AI tools pretend everything is equally reliable. Receipts is explicit: the Gold Standard has sourced quotes from congressional testimony and published essays; the Rapid Receipt is LLM-extracted and says so. The honest label is the product.
Also: the model neutrality decision. Using Gemini to evaluate the CEOs of OpenAI and Anthropic means no vendor is grading its own homework.
What we learned
Elasticsearch's structured query capabilities — term filters, aggregations, date sorting — were more valuable for this use case than the vector search / RAG pattern in the starter repo. When you ask what someone said about open source, you want every statement in chronological order, not semantically similar results. We deliberately chose precision retrieval over similarity search. The right tool for memory depends on what kind of remembering you're doing.
What's next for Receipts
Expanding the Gold Standard corpus beyond AI CEOs to politicians, executives, and public intellectuals. Adding source URL verification so every quote links back to its primary source. And exploring a time-weighted confidence score that accounts for how much evidence exists per topic — a topic with 8 quotes across 5 years is a stronger signal than one with 2 quotes across 6 months.
Built With
- elasticsearch
- gemini
- mastra
- node.js
- openrouter
- typescript
Log in or sign up for Devpost to join the conversation.