-
-
FastMCP server and Strands agent running tool calls over stdio.
-
Automated daily cron workflow running unattended in GitHub Actions.
-
Multi-agent architecture with deterministic scoring and vector verification.
-
Radical honesty: ungrounded claims are openly flagged and rejected.
-
Audit dashboard displaying verified grounding matches and score math.
Inspiration
The "Agents for Humans" brief describes it exactly: a solo researcher facing a mountain of papers. I'm that researcher — following a fast-moving field means dozens of new arXiv papers a day, with no realistic way to read them all. The hackathon's own thesis, "autonomous where safe, deterministic where critical, human where necessary," gave me a concrete shape for solving that without building another chatbot I'd have to remember to open.
What it does
Marginalia runs on a schedule, not a chat window. Every day it:
- Pulls new papers from arXiv, once per topic I'm tracking.
- Has an LLM agent read the abstracts and draft a one-sentence connection to my established research themes — only where a genuine connection exists.
- Runs every claim through two independent checks before it reaches me: a deterministic relevance score (plain code, zero hallucination risk) and a semantic verification pass against my own notes in a vector database.
- Sends survivors through one more adversarial pass — a second "Skeptic" agent that actively looks for reasons the first agent's claim overstates the connection.
- Shows me a dashboard where I approve or skip each paper. Approvals get embedded back into my knowledge base, so next week's digest has more to check new papers against than this week's did.
Claims that don't hold up aren't hidden — they're shown, labeled honestly, next to the ones that do.
How I built it
Built on AWS Strands Agents SDK, with Groq (openai/gpt-oss-120b) as the model provider, arXiv's public API for retrieval, Qdrant Cloud for vector verification, sentence-transformers for local embeddings, GitHub Actions as the actual scheduled trigger, and a small Flask app for the human-approval dashboard. The paper-fetching tool is also wrapped as a real, standalone MCP server, independently reusable by any MCP-compliant client, not just this agent.
I worked with Claude as an architecture and design mentor throughout, and wrote the code in Google Antigravity, running Gemini — a genuinely AI-assisted build from planning through implementation, with every phase tested against real output before moving to the next.
Challenges I ran into
- The Groq model I planned around got decommissioned mid-build. The fix was querying Groq's live model list directly instead of trusting a hardcoded assumption.
- arXiv's search API doesn't automatically phrase-match multi-word queries — an unquoted search for "large language model agents" silently returned whatever was newest across all of arXiv, with zero errors. It looked like it was working. It wasn't. That gap between "ran without crashing" and "produced correct output" turned out to be the recurring lesson of this build.
- My first keyword list used spelled-out phrases like "language model agents," but the literature overwhelmingly writes "LLM." The deterministic scorer was quietly missing most of what it should have caught, until I compared the list against what real abstracts actually said.
Accomplishments that I'm proud of
- A genuinely background-triggered agent — it runs on a GitHub Actions schedule and produces a digest with nobody prompting it.
- A real MCP server, verified end-to-end by a live client actually discovering and calling its tool over the protocol.
- An adversarial verification tier where a second agent has to independently agree a claim holds up, and says exactly why when it doesn't.
- A feedback loop that's actually closed: every approval makes the next digest a little sharper, not just a nicer UI.
What I learned
The most useful design decision in this project wasn't a clever prompt — it was deciding, explicitly, which parts of the pipeline the model is allowed discretion over, and which parts always run in plain code no matter what the model does or says. The scorer and verifier are never things the agent can choose to skip. That distinction is the entire safety story, and it's simpler to build than it sounds.
What's next for Marginalia
- A small evaluation harness reporting real precision/recall against a synthetic benchmark, instead of "it seems to work."
- Lightweight linking between approved notes as the corpus grows.
- Enriching candidates with citation counts from a free, public API like OpenAlex — extra signal, never a new gate.
Log in or sign up for Devpost to join the conversation.