Why We Built It
AI agents are becoming increasingly capable of taking actions, making recommendations, and synthesizing large amounts of information. But capability is not the same as reliability.
Selection Lab began with a simple question: What happens when an AI agent encounters incomplete evidence, conflicting sources, missing context, or uncertainty — but is still expected to produce an answer?
The project grew from more than two decades of human-centered visual and pattern research and our current work on Human-Grounded AI Reliability Infrastructure. We wanted to translate those principles into an agent that does not simply generate a plausible response, but actively investigates the quality of the information behind it.
What It Does
The Selection Lab agent acts as an evidence-first reliability layer.
Instead of immediately collapsing information into a single answer, it examines the available evidence and looks for:
- contradictions between sources or claims
- missing information that could materially change the conclusion
- uncertainty that should be made explicit
- provenance and evidence quality
- alternative explanations or hypotheses
- decisions that should remain under human review
The goal is not to replace human judgment. It is to make human judgment better informed.
How We Are Building It
For the Agents for Humans Hackathon, we are translating this research architecture into a working agent using the Strands Agents SDK.
The workflow separates investigation from final synthesis: the agent gathers and evaluates evidence, identifies conflicts and gaps, and then produces a structured output designed for human review rather than hiding uncertainty behind a confident answer.
This hackathon is an opportunity to turn a broader reliability methodology into a focused, testable agent architecture.
What We Learned
One of the most important lessons is that useful AI does not always need to provide more answers. Sometimes it needs to know when an answer is not yet justified.
Reliability requires preserving uncertainty, exposing contradictions, and keeping humans in the loop when evidence is insufficient.
Challenges
The central challenge is balancing usefulness with epistemic restraint. An agent that refuses everything is not useful; an agent that confidently fills every gap can be dangerous.
Our challenge is to design the space between those extremes: an agent that can act and reason effectively while showing people what it knows, what conflicts, what is missing, and where human judgment is still required.
Log in or sign up for Devpost to join the conversation.