Why We Built It

AI agents are becoming increasingly capable of taking actions, making recommendations, and synthesizing large amounts of information. But capability is not the same as reliability.

Selection Lab began with a simple question: What happens when an AI agent encounters incomplete evidence, conflicting sources, missing context, or uncertainty — but is still expected to produce an answer?

The project grew from more than two decades of human-centered visual and pattern research and our current work on Human-Grounded AI Reliability Infrastructure. We wanted to translate those principles into an agent that does not simply generate a plausible response, but actively investigates the quality of the information behind it.

What It Does

The Selection Lab agent acts as an evidence-first reliability layer.

Instead of immediately collapsing information into a single answer, it examines the available evidence and looks for:

  • contradictions between sources or claims
  • missing information that could materially change the conclusion
  • uncertainty that should be made explicit
  • provenance and evidence quality
  • alternative explanations or hypotheses
  • decisions that should remain under human review

The goal is not to replace human judgment. It is to make human judgment better informed.

How We Are Building It

For the Agents for Humans Hackathon, we are translating this research architecture into a working agent using the Strands Agents SDK.

The workflow separates investigation from final synthesis: the agent gathers and evaluates evidence, identifies conflicts and gaps, and then produces a structured output designed for human review rather than hiding uncertainty behind a confident answer.

This hackathon is an opportunity to turn a broader reliability methodology into a focused, testable agent architecture.

What We Learned

One of the most important lessons is that useful AI does not always need to provide more answers. Sometimes it needs to know when an answer is not yet justified.

Reliability requires preserving uncertainty, exposing contradictions, and keeping humans in the loop when evidence is insufficient.

Challenges

The central challenge is balancing usefulness with epistemic restraint. An agent that refuses everything is not useful; an agent that confidently fills every gap can be dangerous.

Our challenge is to design the space between those extremes: an agent that can act and reason effectively while showing people what it knows, what conflicts, what is missing, and where human judgment is still required.

Built With

Share this project:

Updates

Submission history