Dredge Echo

Ask the live web. Hear the evidence answer back.

Dredge Echo is a grounded, multi-model research agent that does more than generate a fluent answer. It retrieves current evidence, constructs a source-grounded response, independently challenges the material claims, and exposes the reasoning route behind the final result.

Its governing principle is simple:

Evidence outranks model agreement.

Inspiration

The larger idea behind Dredge came to me in a vision on December 25, 2025, while I was homeless with my husband and son and building with only an iPhone 13.

I kept returning to one question: what if an intelligent system did not merely produce answers, but actively organized the intelligence required by each problem?

Dredge Echo is that idea made operational for live research. It is a new research-agent implementation developed during this hackathon on top of a reusable engineering foundation.

The problem

A fluent AI answer can still contain unsupported claims, hide disagreements between sources, or express more certainty than the evidence permits.

Manually checking every statement is slow. Sending every routine question through the most expensive reasoning model is also wasteful.

Dredge Echo addresses both problems by separating four responsibilities:

retrieving evidence; constructing the initial answer; independently challenging its claims; escalating only when the evidence reveals a serious problem.

This turns verification into an executable stage of the system—not a sentence buried inside a prompt.

How Dredge Echo works

Tavily — Scout

Tavily retrieves current web evidence at runtime and creates the source packet used by every downstream model.

GLM-5.3 — Architect

GLM produces the initial grounded synthesis through Nebius Token Factory. It handles the routine path with a fast, economical model.

NVIDIA Nemotron — Challenger

Nemotron independently evaluates material claims against the retrieved evidence. It identifies:

supported claims; partial support; conflicting evidence; unsupported claims; missing context; excessive certainty.

Nemotron is deliberately not used as a second answer writer. Its responsibility is to challenge the answer—not imitate it.

Dredge Echo — Router and Arbiter

Dredge Echo decides what happens next:

A supported answer passes without unnecessary rewriting. A partially supported claim returns to GLM for a focused repair. A conflicted or unsupported claim escalates to Kimi K3. Any revised answer receives a final Nemotron evidence check.

Kimi K3 — Escalation model

Kimi is reserved for cases where conflicting or unsupported claims justify deeper reasoning. It reconciles the draft and critique against the evidence packet rather than simply choosing which model sounds more convincing.

What the user receives

Dredge Echo returns:

a grounded final answer; live Tavily sources; claim-level evidence status; the Nemotron verification result; stage-level routing and latency information; the arbitration route used—or a visible SKIPPED state; an insufficient-evidence response when the sources cannot support an answer.

The result is not just an answer. It is an inspectable decision path.

Routing behavior and benchmark evidence

The routing policy is covered by deterministic tests using controlled provider responses.

For the tested routes:

Supported draft: two model calls instead of four under the always-escalate baseline—a 50% reduction in model calls for that path. Partial correction: GLM performs the economical repair before the final Nemotron check. Serious dispute: Dredge Echo preserves Kimi escalation and final verification.

The benchmark harness records routing decisions, model call counts, provider-reported token usage, measured latency, and final evidence status.

In the final live three-question paired benchmark, each comparison used the same Tavily evidence packet, GLM draft, and initial Nemotron assessment. Dredge Echo completed all three pairs and produced evidence-supported final answers in 3/3 cases, compared with 1/3 for the always-Kimi baseline. Dynamic routing used 9 model calls instead of 13 (30.8% fewer), 47,172 total tokens instead of 65,567 (28.1% fewer), and 211.9 seconds of aggregate latency instead of 275.5 seconds (23.1% less).

Dredge Echo skipped unnecessary escalation on two supported questions and spent additional reasoning on the one answer that required repair. The benchmark is a transparent three-question controlled sample; it does not claim general production accuracy or exact dollar savings because provider pricing differs by model.

Why Nebius and NVIDIA matter

Nebius Token Factory made this multi-model architecture practical through one OpenAI-compatible inference interface.

Instead of building separate integrations for every model, I could assign distinct responsibilities to GLM, NVIDIA Nemotron, and Kimi while preserving one observable orchestration path.

NVIDIA Nemotron is central to the product—not decorative. It acts as the independent evidence critic on every completed research path and checks revised answers before they are returned to the user.

How I built it

I used Codex as a collaborative engineering partner while I remained the system architect.

I directed the separation of retrieval, synthesis, verification, repair, and escalation into explicit components. Together, we implemented the provider adapters, strengthened failure boundaries, created routing and verification tests, diagnosed live deployment problems, and refined the judge-facing experience.

When live behavior showed that using Kimi for routine synthesis was too expensive and slow, I redesigned the architecture:

GLM handles routine synthesis. Nemotron challenges the evidence. Kimi is reserved for consequential disputes. Dredge Echo decides which intelligence the problem deserves.

Built during the hackathon

The hackathon build added:

Nebius Token Factory inference; Tavily runtime retrieval; GLM grounded synthesis; NVIDIA Nemotron claim verification; evidence-aware routing; GLM repair and Kimi escalation; final-answer re-verification; a Gradio interface on Hugging Face Spaces; execution traces and failure visibility; deterministic routing and benchmark tests; live GitHub Actions verification.

Technology

Python, Gradio, Nebius Token Factory, NVIDIA Nemotron 3 Nano 30B A3B, GLM-5.3, Kimi K3, Tavily, Hugging Face Spaces, and GitHub Actions.

Links

Working demo Public source code Demo video

Built With

  • github-actions
  • gradio
  • hugging-face-spaces
  • kimi
  • nebius-token-factory
  • nvidia-nemotron
  • python
  • tavily
Share this project:

Updates

Submission history