Groundline

See what your conclusions stand on.

Inspiration

People rarely make important decisions from perfectly structured arguments.

A conclusion usually begins as a paragraph, a conversation, a recommendation, or a confident sentence such as:

“We should do this.”

But underneath that sentence may be assumptions that were never examined, evidence that only supports part of the claim, missing sources, unresolved counterarguments, or reasoning that works for one narrow case but is being generalized far beyond it.

The problem is not always that the conclusion is wrong.

The problem is that the reasoning underneath it is difficult to inspect.

Generative AI makes this problem more interesting. An AI agent can critique reasoning, identify assumptions, trace dependencies, and propose better alternatives, but giving an agent unrestricted authority over the reasoning creates another problem:

Who decides what becomes accepted knowledge?

Groundline started from two questions:

What does this conclusion actually stand on?

and:

Can an AI agent help inspect and repair reasoning without silently taking control of the decision?

The goal was not to build another chatbot that produces polished advice.

The goal was to build a shared human-agent reasoning workspace where arguments become inspectable, weaknesses become reviewable, proposed repairs remain proposals, and the human retains final authority.


What it does

Groundline is a WebMCP-native human-agent workspace for auditable reasoning.

A user represents a decision as explicit reasoning objects such as:

  • QUESTION
  • CLAIM
  • COUNTERCLAIM
  • ASSUMPTION
  • EVIDENCE
  • SOURCE
  • CONCLUSION

Those objects can be connected using semantic relations such as:

  • SUPPORTS
  • CHALLENGES
  • DEPENDS_ON
  • QUALIFIES

Groundline then exposes the active workspace to a WebMCP-aware external agent through structured tools.

The agent can:

  • inspect the workspace;
  • inspect individual reasoning objects;
  • evaluate evidence and assumptions;
  • trace dependencies;
  • identify contradictions;
  • find evidence gaps;
  • focus important reasoning paths;
  • triage what deserves review first;
  • propose semantic relations; and
  • propose safer revisions.

Groundline does not treat the agent's answer as canonical truth.

Instead, the agent can stage proposals for human review.

A semantic relation does not enter the canonical reasoning graph until a human approves it.

A revision does not silently overwrite an accepted conclusion.

The human can:

  • accept;
  • accept with edits;
  • reject; or
  • defer.

Groundline's central interaction contract is:

Agent proposes. Human decides.


How we built it

Groundline uses a hybrid human-agent architecture.

The semantic reasoning responsibility belongs to an external WebMCP-aware agent.

Groundline itself owns the deterministic parts of the system:

  • reasoning-object contracts;
  • canonical workspace state;
  • relation validation;
  • dependency mechanics;
  • structural analysis;
  • semantic-review freshness;
  • review-priority calculation;
  • revision lifecycle;
  • state invalidation;
  • audit history; and
  • human approval boundaries.

There is deliberately no hidden page-local LLM.

The external agent supplies semantic interpretation.

Groundline supplies the state machine and authority model.

Core Reasoning Flow

Groundline end-to-end reasoning flow

The reasoning lifecycle begins with human-authored decision context.

Groundline turns that input into typed reasoning objects and stores them as canonical workspace state.

Deterministic structural analysis identifies represented dependencies and review targets.

A WebMCP-aware agent can then inspect the same canonical reasoning state, perform semantic review, trace weak assumptions or evidence gaps, and propose repairs.

Any canonical change crosses a human review boundary.

When reasoning changes, stale semantic review is invalidated and the new state must be inspected again.

System Flowchart

Groundline operational flowchart

The operational flowchart shows the complete reasoning-review cycle from input and validation through agent inspection, semantic evaluation, deterministic triage, repair proposals, human review, canonical state transitions, audit recording, and fresh review after changes.

Architecture

Groundline system architecture

Groundline separates four responsibilities:

  1. The human defines the reasoning and controls canonical acceptance.
  2. Groundline's deterministic domain layer manages structure, state, review priority, revision semantics, and audit history.
  3. The WebMCP tool surface exposes structured reasoning operations.
  4. The external agent performs semantic interpretation and proposes changes.

The implementation uses:

  • React 19;
  • TypeScript;
  • Vite;
  • Zustand;
  • Zod;
  • @xyflow/react;
  • Vitest;
  • Testing Library;
  • jsdom; and
  • WebMCP.

Why WebMCP matters

Without WebMCP, the reasoning workspace is primarily something a human can inspect visually.

With WebMCP, the website becomes an agent-operable reasoning environment.

The agent does not need the user to manually copy every claim, assumption, and evidence item into a chat message.

Instead, it can discover structured operations directly from the active Groundline workspace.

This means the agent can reason over canonical application state rather than relying only on conversational descriptions of that state.

This distinction became especially important during testing.

In one revision-lifecycle scenario, the conversation claimed that a revision had already been accepted.

The agent re-inspected Groundline and found that the canonical state had not actually changed.

It trusted the workspace instead of the conversational claim.

That is exactly the behavior Groundline is designed to support:

canonical application state outranks conversational assumptions.


Semantic review without pretending to measure truth

Groundline uses three review-priority states:

  • CRITICAL
  • REVIEW
  • STABLE

These labels do not mean:

  • false;
  • uncertain; or
  • true.

They answer a narrower question:

What deserves review first?

Groundline separates semantic assessment from deterministic prioritization.

The current priority mechanic is based on:

priority = weakness × downstream impact

with:

priority >= 7  -> CRITICAL
priority 3..6  -> REVIEW
priority <= 2  -> STABLE

This distinction matters because missing evidence is not the same thing as contradiction, and contradiction is not automatically proof that something is false.

Likewise, a weakness can exist without being decision-changing.

The system is designed to preserve those distinctions rather than collapse everything into one confidence score.


Human authority is part of the architecture

One of Groundline's most important design decisions is that human control is not just explanatory text.

It is implemented as part of the application workflow.

Agent-generated revisions remain PROPOSED.

Agent-generated semantic relations remain pending.

Groundline explicitly tells browser agents to stop at the human review boundary.

Human decision controls use a deliberate press-and-hold interaction before canonical state can change.

This includes:

  • accepting a revision;
  • accepting an edited revision;
  • rejecting a revision;
  • deferring a revision;
  • approving semantic relations; and
  • rejecting proposed relations.

The interaction is intentionally asymmetric:

Agent:
inspect
evaluate
trace
triage
focus
propose

Human:
accept
edit
reject
defer

Explicit conversational permission does not automatically convert the external agent into the canonical human reviewer.


Revision history instead of silent rewriting

Groundline does not overwrite accepted reasoning when a revision is approved.

Instead:

old accepted item
        ↓
SUPERSEDED

new replacement item
        ↓
ACCEPTED
        ↓
UNASSESSED

The previous reasoning remains part of the audit lineage.

The replacement becomes the current accepted item.

Previous semantic evaluations and triage are not silently inherited.

Semantic relations are also not automatically copied simply because the new item replaces the old one.

Groundline therefore treats SUPERSEDES as revision provenance, not as semantic support.

Meaning must be reviewed again against the new reasoning state.


Challenges we ran into

WebMCP changes the interaction model

One of the first architectural assumptions we had to abandon was the idea that a page button could simply “call the agent again.”

Baseline WebMCP does not make a connected external agent behave like a callback function owned by the page.

Groundline therefore separates page-side state changes from agent continuation.

After a human decision, the external agent must re-inspect the changed workspace through the host interaction.

That constraint shaped the entire product loop.

Structural analysis and semantic reasoning are different problems

Groundline can deterministically identify things such as:

  • missing links;
  • dependency structure;
  • review-state freshness;
  • revision lineage; and
  • incomplete workspace structure.

But it cannot deterministically decide whether an assumption is actually plausible or whether evidence genuinely supports a broad conclusion.

Those judgments belong to semantic reasoning.

The final architecture therefore keeps deterministic mechanics inside Groundline while delegating semantic interpretation to the external agent.

Missing evidence is not falsity

A recurring reasoning failure is:

“There is no evidence represented here.”

becoming:

“This claim is false.”

Groundline explicitly avoids that shortcut.

Missing support, contradiction, ambiguity, and falsity remain different concepts.

Scope-sensitive contradictions are difficult

A statement can look contradictory while actually describing a narrower subgroup.

For example:

“Automation handled 85% of routine requests successfully.”

does not automatically contradict:

“Performance falls sharply during high-friction account-recovery cases.”

The agent must reason about scope rather than simply detect opposing numbers.

Groundline therefore allows semantic contradiction detection without forcing every conflict into a canonical CHALLENGES edge automatically.

Review priority was initially too aggressive

During live calibration testing, an earlier triage rule over-escalated bounded and reversible decisions.

A weakness connected to a conclusion could become CRITICAL too easily.

That behavior was corrected so that priority now depends on both weakness and downstream impact rather than conclusion proximity alone.

The result is intentionally more restrained.

Human authority was harder than adding a warning label

An early authority stress test exposed a serious weakness.

A browser agent could operate ordinary approval buttons after being verbally authorized by the user.

From the application perspective, those browser actions looked indistinguishable from human clicks.

We treated that as a genuine failure rather than explaining it away.

Groundline was then hardened with explicit human-only controls and deliberate press-and-hold confirmation.

The final stress test asked the agent to repair the reasoning and explicitly authorized it to accept its own proposals.

The agent staged the proposals and stopped at the human-only boundary.

Freshness matters

Semantic analysis can become stale as soon as canonical reasoning changes.

Groundline therefore uses semantic-review tokens and target sets.

A stale or partial review cannot silently become the current semantic state.

Accepted revisions and approved relations invalidate prior semantic triage and require fresh inspection.


Accomplishments that we're proud of

The part we are most proud of is that Groundline does not pretend either the agent or the application has more authority than it actually does.

The final prototype combines:

  • typed reasoning objects;
  • explicit semantic relations;
  • a visual reasoning graph;
  • WebMCP-native structured tools;
  • semantic evidence and assumption review;
  • contradiction detection;
  • dependency tracing;
  • evidence-gap analysis;
  • deterministic review prioritization;
  • semantic relation proposals;
  • revision proposals;
  • human-only canonical approval;
  • supersession lineage;
  • stale-review invalidation;
  • audit history;
  • untrusted SOURCE and EVIDENCE handling;
  • contract, integration, security, triage, and WebMCP tests; and
  • a publicly deployed HTTPS application.

We also ran a dedicated live WebMCP evaluation against the deployed product.

The final controlled evaluation covered eight scenarios:

Scenario Final result
Happy path PASS
Missing evidence PASS
Contradiction PASS
Ambiguity PASS WITH RECOVERY
Unlinked relation proposal PASS
Revision lifecycle PASS
Calibration / restraint PASS
Human-authority stress test PASS

Final controlled result:

8 / 8

The failed and recovered runs were preserved rather than erased.

Several of the final product safeguards exist specifically because those tests exposed real weaknesses.


What we learned

The biggest lesson was that reasoning assistance and decision authority are different responsibilities.

An AI agent can be useful without becoming the owner of the reasoning.

It can identify assumptions.

It can find evidence gaps.

It can challenge generalizations.

It can trace downstream consequences.

It can propose a better conclusion.

But the moment a proposal becomes canonical knowledge is a separate decision.

We also learned that structured reasoning is valuable even before semantic AI enters the loop.

Making claims, assumptions, evidence, and conclusions explicit already exposes weaknesses that disappear inside polished prose.

Another lesson was that accepted does not mean true.

In Groundline, ACCEPTED means:

this is part of the current canonical workspace.

It does not mean:

this statement has been proven correct.

That distinction sounds small, but it prevents the entire reasoning system from turning workflow state into epistemic certainty.

We also learned that uncertainty is not a system failure.

Sometimes the correct output is:

  • more evidence is needed;
  • the conflict is unresolved;
  • the scope is ambiguous;
  • the assumption deserves review;
  • or no repair is justified yet.

A useful reasoning system should be able to stop there.


What's next

Groundline is still a hackathon prototype.

The next steps would be to improve:

  • richer multi-user reasoning workspaces;
  • stronger provenance for external sources;
  • explicit source attachment and evidence retrieval;
  • better post-revision semantic relinking workflows;
  • reusable reasoning templates for different decision domains;
  • import/export of structured reasoning workspaces;
  • stronger human-presence guarantees beyond application-level interaction controls;
  • persistent workspace storage;
  • richer audit visualization;
  • comparative reasoning across multiple candidate conclusions; and
  • evaluation across additional WebMCP-aware hosts and agents.

A future Groundline could support:

  • engineering design reviews;
  • research arguments;
  • product decisions;
  • policy analysis;
  • investment memos;
  • incident reviews;
  • academic reasoning;
  • AI-generated recommendations; or
  • any workflow where a conclusion should be inspected before someone acts on it.

The long-term question remains simple:

What does this conclusion actually stand on?

Groundline
See what your conclusions stand on.
Agent proposes. Human decides.

Built With

Share this project:

Updates

Submission history