Inspiration

Security and RFP questionnaires look like writing work, but the expensive part is evidence control. Teams repeatedly search policies, audit reports, architecture documents, and previously approved responses to prove that every statement is current and consistent.

A polished answer without a source is worse than a blank field: it can create contractual, audit, and trust risk. We built EvidenceOps around one operating principle:

Uploaded documents are evidence. Generated text is not.

EvidenceOps is for security, GRC, trust, and solutions engineering teams that must answer buyer questionnaires under deadline without weakening evidence discipline. It turns document search into an exception queue: supported drafts, source conflicts, and precise evidence requests, with a person deciding what leaves the system.

What it does

EvidenceOps handles the questionnaire workflow end to end:

  1. Ingest a questionnaire in PDF, XLSX, CSV, DOCX, Markdown, or text form.
  2. Parse it into traceable question records with source locations.
  3. Search only the uploaded evidence library for relevant passages.
  4. Use a Strands agent to draft a concise answer with selected citation indexes.
  5. Check every draft for unsupported claims, citation integrity, numeric changes, and conflicting evidence.
  6. Turn missing support into a specific evidence request instead of plausible prose.
  7. Require a person to approve, edit, or reject every answer.
  8. Export approved work as XLSX, CSV, or JSON with citations and reviewer notes.

The application does not expose generic prompt forwarding, distribute model credentials, or auto-submit answers to a customer.

What makes it different

EvidenceOps is the agent that knows when not to answer.

Retrieval alone is not treated as proof. Deterministic controls keep ownership of document identity, source locators, contradiction detection, approval state, and approved-only export. Strands owns the model-assisted drafting step through a narrow typed result: answer text plus the evidence indexes that support it.

If a provider fails, the system degrades to evidence-only text. If evidence is missing, the item stays blocked. If two policies conflict, both remain visible until a reviewer records a disposition.

What the working demo proves

The synthetic CloudDesk review contains eight questions and four evidence documents. It deliberately exercises the cases that matter most:

  • A supported encryption answer resolves to its exact source quote.
  • Two incident-response documents disagree on a 48-hour versus 72-hour notification timeline. EvidenceOps displays both instead of silently choosing one.
  • No current subprocessor register exists in the evidence pack. The item becomes a missing-evidence request rather than an invented answer.
  • Editing an approved answer clears the previous approval.
  • Approved-only export prevents drafts and unresolved items from being presented as final.

This is a working application, not a slide path. The repository includes the responsive Web workspace, API, synthetic fixtures, tests, architecture, restricted verification action, and export implementation.

How we built it

The MVP uses Python, FastAPI, SQLite, Pydantic, and the Strands Agents SDK. A reusable EvidenceOpsService coordinates parsing, evidence retrieval, grounded drafting, verification, workflow state, human review, and export.

The provider boundary supports an operator-controlled OpenAI-compatible endpoint or Amazon Bedrock. It applies request throttling, bounded retries, a circuit breaker, and a deterministic grounded fallback.

A release-gate integration test invokes the real Strands Agent and OpenAIModel, uses a typed AgentDraft output tool, and then runs the result through EvidenceOps retrieval and grounding. The loopback protocol fixture proves the SDK integration without presenting it as hosted-model or Bedrock inference.

The same service exposes a restricted verify-evidence action through structured HTTP and an optional MCP adapter. It accepts a tenant-scoped project, proposed answer, and exact citations. It is a fixed business action, not a general chat proxy.

Challenges we ran into

Retrieval is not proof

A passage can be related to a question without supporting every clause in a draft. We separated retrieval from grounding and made evidence coverage independent from model confidence.

Contradictions must remain visible

Policies age at different rates. Automatically choosing the more convenient source would hide risk, so EvidenceOps preserves both passages and requires a human disposition.

Approval changes the data model

Approval cannot be a decorative final button. Draft text, citations, checks, decision status, and reviewer notes are persisted separately. Any later edit or new evidence invalidates stale approval.

Provider failure must not become factual failure

Retries and fallback are useful only when fallback stays grounded. EvidenceOps degrades to cited evidence or an explicit gap, never an uncited best guess.

Accomplishments

  • Completed the full workflow from heterogeneous upload to reviewable export.
  • Made citations, conflicts, and evidence gaps first-class workflow objects.
  • Added persisted human approval with invalidation after edits or new evidence.
  • Passed 28 automated tests with 88% application coverage.
  • Verified a real Strands SDK typed-output path and a real MCP call_tool path.
  • Shipped a responsive workspace, public MIT repository, synthetic fixtures, architecture package, and reproducible validation.

What we learned

For high-stakes document work, the most valuable agent behavior is often a structured refusal to guess: identify the unsupported claim, show the evidence that was found, and request the exact missing source.

Trustworthy citations also require stable document identity and locators. A generated footnote is not evidence. Human review works best when it is part of workflow state, not a disclaimer after generation.

What's next

The optional AWS deployment path preserves the same trust boundary: Amazon S3 for encrypted documents and exports, DynamoDB for workflow state, CloudWatch for redacted telemetry, and AgentCore or a container service for the Strands runtime. Larger evidence sets could move to OpenSearch Service or Knowledge Bases for Amazon Bedrock.

Production hardening would add tenant authentication, role-based approval, malware scanning, retention controls, immutable audit logging, and organization-specific evaluation sets.

Disclosure

EvidenceOps was created during the 2026 Agents for Humans Hackathon submission period with implementation and documentation assistance from OpenAI Codex. It uses synthetic demonstration fixtures and contains no private customer evidence. EvidenceOps assists document preparation; it is not a certification, audit opinion, legal determination, or autonomous submission system.

Built With

Share this project:

Updates

Submission history