GenAI Evidence Workbench has evolved from a research prototype into a reproducible workflow for inspecting, verifying, and exporting AI-generated evidence.
Recent improvements include:
• deterministic evidence exports • clearer provenance and verification records • reproducible forensic reports • vendor-neutral evaluation workflows • improved documentation and deployment
The project is designed around a simple principle: plausible AI output is not sufficient. Claims should remain traceable to observable evidence, execution history, and reproducible artifacts.
I also published a live deployment and expanded the project documentation so evaluators and developers can inspect the workflow directly.
Live app: https://genai-evidence-workbench.hiro4ever.chatgpt.site/
GitHub: https://github.com/hiroki-tamba-research/GenAI-Evidence-Workbench
The next phase will focus on long-horizon agent evidence, including interruption, compaction, resume behavior, and trajectory-level verification.
Log in or sign up for Devpost to join the conversation.