posted an update

GenAI Evidence Workbench has evolved from a research prototype into a reproducible workflow for inspecting, verifying, and exporting AI-generated evidence.

Recent improvements include:

• deterministic evidence exports • clearer provenance and verification records • reproducible forensic reports • vendor-neutral evaluation workflows • improved documentation and deployment

The project is designed around a simple principle: plausible AI output is not sufficient. Claims should remain traceable to observable evidence, execution history, and reproducible artifacts.

I also published a live deployment and expanded the project documentation so evaluators and developers can inspect the workflow directly.

Live app: https://genai-evidence-workbench.hiro4ever.chatgpt.site/

GitHub: https://github.com/hiroki-tamba-research/GenAI-Evidence-Workbench

The next phase will focus on long-horizon agent evidence, including interruption, compaction, resume behavior, and trajectory-level verification.

Log in or sign up for Devpost to join the conversation.