Inspiration

I was working on a benchmark where I needed to repeatably run the same process over documents and wanted to script/orchestrate codex subagents, a la Claude workflows.

What it does

gpt-workflow lets you write multi-agent workflows in JavaScript. Your code controls loops, branches, retries, and parallelism, while Codex agents handle bounded judgment. Runs can produce structured output and can resume without repeating completed work.

How we built it

First, I had Claude generate an authoritative set of workflows.

Then, we spec'd out an implementation with the Codex App Server.

With a goal, I had it build it by delegating to Codex 5.6 Sol subagents for implementation. This got us 90% there.

I dogfooded it in my benchmark project and, using Codex, refined the package, the skills, created the plugin, published it, and created a Python wrapper.

Challenges we ran into

  1. There is no authoritative Claude workflows spec — the initial set of replication scripts were insufficient
  2. Codex is too eager to put too much logic into the scripts, I had to update the skills to tell it to chill

Accomplishments

Using this, I was able to script/standardize my data preparation pipeline for a first-of-its-kind deterministic long-horizon medical evidence synthesis benchmark. I’m also proud that the result is a standalone CLI, library, and Codex plugin rather than a one-off script.

What's next

Launching gpt-workflow and the medical evidence synthesis benchmark =)

Built With

Share this project:

Updates

Submission history