Inspiration
I was working on a benchmark where I needed to repeatably run the same process over documents and wanted to script/orchestrate codex subagents, a la Claude workflows.
What it does
gpt-workflow lets you write multi-agent workflows in JavaScript. Your code controls loops, branches, retries, and parallelism, while Codex agents handle bounded judgment. Runs can produce structured output and can resume without repeating completed work.
How we built it
First, I had Claude generate an authoritative set of workflows.
Then, we spec'd out an implementation with the Codex App Server.
With a goal, I had it build it by delegating to Codex 5.6 Sol subagents for implementation. This got us 90% there.
I dogfooded it in my benchmark project and, using Codex, refined the package, the skills, created the plugin, published it, and created a Python wrapper.
Challenges we ran into
- There is no authoritative Claude workflows spec ā the initial set of replication scripts were insufficient
- Codex is too eager to put too much logic into the scripts, I had to update the skills to tell it to chill
Accomplishments
Using this, I was able to script/standardize my data preparation pipeline for a first-of-its-kind deterministic long-horizon medical evidence synthesis benchmark. Iām also proud that the result is a standalone CLI, library, and Codex plugin rather than a one-off script.
What's next
Launching gpt-workflow and the medical evidence synthesis benchmark =)
Built With
- bun
- codex
- python

Log in or sign up for Devpost to join the conversation.