Searches complete model-routing combinations over a fixed twenty-case evaluation and selects the lowest-cost route that clears quality and coverage gates. I used Codex with GPT-5.6 to implement exhaustive search, deterministic scoring, tests, and the product experience. GPT-5.6 is represented as the high-capability baseline in the tested fixture. The public benchmark uses fixed synthetic scores and does not make live model calls.
Built With
- cloudflare-workers
- codex
- css
- gpt-5.6
- html
- javascript
- python
Log in or sign up for Devpost to join the conversation.