Inspiration

What makes Claude Code good is not a clever instruction anywhere in it, it is scaffolding: what enters context, what is held back, what survives a compaction. That layer sits between the model and the world, and it is where I like working.

Programmatic tool calling is the sharpest recent example. The model writes one program that calls every tool instead of asking for them one at a time, and Anthropic measured complex research tasks falling from 43,588 tokens to 27,297. The saving is context. The clock is untouched. Inside that program the calls still run in written order, and none of them start until the model types the last line.

That gap is the whole opportunity, and on this platform it is enormous, because a single tool call here is not an API request. It is an entire agent turn: a container, a Codex session, its own model calls. A plan with several of them spends almost all of its time as a platform sitting still, waiting on work it already had everything it needed to start.

Eager PTC is the layer I built to close that gap. It sits between an agent and its tools, reads the program while the model is still writing it, and starts every call it can prove is safe to start.

What it does

The agent calls plan(goal), a tool that sits in its own tool list and is served over MCP rather than suggested in a prompt file. Everything after that is the layer's job, and from the outside it does four things.

It starts work before the plan exists. As the model writes, the layer is already reading. The moment a call is complete enough to be understood, and only if it can prove that call is safe, it launches it. The rest of the plan is still being typed.

It refuses, out loud. Every call gets one of three answers: speculate, defer or refuse. Each carries a reason and the line of source that caused it. A call that cannot be started early still runs, in order, when execution reaches it — refusal governs earliness, never whether a call happens.

It runs the plan at its real shape. Calls that do not depend on each other run together, regardless of the order they were written in, and anything held back acts as a barrier so nothing can reorder itself past a side effect.

It shows its work. Every run renders a waterfall where the pale band is the model still generating and each bar begins inside it. A Compare button re-runs the identical plan serially on the same axis, so the difference is a length you can see rather than a claim you are asked to accept.

One rule holds the design together: the model writes the plan, but nothing the model writes decides what is safe to run early. A plan that asks to be trusted gets exactly the same treatment as one that does not.

How I built it

TypeScript, inside the Starter Kit's existing Fastify control plane, in five pieces.

A registry gives every tool three independent flags: speculatable, sideEffectFree, deterministic. They answer three different questions, and a single "pure" flag would lose the difference between safe to run early and safe to reuse a result from.

An analyser built on acorn re-parses the growing plan after every token, classifying each argument as literal, purely computable, resolved from an earlier call, or tainted. Taint propagates transitively, and anything it cannot identify is treated as tainted rather than assumed safe. That last default is the one that matters: unknown means unsafe, so the failure mode of a parser gap is a missed speedup, never a call that should not have run.

An admission gate turns those classifications into the three answers, and a scheduler holds the dependency graph. The worker pool exists because the platform forces it: agent-service returns 409 on a second concurrent run of the same Agent, so parallel calls have to be leased out to distinct ones.

A promise store keyed on (tool, argsHash, occurrence) deduplicates identical calls, but only when the tool has declared itself deterministic.

plan() reaches the agent as a real MCP tool registered into Codex's tool list, not as an instruction in a prompt file, so an agent that ignores its instructions still gets the layer.

The boundary that holds all of it together: the analyser, the gate and the scheduler are pure functions of the source text and the registry. None of them read the model's prose. Owning the model call also means the layer can shape it, so the planner runs with reasoning disabled and emits its calls ahead of its rationale, which is what turns a theoretical early start into a usable one.

Challenges I ran into

The measurement almost buried the project. My first speedup metric summed the durations of awaited calls, which meant the number fell when speculation worked, because successful speculation removes time from the very sum being measured. The same workload read 2.32× under the broken metric and 5.03× once it counted work actually performed. I spent an afternoon convinced the idea did not pay.

That produced the rule the rest of the project follows: measure the two savings separately and never average them. Concurrency and early starting respond to completely different plan shapes, so a single headline number hides the case where one of them is doing all the work. Both now carry invariant tests, including one asserting that a call's elapsed time matches its measured work within 50 ms, which is what caught a second bug where start times were recorded at claim rather than at execution.

Side effects were the harder design problem. Running calls early is easy; the difficulty is proving you are allowed to. An if whose condition depends on a tool result cannot be resolved before that result exists, so anything inside it has to wait. Rather than trying to reason about which side effects commute, I made every defer and refuse a hard barrier in the scheduler. It gives up some speed on plans with an early unresolvable branch, and in exchange the ordering guarantee is something I can state in one sentence and test directly.

Accomplishments that I'm proud of

The result I am most pleased with is the least dramatic one. On a chain of three dependent calls, where concurrency can contribute literally nothing and the sequential and parallel baselines land within five milliseconds of each other, the layer still returns 1.28× — entirely from starting the first call before the plan had finished being written. That case isolates the mechanism completely: there is no parallelism available to borrow credit from.

The generation contract is the part I would defend hardest. Running the planner with reasoning disabled took the first usable token from 26.9 s to 0.43 s and the head start from 0.46 s to 8.82 s, and the plans got better rather than worse: protocol adherence went from zero valid calls out of three to three out of three.

The layer does not say yes to everything, which is easiest to see in the case where it says no. A notify call is registered as non-speculatable, because you cannot un-send a message, so the gate refuses to start it early and labels it with the reason, "tool notify is not speculatable", together with its argument class and the exact source position that produced it. It still runs when execution reaches it: refusal governs earliness, not whether a call happens. The rest of that same plan still came back 2.54× faster, which is the point. Being safe and being fast are not the trade they are usually assumed to be.

npm run check passes: typecheck, 119 tests, and both production builds.

What I learned

That the honest version of a performance result is two numbers rather than one. Running calls concurrently and starting them early are independent savings, they respond to completely different plan shapes, and averaging them into a single headline hides the case where one of them is doing all the work. So the interface reports both separately, and the README says plainly which one has preconditions.

And that much of an agent platform's latency is structural rather than computational. Nothing here made the model faster or the tools cheaper. It only stopped the platform waiting for permission it already had.

What's next for Eager PTC

Adaptive concurrency is unit-tested against injected faults but has never fired against a real provider throttle, so the next honest step is running it under one. After that: widening the plan grammar beyond its current restricted subset, and letting a tool declare its own cost, so the admission gate can weigh what a wasted speculation actually costs rather than treating every discarded call as equivalent.

Built With

Share this project:

Updates

Submission history