Inspiration

AI infrastructure has become very good at making models fast.

Then we started looking at agents.

An agent does not just send one prompt and wait for an answer. It asks the model something, calls a tool, fetches data, parses a response, updates its state, maybe waits on a network request, then finally goes back to the model. Do that across dozens of agents and something strange happens: you can have a fast inference server sitting there while the rest of the system is still trying to prepare its next useful request.

That felt backwards.

A lot of performance work focuses on what happens inside the model server. That work matters. But engines such as vLLM already put serious effort into serving many requests efficiently once those requests arrive. The question that interested us was what happens before that point.

What if the model is ready, but the agent is not?

That became AgentFeeder.

The idea is simple enough to say in one line:

We did not make the model faster. We stopped making it wait.

What it does

AgentFeeder sits between an agent workflow and its inference server.

Instead of treating every agent as one long sequence of waiting, AgentFeeder breaks each session into explicit pieces of work: model inference, tool calls, CPU work, validation and state updates.

Then it asks a very practical question:

What useful work is ready right now?

If one agent is waiting for logs from a tool, another agent can move forward with inference. If another session finishes parsing data, its next model request can become ready immediately. Tool work, CPU work and inference each get their own bounded queues, so one slow resource does not blindly flood everything else.

We also keep the boring parts that become very important once concurrency gets real: fairness, backpressure, cancellation, deadlines, retries and idempotency.

AgentFeeder is not trying to replace vLLM or rewrite model scheduling. The boundary is different. The model server decides how to serve requests that have already arrived. AgentFeeder works on the host side to make sure useful requests become ready in the first place.

The main metric follows the same philosophy. We care about successful agent tasks per minute, not an impressive tokens-per-second number that says nothing about whether the agent actually finished its job.

How we built it

We designed AgentFeeder around explicit state machines instead of hiding the workflow inside a pile of asynchronous callbacks.

Every agent session moves through a graph of work such as:

INFER → TOOL_IO → CPU_TASK → INFER → VALIDATE → COMPLETE

Ready work goes into separate bounded queues for inference, tools and CPU tasks. A resource-aware scheduler decides what can run, advances sessions as work finishes and keeps other sessions moving when one of them blocks.

For the main benchmark, we use a repeatable multi-agent incident-response workload. An agent classifies an incident, retrieves logs, parses them, diagnoses the problem, fetches a runbook, chooses an action and passes the result through a validator. It gives us model work, network-style waiting and real CPU-side processing in the same flow.

The Arm side is structural, not a logo on the README. The target runtime is Linux on Arm64 cloud hardware. We profile the host path, look for actual CPU bottlenecks, change one thing at a time, then run the same workload again. That can include parsing and serialization costs, memory movement, connection reuse, queue contention, worker topology and CPU affinity.

And we keep receipts.

The baseline and optimized runs use the same model, prompts, tool fixtures, seeds, workload, validation rules and machine configuration. Benchmark results are stored as machine-readable evidence so the charts can be rebuilt from the raw runs instead of typed into a dashboard.

No magic numbers.

If an optimization does not help, we want the benchmark to tell us that too.

Challenges we ran into

The hardest problem was defining what AgentFeeder should not claim.

Agent scheduling already exists. Inference scheduling already exists. Async runtimes obviously exist. Calling AgentFeeder the first system to do any of those things would make the project sound bigger for about thirty seconds and much weaker the moment somebody technical opened the repository.

So we had to draw the boundary carefully.

AgentFeeder focuses on the host-side gap between one model turn finishing and the next useful inference request becoming ready.

Measurement creates another trap. It is very easy to look at gaps between requests and call them "GPU idle time." That sounds great in a pitch, but unless we have direct accelerator telemetry, we do not actually know that. So we call the metric what it really is: model-request supply gap.

Concurrency was the other uncomfortable part. Making more things run at once is easy. Making them run at once without duplicate tool actions, broken dependencies, starvation, bad retries or sessions mysteriously completing after cancellation is much harder.

That forced us to treat correctness as part of performance.

A faster wrong agent is still wrong.

Accomplishments that we're proud of

The thing we are most proud of is that AgentFeeder has a falsifiable idea behind it.

We are not starting with a dashboard and working backwards toward a benchmark that makes it look good. The comparison is defined first. Same hardware. Same workload. Same model. Same validation. Then baseline versus AgentFeeder.

We also resisted turning it into another giant agent framework.

The useful artifact is much smaller: a host-plane scheduler, a repeatable workload harness, structured traces and an evidence format developers can inspect for themselves.

That matters because another developer should be able to look past our demo and ask, "Would this scheduling idea help my own agent workload?"

They should not have to trust us.

They should be able to run it.

And perhaps our favorite part is the visualization itself. Instead of showing a vague "AI optimized" badge, AgentFeeder can show the actual execution timeline. You can see where sessions wait, where useful work becomes ready, and what changes when the scheduler starts filling those gaps.

The performance story becomes visible.

What we learned

We started with a very normal assumption: if you want faster AI, make inference faster.

Agent systems made that assumption feel incomplete.

Once a workflow starts calling tools, querying services, parsing data and bouncing between model turns, the speed of the model becomes only one piece of the final response time. A fast engine cannot serve a request that the rest of the application has not prepared yet.

We also learned that concurrency and throughput are not the same thing.

You can spawn hundreds of tasks and create a beautiful mess. More workers can increase contention. More in-flight requests can destroy tail latency. Aggressive retries can accidentally execute a real-world action twice.

Good scheduling is mostly about knowing when not to run something.

The final lesson was about benchmarking.

One big percentage is tempting. It is also weak evidence by itself.

Repeated runs, raw receipts, correctness checks, hardware fingerprints and ablations are less glamorous, but they let us answer the question that actually matters:

What changed, and did that change really help?

What's next for AgentFeeder

The first goal is to make AgentFeeder excellent at one job: keeping the host side of multi-step agent workloads moving efficiently on Arm64 infrastructure.

From there, the interesting work starts.

We want to test scheduling policies that react to queue pressure and deadlines instead of using one fixed concurrency setting. We want to import traces from real agent applications so developers can replay their own workloads instead of relying only on our incident-response fixture. We also want to expand the model-server adapter layer so AgentFeeder can sit in front of different inference stacks without forcing teams to rebuild their agents around it.

There is a bigger idea here too.

As agents become longer-running and more tool-heavy, buying faster inference alone will not fix every delay. The CPU host becomes the place coordinating model requests, retrieval, tools, validation, state and data movement.

That is the part AgentFeeder wants to make better.

Not by adding more agents.

By feeding the ones we already have.

Built With

  • arm64
  • armperformix
  • asyncio
  • awsgraviton
  • docker
  • fastapi
  • hypothesis
  • inferenceoptimization
  • linuxperf
  • llamacpp
  • opentelemetry
  • postgresql
  • prometheus
  • pytest
  • python
  • redis
  • vllm
Share this project:

Updates

Submission history