Inspiration LLM serving on a single GPU can become inefficient when a long request sits at the front of the queue and forces many short requests to wait behind it. We wanted to explore whether we could predict how expensive a request would be before generation and use that information to schedule inference more intelligently. What it does InferQueue is an adaptive LLM inference scheduler for constrained local GPUs. It predicts the expected output length of each prompt, then uses that estimate to decide which request should run next. We compare three policies:
- FIFO
- Shortest Estimated Job First
- An adaptive fairness-aware policy that prioritizes shorter jobs but promotes requests that have waited too long The system runs real Llama 3.1 8B inference on an RTX 3060 Ti, records latency and queueing behavior, stores telemetry in Tiger Data, and visualizes the results in a live dashboard. How we built it We trained a DistilBERT-based predictor to estimate output length before inference. Training labels came from repeated real Llama 3.1 8B generations, and the predictor outputs both an expected token count and uncertainty. The scheduler is built around interchangeable predictor and backend interfaces, so the scheduling logic is independent of the ML model or inference runtime. We implemented FIFO, shortest-estimated-job-first, and an adaptive policy with starvation protection. For inference, we used llama.cpp with a quantized Llama 3.1 8B model on an RTX 3060 Ti. We also built an event and metrics pipeline for queue wait, service time, end-to-end latency, throughput, and starvation. Tiger Data stores scheduler events, while a Next.js dashboard shows queue reordering, execution timelines, and policy comparisons. Challenges we ran into One of the biggest challenges was collecting reliable training labels. We discovered that llama.cpp prompt-cache reuse affected repeated-generation reproducibility, which made our planned extension pass for truncated outputs invalid. We preserved the original measurements, flagged censored examples, and avoided mixing incompatible reruns into the dataset. We also found that fairness thresholds behave very differently depending on system load. A fixed 15-second maximum wait worked under lighter traffic but became too aggressive under heavy load. We changed the adaptive policy so the real threshold is derived from the model's measured service time instead. Another challenge was working under a tight deadline while the GPU was occupied generating labels. We split development across two machines so the scheduler, simulator, dashboard, and Tiger Data integration could be built in parallel with the ML pipeline. Accomplishments that we're proud of We built the entire system end to end rather than stopping at a scheduling simulation. Our DistilBERT predictor achieved lower test error than the simple prompt-length baseline while adding only a few milliseconds of overhead compared with multi-second Llama inference. We also built a scheduler that clearly demonstrates the tradeoff between latency and fairness: shortest-job-first improves responsiveness for short requests but can starve long ones, while our adaptive policy bounds that waiting time. Finally, the system includes real Llama inference, automated testing, persistent telemetry, and a dashboard that clearly separates simulated, mock, and real results. What we learned We learned that inference scheduling is not mainly about increasing raw throughput. With a single non-preemptive GPU worker, reordering requests mostly changes who waits and for how long. We also learned that prediction quality matters differently depending on the scheduler. Underpredicting a long request can be more harmful than overpredicting one because it can place an expensive request ahead of many short jobs. Most importantly, we learned how tightly connected ML serving is to systems engineering. Model prediction, queueing policy, backend behavior, caching, measurement, and fairness all affect the final result. What's next for InferQueue The next step would be to move beyond a single active request and explore batching and controlled concurrency, where scheduling decisions can affect both latency and GPU utilization. We would also like to collect a larger clean training dataset with prompt caching disabled, improve the cost model beyond output length alone, and incorporate prompt-prefill cost and predictor uncertainty directly into scheduling decisions. Longer term, InferQueue could evolve into a lightweight serving layer for self-hosted LLM deployments where latency, fairness, and limited GPU capacity all matter.
Built With
- cuda
- distilbert
- hugging-face
- llama-3.1
- llama.cpp
- llm
- machine-learning
- next.js
- postgresql
- python
- pytorch
- react
- tiger-data
- typescript
Log in or sign up for Devpost to join the conversation.