Track: Cloud AI. Autonomous LLM inference optimization on AWS Graviton4.
Inspiration
Optimizing LLM inference on Arm is expert work. Most developers ship a naive llama.cpp build on Graviton and leave 2x on the table, because knowing which build flags, quant formats, microkernels, and runtime knobs matter for a specific core is a specialist skill. Arm just shipped Performix, which finally exposes the performance data, but reading the counters and deciding what to change next is still entirely manual. We wanted an agent that runs that loop autonomously and shows its reasoning at every step.
What it does
armsmith is a reproducible Graviton LLM-optimization lab:
- A deterministic autotuner sweeps a fixed 5-lever action space on llama.cpp: build flags, KleidiAI microkernels, quant format, threading/affinity, and KV-cache type.
- A Claude analyst reads Arm Performix counter data, prioritizes the search order, and narrates every keep/revert. The LLM proposes; a registry-validated gate decides. The model never emits shell, and off-schema suggestions are dropped before execution.
- A quality guard measures KL divergence vs the baseline and perplexity for every candidate, so a config can never "win" by quantizing the model into mush.
- A rigorous benchmark harness: warmup discarded, N>=5 repeats, median with 95% CI, decode and prefill reported separately. Only CI-significant wins count.
- The winning config is saved as a recipe that
armsmith reproreplays with no LLM in the loop, alongside a self-contained HTML report and a JSONL trajectory of every decision and its evidence.
Result on Graviton4 (r8g, Qwen2.5-7B Q4_0, pp512/tg128, N=7): decode went from 21.6 to 44.7 tok/s (2.1x, non-overlapping 95% CIs) and prefill TTFT from 17.6s to 3.8s, with KL vs baseline of 0.002. The autotuner beat our pre-registered expert hand-tuned config (40.1 tok/s), closing 125% of the naive-to-expert gap. The recipe then replayed on a brand-new instance at 43.9 tok/s — 2% drift, no LLM in the loop. The full run (recipe, trajectory, report, repro result) is committed in the repo under trajectories/.
How we built it
A Python 3.11 control plane on the dev machine drives everything over SSH; all measurement happens on a real Graviton4 target (r8g.4xlarge). llama-bench does the measuring, Arm Performix (via the armlimited/arm-mcp MCP server) supplies the counter story, and a pluggable Brain protocol supplies the analyst: AnthropicBrain (Claude) for narrated runs, NullBrain as the no-LLM control. The optimizer core is deterministic search; the LLM earns its place as diagnostician and narrator, never as the control loop. 194 tests with real parsing fixtures, CI on both x86 and arm64.
Challenges we ran into
- Virtualized Graviton hides most counters: only 2 PMU counters and no SPE, so Arm's microarchitecture recipes hard-fail on EC2. We reframed the loop as try-measure-keep-best, with Performix hotspots as narration rather than the decision signal.
- Real hardware bites: no cpufreq node on EC2, IMDSv2 token dance, llama-bench aborting on
-t 0, engine-illegal KV-cache combos, and a llama.cpp webui asset download that rotted between two runs ten days apart and broke otherwise-identical builds. Each one became a regression test. - Keeping the LLM honest: the brain's prior said "KleidiAI first." The measurement said no: slower decode on this core/model pairing — in both the smoke run and the full submission run. The agent reverted on the evidence and pivoted to the native build. The narrated run converged on the same winner as the no-LLM control run, which is exactly the honesty property we wanted.
Accomplishments that we're proud of
- The autotuner beat the expert config we pre-registered before the discovery run, so the ">=90% of the gap" success metric could not be gamed after the fact.
- The recipe replayed on a bare fresh instance within 2% of the recorded result —
armsmith repro <run-id>from a clean clone, no LLM involved. - The safety gate held against real LLM output: off-schema quant suggestions and a raw hex
cpu_maskwere rejected at the validator, never executed. - Every reported number is CI-significant, measured on real Arm silicon, and committed with its full evidence trail.
What we learned
- On Graviton4, the single biggest lever for llama.cpp decode is building with
-mcpu=native(GGML_NATIVE), worth 2.1x on its own; KleidiAI is workload-dependent and measurement decides, not the datasheet. - An LLM in an optimization loop is most valuable as an analyst and narrator with a hard gate in front of it. Determinism is what makes the result publishable.
- Honest baselines matter: our baseline is a portable build a developer would actually ship, pinned in the run manifest, never a crippled strawman.
What's next
- A second Arm core (Graviton3 c7g) to show the recipes adapt per-core.
- PyPI release so
pip install armsmithworks out of the box. - Bare-metal Graviton, where the full Performix PMU/SPE counter set lights up the diagnosis half of the analyst.
Built With
- amazon-ec2
- arm-performix
- aws-graviton
- claude
- kleidiai
- llama.cpp
- mcp
- python
- ssh
Log in or sign up for Devpost to join the conversation.