neoverse-tune — workload-sensitivity for llama.cpp on Arm
Four users of one AI agent cost 10.9× time-to-first-token on Arm Neoverse N2.
One flag fixes it. We swept 6 prompt shapes to find exactly when it matters —
and when the flag does nothing at all.
Log in or sign up for Devpost to join the conversation.