Inspiration
Smallholder farmers and traders across Africa need reliable market-price answers, but that data is scattered, often offline-inaccessible, and cloud AI assistants require exactly the connectivity and electricity that's least reliable where this would help most. We wanted something that runs on the machine already on the desk, no cloud, no GPU, 8GB of RAM.
What it does
FarmGate is an offline assistant fine-tuned on real WFP VAM African market price data. It compares prices across markets and, just as importantly, knows what it doesn't know: when asked for a specific historical price it can't reliably recall, it says so plainly instead of inventing one.
How we built it
We bake-off tested four candidate models on identical hardware before choosing Qwen3 1.7B, built a 1,200-example fine-tuning set where every gold answer is computed by verifiable arithmetic over WFP's price data (not LLM-judged), and trained via QLoRA on Apple Silicon using MLX-LM after finding bitsandbytes' MPS backend unreliable in direct testing.
Challenges we ran into
Our first trained model looked good by loss curve but scored 0% on direct price recall, confident fabrication, not near-misses. We diagnosed the real cause: two training families sharing an identical prompt shape were teaching the model contradictory lessons (answer confidently vs. refuse). Relabeling one family to a calibrated, honest hedge fixed both at once, accuracy on that capability went from 0% to 100%, with zero new data. We also found and fixed a separate, more dangerous issue during cross-platform testing: a default context-window setting that could push memory use close to the competition's disqualifying 7GB ceiling on some runtimes, patched before submission.
What we learned
The most valuable finding was recognizing that "the model won't tell me the price" and "the model was never taught it's allowed to say that" are different problems with different fixes and it wasn't a bigger model or more training. Loss curves hid this; only generation-based, per-behavior testing surfaced it.
What's next
Extending genuine trend-direction reasoning (currently an honestly-acknowledged weak spot, not hidden) and exploring retrieval-based grounding for a future version, should the evaluation format ever support an application layer beyond the bare model.
Log in or sign up for Devpost to join the conversation.