Inspiration

It started with a maize field in Kara. The leaves went yellow three days after a heavy rain, and the grower didn't know if the rain had washed out the nitrogen or if a disease was starting. That question has a clear answer. It's printed in an ITRA extension sheet in an office in Lomé. But he had no officer to call, no data bundle, no signal. By the time an answer could reach him, the week to act was gone, and part of his harvest with it.

That's the thing that wouldn't leave me: the knowledge already exists. What's missing is knowledge that shows up in time, offline, in a form a villager can actually use. Not another app that needs the cloud, a smartphone and a good signal, three things the people who need this most rarely have at once. So I decided to build the opposite.

What it does

Agritgllm is an offline agricultural adviser that runs on a cheap 8 GB laptop with no internet and no graphics card. One file of 470 MB. Ask it about a crop, an animal, a season, a sick plant, or when to sell, and it answers with advice grounded in Togo's own agronomy: the right sowing window for the north versus the south, whether to hold your maize for the lean season, whether a small local-hen unit pays in its first year, what to check when your hens start dying suddenly.. It names the region, the crop and the season, because a tip that works for Senegal as well as Togo helps nobody.

And it knows where it stops. Ask it for a human antibiotic dose for your cattle and it refuses, names no disease from the signs, and sends you to the veterinary service today. That is the exact question our Round 1 model failed by playing doctor.

How we built it

I started from 60 public extension sheets, ITRA and ICAT on maize, rice, cassava, cowpea, poultry, small ruminants, plus national studies on markets, climate and inputs. Half of them turned out to be scans with no readable text, so I read them as images to save every dose, spacing and calendar by hand.

Those sheets became a knowledge base of 169 pages and 1,076 verified citations, where a page may contain nothing but sentences copied from a sheet, and a linter with no model in it checks every quote word for word against its source. From there, 521 fact bundles, each fact taught from six to nine angles, because a small model only repeats what it has seen in several wordings.

An open model writes the training conversations one bundle at a time; a deterministic checker and a judge from a different model family, which has to quote its proof, throw out 15.5% of what it writes. Then LoRA fine-tuning of LFM2-700M: two adapters, one for knowledge, one for the safety boundary, averaged 50/50 into the shipped one, merged into the base and quantized to Q4_K_M GGUF running through llama.cpp, the exact runtime the challenge scores.

Challenges we ran into

  • Silent documents. The most useful sheets were scanned images. A normal PDF pipeline handed me empty files. Recovering the numbers meant reading the pages visually, one by one.
  • French in, English out. The source knowledge is in French; the exam is in English. I had to ground English answers in French text and prove they were faithful without any shared words to lean on.
  • Honesty at 700M. Small models lie with confidence. Forcing every figure back to its source, and giving an independent model a veto, was the only way I'd trust what came out.
  • Safety is shallow. Teaching a refusal is easy; making it hold to the last word is not. Our model would refuse a dose, then give one forty words later. Every attempt to fix that with preference training broke something else, so we trained two models instead, one for knowledge and one for the boundary, and averaged them.

Accomplishments that we're proud of

The whole thing runs on an 8 GB laptop with no GPU, fully offline, at 17.07 tokens per second with 557 MB of peak memory, measured by the organisers' own profiler, in their Docker image, on a replica of the target machine, not on my development laptop.

But the number I'm proudest of isn't speed. On the 27 questions built from the Round 1 judging report, it scores 20.25 / 27 where the unmodified base scores 4.0; on 27 held-out questions about sheets it has never read, 16.75 / 27 against 2.5. And across 23 adversarial attempts to make it hand out a dose, it gave zero figures. Eight candidates were measured the same way this week; we shipped the one that held.

What we learned

My first instinct was to grab the biggest, smartest model I could fit. The numbers proved me wrong. Speed is only rewarded up to a ceiling:

$$S_{\text{perf}} = \min!\left(\frac{\text{TPS}}{15},\ 1\right) \times 100$$

Past \(15\) tokens per second, extra speed buys you nothing, while a bigger model quietly costs you memory and speed you can't recover. The real lever wasn't the model at all; it was the knowledge poured into it, and the discipline of refusing to let it learn anything we couldn't point to on a page. A fast, empty head loses to a small, well-taught one.

The second lesson was cheaper than it looks: when two abilities fight inside one model, stop training and start averaging. Two adapters, 0.5 and 0.5, no retraining, and the result beats both parents.

What's next for Agritgllm

Togo first, where it earns its keep on real crops and real seasons. Next steps: French answers, then a voice layer in Ewe so a farmer in the Savanes can hear the answer on a basic phone with no connection; wider livestock and market-price coverage; and putting it in the hands of a few extension officers to hear where it's still wrong. Then the same recipe travels, swap the corpus and keep the method, across West Africa and the wider Sahel, wherever good agronomy sits on paper and never reaches the field in time.

Built With

  • dpo
  • gguf
  • gpt-oss
  • groq
  • huggingface
  • lfm2
  • llama.cpp
  • lm-evaluation-harness
  • lora
  • offline
  • ollama
  • on-device-ai
  • peft
  • python
  • pytorch
  • q4-k-m
  • qwen
  • transformers
  • trl
  • unsloth
Share this project:

Updates

Submission history