Inspiration

Nigeria is the largest cassava producer in the world, over 60 million metric tons a year. And we get about 8 tons per hectare. Brazil gets 35.

Cassava fed my family growing up. My mother was a farmer and a teacher. So the gap between what the crop could yield and what it does yield is not abstract to me.

Before I built anything I spoke to three people. An agricultural extension officer at the Ebonyi State ministry told me farmers post their sick crops into WhatsApp groups full of strangers, because that is the only expert they can reach. She was also the one who told me agriculture was too broad and I needed to pick one crop. A production manager at Green Hills in Edo State, who grow cassava for ethanol, told me their two goals are yield and starch content, and their biggest problem is weeds, not disease. A PhD agricultural researcher who has worked with over 400 farmers across Nigeria and Ghana told me most of them cannot read or write English.

The obvious answer is an AI assistant. That fails here. It needs a connection the farmer does not have, data they pay for by the megabyte, and a subscription in dollars.

What it does

An agricultural extension officer that sits on a laptop and needs no internet.

A farmer describes what they see. The system tells them the most likely problem, why, what to check, and what to do — and which trusted guide the answer came from. It covers cassava diseases and pests, weed management, land preparation and planting, improved variety selection, and yield and starch improvement.

It runs on a 2015 MacBook Air with no GPU, offline, at zero cost per question.

How I built it

A 1.1 billion parameter Llama 3.2 quantised to Q4_K_M, running through llama.cpp. Retrieval over 688 passages from IITA, FAO, NAERLS and ASHC guides, embedded offline with all-MiniLM-L6-v2 in GGUF so there is no Python ML stack at inference. A FastAPI web interface with streaming responses and source attribution, and no external assets of any kind.

Two pieces are less obvious. A hand-written vocabulary map bridges how farmers describe symptoms to how agronomists write about them, including Igbo terms, because a guide says "bunchy top" where a farmer says "the top is bunching." And a condensed knowledge digest is embedded directly in the model's GGUF chat template, so the model answers cassava questions correctly even loaded bare in LM Studio with no application around it.

Challenges

A model this small invents things. At one point it recommended copper fungicide for cassava mosaic disease, which is a virus. None of the retrieved passages mentioned any chemical — it invented all of them because it had been asked what to do and had nothing to say. A farmer acting on that spends money they do not have and still loses the crop.

So the system assumes the model will fail. A relevance gate refuses questions outside the corpus before the model is even called. Question type is classified in Python, not by the prompt, because the model could not do it reliably. And any chemical recommendation not present in the retrieved sources is deleted before display, not merely flagged.

Retrieval was harder than generation. A question about leaves "curling near the top" retrieved nothing useful, because the guides say "bunchy tops." Fixing that took a vocabulary bridge. A question asking what chemical to spray retrieved chemical-dense passages regardless of what those chemicals treated, which meant rewriting the corpus so every control measure names its target explicitly.

My own hardware fought me. PyTorch dropped Intel Mac support at 2.2.2, so sentence-transformers was unavailable. That forced embedding through llama.cpp, which made the system lighter. And uploading an 800 MB model from my connection projected 20 hours, so the model is built and pushed from a Kaggle notebook instead.

What I learned

Measure on your own hardware. Q4_K_M ran at 18 tokens per second while the smaller Q3_K_L ran at 11, because 4-bit kernels are better optimised on CPU. The smallest file was not the fastest.

Write your corpus in the shape of the question. A list of variety yields ranked 71st for "which variety gives the highest yield." The same facts written as a sentence ranked in the top four.

Fine-tuning is not always the answer. I trained a LoRA on 206 curated examples. It converged cleanly and was worse: it garbled the yields, invented a crop that does not exist, and answered an out-of-scope question with fabricated veterinary advice despite a refusal for that exact question being in its training data. At this data scale, knowledge present at inference time beat knowledge diffused into weights.

Almost every improvement came from cheap, transparent, deterministic components rather than from swapping models. I tested four models and rejected three.

What's next

Weed management, soil correction, voice input and audio output in Igbo, photo diagnosis of weeds and soil, and packaging that a farmer can install without a terminal. Every item came from those three conversations.

Built With

  • agriculture
  • css
  • edge-ai
  • fastapi
  • gguf
  • html
  • huggingface
  • igbo
  • javascript
  • kaggle
  • llama-3.2
  • llama.cpp
  • lora
  • nigeria
  • numpy
  • offline-first
  • peft
  • python
  • quantization
  • rag
  • sentence-embeddings
  • server-sent-events
Share this project:

Updates

Submission history