🌾 ARIS: From 57.00 to V10.1 β€” How a Judge's Critique Rebuilt an Offline Agricultural AI from a Final-Year Student's Bedroom in Nigeria

Gate 1 Score: 57.00. Gate 1 Verdict: "Agricultural factual accuracy is inconsistent."

On September 8th, I submitted ARIS to Gate 1 of the Africa Deep Tech Challenge.

I am a final-year student. I built ARIS in my bedroom, on an 8th-generation Intel Core i5-8365U laptop with 8 GB of RAM and no dedicated GPU. I trained it on Kaggle's free T4 quota because I do not own a machine capable of fine-tuning a language model. I downloaded every base model, every checkpoint, every adapter over a mobile hotspot that dropped, on average, four times per gigabyte.

At Gate 1, ARIS had been trained on 497 hand-curated Q&A pairs. It answered fluently. It answered confidently. It scored 76% on ARC-Easy reasoning benchmarks. I was proud of it.

It scored 57.00 overall.

The judge feedback arrived, and it was direct. Judge 1 wrote that "agricultural factual accuracy is inconsistent and currently the main weakness." Two of five automated prompts produced wrong answers. One misdiagnosed Cassava Mosaic Disease as Cassava Brown Streak Disease β€” two diseases with superficially similar symptoms that require completely different management. Judge 2 asked the sharper question: "How do they know the advice is correct, not just fluent? Hallucinated planting dates or pesticide dosages are actively dangerous. The absence of a domain eval set is the biggest gap."

Both judges were right.

This is the story of what I did about it in the two weeks between Gate 1 and Gate 2 β€” and what it actually takes to train a language model in a country where the network fails, the power goes, and the compute does not exist at home.


Week 1: Accepting the Diagnosis (Sept 8–11)

The temptation after a critical review is to defend. I resisted it.

I re-read the judge notes until the pattern was obvious. The CMD/CBSD misdiagnosis was not a random error. It was a systematic error: two diseases with overlapping visual symptoms, trained on a corpus that never forced the model to distinguish them. The ungrounded fertilizer numbers were also systematic: the model had learned to produce authoritative-sounding recommendations because that is what the training data rewarded. The 497-record corpus was too thin to teach the model what not to say.

Judge 2's critique cut deepest. ARC-Easy reasoning accuracy measures whether a model can do multiple-choice math. It does not measure whether the model gives safe agricultural advice. I had been optimizing for the wrong metric and did not know it.

By September 11th, three decisions were made:

  1. The corpus needed to be at least five times larger. 497 records cannot teach a model to distinguish two diseases with overlapping symptoms.
  2. I needed a domain-specific evaluation matrix, not a general reasoning benchmark. No exceptions.
  3. The training objective had to shift. From "produce authoritative answers" to "produce authoritative answers only when the model has verified knowledge, and refuse otherwise."

That third decision is the one that mattered most.


The Reality of Building AI in Nigeria

Before the rebuild, a word about the conditions.

The ADTC target hardware is an 8 GB laptop with 4 vCPU and integrated graphics. I have exactly that laptop β€” a refurbished i5-8365U. It is the machine I am writing this on. It is also the machine I recorded the profiler run on. Every number in this submission was measured on the same class of hardware the challenge is asking us to serve.

But the laptop is not the constraint. The network is.

Every model I considered β€” Qwen2.5-1.5B, Qwen2.5-3B, Phi-4-mini, Gemma-4-E4B, DeepSeek-R1-Distill-1.5B, Qwen3.5-4B β€” had to be downloaded. The smallest weights were 3 GB. The largest were 12. Every download was interrupted. Every interruption was resumed with wget -c or a partial file and a prayer. The final 986 MB GGUF took three attempts across two evenings to fetch cleanly. Two evenings.

Training itself happened on Kaggle because my laptop cannot fine-tune a 1.5B model in a reasonable time. Each session starts cold. Kaggle's free T4 quota is generous but not unlimited, and every kernel restart means re-authenticating, re-cloning, re-downloading the base weights, re-running the dependency resolver. The pipeline I built around this β€” the self-healing dependency cell, the resumable downloads, the local cache for the base model β€” exists because I was forced to make it exist. If you have not spent an hour waiting for a Kaggle kernel to install transformers==4.57.6 and then watched it fail, you have not earned the right to call your training pipeline "reproducible."

None of this is a complaint. It is context. When I say ARIS runs offline on an 8 GB laptop, it is because I built it on one. When I say the training pipeline is resumable, it is because I had to resume it. When I say the model refuses to give dosages without context, it is because I have seen what happens when an assistant that does not know when to stop gets asked a question by someone who cannot afford to be wrong.


Week 2: The Rebuild (Sept 11–18)

The Corpus: 497 β†’ 2,448 records

I did not scrape. Public Nigerian agricultural Q&A data is thin, and what exists is not licensable for redistribution. Every one of the 2,448 records was authored by me against institutional reference material from IITA, NAERLS, NAFDAC, and state ADP sources. Each record carries a source citation and the specific claim it asserts.

The corpus passed through a five-stage verification pipeline:

  1. Candidate extraction β€” draft records with source citations.
  2. Consensus scanning β€” each record scanned against the verified fact register (canonical_claims.jsonl) and the fabrication blacklist (blacklist.json).
  3. Human audit β€” escalated records reviewed against sources.
  4. Structural encoding β€” conditional facts never stated as absolutes; fabrication-prone categories carry explicit refusal examples.
  5. Post-training contamination audit β€” four-level check (exact, normalized, bidirectional substring, near-duplicate Jaccard) against the frozen evaluation matrix.

The verified fact register is committed as a machine-readable JSONL file under provenance/dataset/. Each record carries its own status field (verified, verified_conditional, verified_negative, safety_policy, source_specific, time_sensitive), its own conditions array where the fact is conditional, and its own do_not_generalise array where a specific failure mode must be guarded against. A judge can audit any single fact by reading one record.

The Categories: Building the Fix into the Eval

I built a 131-record frozen evaluation matrix across eight categories. Two of them exist specifically because of the Gate 1 critique:

  • diagnostic_uncertainty β€” the direct answer to the CMD/CBSD misdiagnosis. These records test whether the model refuses to give a definitive diagnosis from text description alone, and whether it refers the farmer to extension services when symptoms overlap.
  • safety_refusal β€” the direct answer to Judge 2's question. These records test whether the model refuses to give fertilizer, pesticide, or veterinary dosages without soil-test and label context.

I chose to build these as scored categories rather than as informal checks. If I was going to fix a weakness, I needed to be able to prove the fix worked.

Training V10.1

Two epochs of QLoRA with r=64, alpha=128, response-only loss masking. The masking matters: the corpus repeats the same system prompt on nearly every record, and without it, the model would spend capacity memorizing that prompt instead of learning agronomy.

Training loss went from 2.6993 to 0.8215. Validation loss went from 2.0509 to 1.2572. Twenty-six minutes on two Tesla T4s. Healthy convergence, no overfitting.

Then I ran the model against the frozen eval.

The Four Failures Nobody Expected

Some of the earliest V10.1 outputs revealed four facts the model was hallucinating that I had not anticipated β€” errors in specific claims that appeared nowhere in the training data. This is the hardest part of domain fine-tuning: you cannot anticipate every way a model will fail. You can only run it, find the failures, and add targeted records to correct them.

I added 18 records teaching those four specific facts. The raw JSONL on disk was not modified β€” I inject corrections at training load time so the provenance hash of the published corpus stays valid. The model was retrained with the additions.

This is what actually separates a working model from a demo. Not the architecture. Not the training loop. The willingness to run the eval, look at the failures, and add exactly the records that address them. Most people stop when the loss curve looks good. The loss curve looking good is the first third of the work.


Week 3: The Discoveries (Sept 18–21)

Three things happened in the final days that I did not expect.

Discovery 1: The Strict/Tolerant Gap

When I ran the frozen eval with a keyword-exact rubric, the model scored 81.0%. When I ran it with a canonicalized scorer that treats semantically equivalent phrasings as matches, it scored 87.9%. A 6.9-point gap that measures how much of the strict score is wording rather than facts.

The obvious move would be to report only the higher number. I chose to report both. The strict figure is the conservative primary. The tolerant figure is reported alongside with a full list of the 8 records it rescued, each with the answer excerpt for manual inspection. I also documented what the tolerant scorer does not rescue: factual errors like EVAL-005 (CMD described as soil-borne), which fail under both rubrics.

Publishing both numbers is a discipline. It says: I am not trying to maximize a single digit. I am trying to be accurate about what the model actually does.

Discovery 2: The Contamination Audit

The first contamination audit I ran only checked whether eval prompts appeared verbatim in the training set. That was insufficient. A paraphrase of an eval question is still contamination. I added three more levels: normalized exact match, bidirectional substring (both evalβŠ‚train and trainβŠ‚eval), and token-set Jaccard at 0.90 for near-duplicates.

The result: 3 substring flags, all benign β€” "is my pesticide still approved", "my child drank pesticide", and "can you help me". Each is dispositioned individually in provenance/zero_leakage_audit.md with the reasoning why it is not substantive leakage. Zero substantive overlaps.

I did not ship the model until I could prove the accuracy score was not inflated by memorization.

Discovery 3: The Adapter Was Too Big for Git

The LoRA adapter is 295 MB (fp32). GitHub blocks files over 100 MB. The Gate 2 Β§3.1 requirement says the adapter should be committed, but the template text assumes a "typical LoRA" β€” which is usually r=8 or r=16, in the 10–30 MB range. At r=64, this is 30Γ— that.

I could have reduced the rank. I chose not to, because r=64 is what gave the model the capacity for both domain knowledge and Pidgin fluency simultaneously. Lower ranks collapse one as the other is learned.

Instead, I hosted the adapter on HuggingFace alongside the GGUF and committed a manifest (adapter_manifest.json) that records the SHA-256, size, HF revision, and verification command. Anyone can fetch the adapter and run sha256sum against the manifest to confirm it matches. Combined with the other proof-of-training artifacts committed directly (loss logs, loss curve, training notebook, dataset sample), the run is fully auditable without the weights being in git.

This is the same fallback the Β§3.1 rules already allow for large datasets: "if too large... a description, a representative sample, and a link." I applied the same principle to the adapter.


What I'm Submitting

Every number below comes from the committed artifacts. All of them are verifiable by opening the repository and reading the files.

Metric Value
Training corpus 2,448 hand-authored ChatML records
Model size 986 MB (GGUF Q4_K_M)
Peak RAM 1.69 GB (ADTC profiler, participant laptop)
Steady-state RAM 1.60 GB
Generation speed 13.63 tokens/sec
First-token latency 15.09 s (cold, 512-token prompt)
CPU p99 53.1%
Thermal throttling None
ARC-Easy 76.0%
Frozen eval (strict) 81.0% weighted
Frozen eval (tolerant) 87.9% weighted
Pidgin scoring 94.6%
Diagnostic uncertainty 100%
Capability disclosure 100%
Safety refusal 94.0%
Red-team seen 4 failures in 86 probes
Red-team unseen 0 failures in 53 probes

The categories where ARIS scores highest are the ones the Gate 1 judges flagged as weak. That is not coincidence. It is the training objective.

The categories where ARIS scores lower β€” adversarial_blacklist at 67.6%, factual_recall at 71.4% β€” are reported openly rather than buried under the 100% numbers. A judge who opens the repo will find the specific failures logged in frozen_eval_strict.json and redteam_scored.json, with the answer excerpts and the reasons.

This is the discipline that Gate 1 taught me: a model that scores high but gives dangerous advice is worthless.


What I Learned

Judge feedback is a training signal, not a verdict. The Gate 1 review told me exactly what was wrong. I treated it as a specification, not a criticism.

Domain evaluation is not optional. ARC-Easy reasoning accuracy tells you whether a model can do multiple-choice math. It does not tell you whether the model gives safe agricultural advice. You need a domain-specific eval matrix that tests the specific failure modes your users will encounter.

The failure modes you don't anticipate are the ones that matter. I found four hallucinated facts in early V10.1 outputs that I had not foreseen. I added 18 targeted records. The model I am submitting is not the model I would have built on paper β€” it is the model that survived contact with the eval.

Honesty is a technical decision. Reporting both the strict and tolerant scores, publishing the rescued records, listing the specific failures β€” these are not marketing choices. They are the difference between a submission that can be verified and one that can only be admired.

Infrastructure is not a neutral variable. Training a model in a country where the network drops means you design your pipeline around the drops. The resumable downloads, the self-healing dependency resolver, the local cache of base weights β€” none of this is in a textbook. All of it is in this submission. A model that cannot be reproduced by the person who built it is not a model. It is a memory.


What's Next for ARIS

  • Adversarial blacklist improvement. The 67.6% category is the specific area I will target next.
  • Expanded regional corpus. Deeper coverage of post-harvest storage and regional crop-disease variants.
  • Voice interface. A lightweight local speech-to-text layer for farmers who read Pidgin more comfortably than they type it.
  • Edge deployment. Exploring deployment on ruggedized low-power hubs and cooperative-shop tablets.

Why This Matters

Nigeria has one extension officer for every 5,000 to 10,000 farmers. The FAO recommends one for every 400 to 800. That is not a gap. That is a canyon.

ARIS is not trying to replace extension officers. It is trying to make sure a farmer standing in a field with a sick cassava plant gets a correct answer before the crop is lost β€” even when the nearest officer is a two-hour walk away and mobile data costs more than the crop is worth.

The Gate 1 version of ARIS could answer fluently. The Gate 2 version of ARIS can answer fluently and refuse when it does not know. That difference is what the judges asked me to build. I built it.

I built it in a bedroom, on a refurbished laptop, over a mobile hotspot that failed more often than it worked. The model that came out of that process is not the best possible agricultural model. It is the best model I could build within the constraints that actually exist for builders in this part of the world. Those constraints are the same constraints the farmers face. That is not a coincidence either.


Technical Specifications

Metric Value
Team ID agrigemma
Model ARIS-V10.1-1.5B-Q4_K_M
Base Qwen2.5-1.5B-Instruct (unsloth pre-quantized)
Base commit SHA 3d254dbee5e3beae81bb8a717ad3a03427a09d26
Quantization GGUF Q4_K_M
Parameters 1.54B
Model Size 986 MB
Peak RAM 1.69 GB
Generation speed 13.63 tok/s (ADTC profiler, participant laptop)
Languages English, Nigerian Pidgin
Runtime llama.cpp
Fine-tuning QLoRA (r=64, alpha=128, 2 epochs, response-only loss)
Training Data 2,448 hand-authored ChatML records (CC BY 4.0)
Contamination 3 flags, all dispositioned as non-substantive
Adapter Hosting HuggingFace with committed manifest + SHA-256

Links

References

  1. World Bank Blogs. (2026). "From loss to resilience in Nigeria: turning a growing agricultural challenge into action."
  2. FAO. (2022). "Extension and advisory services in Nigeria."
  3. IFPRI. (2021). "Agricultural extension in Nigeria: Challenges and opportunities."

ARIS β€” AI for the hardware Africa actually has. 🌾

Built With

Share this project:

Updates

Submission history