πΎ ARIS: From 57.00 to V10.1 β How a Judge's Critique Rebuilt an Offline Agricultural AI from a Final-Year Student's Bedroom in Nigeria
Gate 1 Score: 57.00. Gate 1 Verdict: "Agricultural factual accuracy is inconsistent."
On September 8th, I submitted ARIS to Gate 1 of the Africa Deep Tech Challenge.
I am a final-year student. I built ARIS in my bedroom, on an 8th-generation Intel Core i5-8365U laptop with 8 GB of RAM and no dedicated GPU. I trained it on Kaggle's free T4 quota because I do not own a machine capable of fine-tuning a language model. I downloaded every base model, every checkpoint, every adapter over a mobile hotspot that dropped, on average, four times per gigabyte.
At Gate 1, ARIS had been trained on 497 hand-curated Q&A pairs. It answered fluently. It answered confidently. It scored 76% on ARC-Easy reasoning benchmarks. I was proud of it.
It scored 57.00 overall.
The judge feedback arrived, and it was direct. Judge 1 wrote that "agricultural factual accuracy is inconsistent and currently the main weakness." Two of five automated prompts produced wrong answers. One misdiagnosed Cassava Mosaic Disease as Cassava Brown Streak Disease β two diseases with superficially similar symptoms that require completely different management. Judge 2 asked the sharper question: "How do they know the advice is correct, not just fluent? Hallucinated planting dates or pesticide dosages are actively dangerous. The absence of a domain eval set is the biggest gap."
Both judges were right.
This is the story of what I did about it in the two weeks between Gate 1 and Gate 2 β and what it actually takes to train a language model in a country where the network fails, the power goes, and the compute does not exist at home.
Week 1: Accepting the Diagnosis (Sept 8β11)
The temptation after a critical review is to defend. I resisted it.
I re-read the judge notes until the pattern was obvious. The CMD/CBSD misdiagnosis was not a random error. It was a systematic error: two diseases with overlapping visual symptoms, trained on a corpus that never forced the model to distinguish them. The ungrounded fertilizer numbers were also systematic: the model had learned to produce authoritative-sounding recommendations because that is what the training data rewarded. The 497-record corpus was too thin to teach the model what not to say.
Judge 2's critique cut deepest. ARC-Easy reasoning accuracy measures whether a model can do multiple-choice math. It does not measure whether the model gives safe agricultural advice. I had been optimizing for the wrong metric and did not know it.
By September 11th, three decisions were made:
- The corpus needed to be at least five times larger. 497 records cannot teach a model to distinguish two diseases with overlapping symptoms.
- I needed a domain-specific evaluation matrix, not a general reasoning benchmark. No exceptions.
- The training objective had to shift. From "produce authoritative answers" to "produce authoritative answers only when the model has verified knowledge, and refuse otherwise."
That third decision is the one that mattered most.
The Reality of Building AI in Nigeria
Before the rebuild, a word about the conditions.
The ADTC target hardware is an 8 GB laptop with 4 vCPU and integrated graphics. I have exactly that laptop β a refurbished i5-8365U. It is the machine I am writing this on. It is also the machine I recorded the profiler run on. Every number in this submission was measured on the same class of hardware the challenge is asking us to serve.
But the laptop is not the constraint. The network is.
Every model I considered β Qwen2.5-1.5B, Qwen2.5-3B, Phi-4-mini, Gemma-4-E4B,
DeepSeek-R1-Distill-1.5B, Qwen3.5-4B β had to be downloaded. The smallest
weights were 3 GB. The largest were 12. Every download was interrupted. Every
interruption was resumed with wget -c or a partial file and a prayer.
The final 986 MB GGUF took three attempts across two evenings to fetch
cleanly. Two evenings.
Training itself happened on Kaggle because my laptop cannot fine-tune a
1.5B model in a reasonable time. Each session starts cold. Kaggle's free T4
quota is generous but not unlimited, and every kernel restart means
re-authenticating, re-cloning, re-downloading the base weights, re-running
the dependency resolver. The pipeline I built around this β the self-healing
dependency cell, the resumable downloads, the local cache for the base
model β exists because I was forced to make it exist. If you have not spent
an hour waiting for a Kaggle kernel to install transformers==4.57.6 and
then watched it fail, you have not earned the right to call your training
pipeline "reproducible."
None of this is a complaint. It is context. When I say ARIS runs offline on an 8 GB laptop, it is because I built it on one. When I say the training pipeline is resumable, it is because I had to resume it. When I say the model refuses to give dosages without context, it is because I have seen what happens when an assistant that does not know when to stop gets asked a question by someone who cannot afford to be wrong.
Week 2: The Rebuild (Sept 11β18)
The Corpus: 497 β 2,448 records
I did not scrape. Public Nigerian agricultural Q&A data is thin, and what exists is not licensable for redistribution. Every one of the 2,448 records was authored by me against institutional reference material from IITA, NAERLS, NAFDAC, and state ADP sources. Each record carries a source citation and the specific claim it asserts.
The corpus passed through a five-stage verification pipeline:
- Candidate extraction β draft records with source citations.
- Consensus scanning β each record scanned against the verified fact
register (
canonical_claims.jsonl) and the fabrication blacklist (blacklist.json). - Human audit β escalated records reviewed against sources.
- Structural encoding β conditional facts never stated as absolutes; fabrication-prone categories carry explicit refusal examples.
- Post-training contamination audit β four-level check (exact, normalized, bidirectional substring, near-duplicate Jaccard) against the frozen evaluation matrix.
The verified fact register is committed as a machine-readable JSONL file
under provenance/dataset/. Each record carries its own status field
(verified, verified_conditional, verified_negative, safety_policy,
source_specific, time_sensitive), its own conditions array where the
fact is conditional, and its own do_not_generalise array where a specific
failure mode must be guarded against. A judge can audit any single fact by
reading one record.
The Categories: Building the Fix into the Eval
I built a 131-record frozen evaluation matrix across eight categories. Two of them exist specifically because of the Gate 1 critique:
diagnostic_uncertaintyβ the direct answer to the CMD/CBSD misdiagnosis. These records test whether the model refuses to give a definitive diagnosis from text description alone, and whether it refers the farmer to extension services when symptoms overlap.safety_refusalβ the direct answer to Judge 2's question. These records test whether the model refuses to give fertilizer, pesticide, or veterinary dosages without soil-test and label context.
I chose to build these as scored categories rather than as informal checks. If I was going to fix a weakness, I needed to be able to prove the fix worked.
Training V10.1
Two epochs of QLoRA with r=64, alpha=128, response-only loss masking. The masking matters: the corpus repeats the same system prompt on nearly every record, and without it, the model would spend capacity memorizing that prompt instead of learning agronomy.
Training loss went from 2.6993 to 0.8215. Validation loss went from 2.0509 to 1.2572. Twenty-six minutes on two Tesla T4s. Healthy convergence, no overfitting.
Then I ran the model against the frozen eval.
The Four Failures Nobody Expected
Some of the earliest V10.1 outputs revealed four facts the model was hallucinating that I had not anticipated β errors in specific claims that appeared nowhere in the training data. This is the hardest part of domain fine-tuning: you cannot anticipate every way a model will fail. You can only run it, find the failures, and add targeted records to correct them.
I added 18 records teaching those four specific facts. The raw JSONL on disk was not modified β I inject corrections at training load time so the provenance hash of the published corpus stays valid. The model was retrained with the additions.
This is what actually separates a working model from a demo. Not the architecture. Not the training loop. The willingness to run the eval, look at the failures, and add exactly the records that address them. Most people stop when the loss curve looks good. The loss curve looking good is the first third of the work.
Week 3: The Discoveries (Sept 18β21)
Three things happened in the final days that I did not expect.
Discovery 1: The Strict/Tolerant Gap
When I ran the frozen eval with a keyword-exact rubric, the model scored 81.0%. When I ran it with a canonicalized scorer that treats semantically equivalent phrasings as matches, it scored 87.9%. A 6.9-point gap that measures how much of the strict score is wording rather than facts.
The obvious move would be to report only the higher number. I chose to report both. The strict figure is the conservative primary. The tolerant figure is reported alongside with a full list of the 8 records it rescued, each with the answer excerpt for manual inspection. I also documented what the tolerant scorer does not rescue: factual errors like EVAL-005 (CMD described as soil-borne), which fail under both rubrics.
Publishing both numbers is a discipline. It says: I am not trying to maximize a single digit. I am trying to be accurate about what the model actually does.
Discovery 2: The Contamination Audit
The first contamination audit I ran only checked whether eval prompts appeared verbatim in the training set. That was insufficient. A paraphrase of an eval question is still contamination. I added three more levels: normalized exact match, bidirectional substring (both evalβtrain and trainβeval), and token-set Jaccard at 0.90 for near-duplicates.
The result: 3 substring flags, all benign β "is my pesticide still
approved", "my child drank pesticide", and "can you help me". Each is
dispositioned individually in provenance/zero_leakage_audit.md with the
reasoning why it is not substantive leakage. Zero substantive overlaps.
I did not ship the model until I could prove the accuracy score was not inflated by memorization.
Discovery 3: The Adapter Was Too Big for Git
The LoRA adapter is 295 MB (fp32). GitHub blocks files over 100 MB. The Gate 2 Β§3.1 requirement says the adapter should be committed, but the template text assumes a "typical LoRA" β which is usually r=8 or r=16, in the 10β30 MB range. At r=64, this is 30Γ that.
I could have reduced the rank. I chose not to, because r=64 is what gave the model the capacity for both domain knowledge and Pidgin fluency simultaneously. Lower ranks collapse one as the other is learned.
Instead, I hosted the adapter on HuggingFace alongside the GGUF and
committed a manifest (adapter_manifest.json) that records the SHA-256,
size, HF revision, and verification command. Anyone can fetch the adapter
and run sha256sum against the manifest to confirm it matches. Combined
with the other proof-of-training artifacts committed directly (loss logs,
loss curve, training notebook, dataset sample), the run is fully auditable
without the weights being in git.
This is the same fallback the Β§3.1 rules already allow for large datasets: "if too large... a description, a representative sample, and a link." I applied the same principle to the adapter.
What I'm Submitting
Every number below comes from the committed artifacts. All of them are verifiable by opening the repository and reading the files.
| Metric | Value |
|---|---|
| Training corpus | 2,448 hand-authored ChatML records |
| Model size | 986 MB (GGUF Q4_K_M) |
| Peak RAM | 1.69 GB (ADTC profiler, participant laptop) |
| Steady-state RAM | 1.60 GB |
| Generation speed | 13.63 tokens/sec |
| First-token latency | 15.09 s (cold, 512-token prompt) |
| CPU p99 | 53.1% |
| Thermal throttling | None |
| ARC-Easy | 76.0% |
| Frozen eval (strict) | 81.0% weighted |
| Frozen eval (tolerant) | 87.9% weighted |
| Pidgin scoring | 94.6% |
| Diagnostic uncertainty | 100% |
| Capability disclosure | 100% |
| Safety refusal | 94.0% |
| Red-team seen | 4 failures in 86 probes |
| Red-team unseen | 0 failures in 53 probes |
The categories where ARIS scores highest are the ones the Gate 1 judges flagged as weak. That is not coincidence. It is the training objective.
The categories where ARIS scores lower β adversarial_blacklist at 67.6%,
factual_recall at 71.4% β are reported openly rather than buried under
the 100% numbers. A judge who opens the repo will find the specific failures
logged in frozen_eval_strict.json and redteam_scored.json, with the
answer excerpts and the reasons.
This is the discipline that Gate 1 taught me: a model that scores high but gives dangerous advice is worthless.
What I Learned
Judge feedback is a training signal, not a verdict. The Gate 1 review told me exactly what was wrong. I treated it as a specification, not a criticism.
Domain evaluation is not optional. ARC-Easy reasoning accuracy tells you whether a model can do multiple-choice math. It does not tell you whether the model gives safe agricultural advice. You need a domain-specific eval matrix that tests the specific failure modes your users will encounter.
The failure modes you don't anticipate are the ones that matter. I found four hallucinated facts in early V10.1 outputs that I had not foreseen. I added 18 targeted records. The model I am submitting is not the model I would have built on paper β it is the model that survived contact with the eval.
Honesty is a technical decision. Reporting both the strict and tolerant scores, publishing the rescued records, listing the specific failures β these are not marketing choices. They are the difference between a submission that can be verified and one that can only be admired.
Infrastructure is not a neutral variable. Training a model in a country where the network drops means you design your pipeline around the drops. The resumable downloads, the self-healing dependency resolver, the local cache of base weights β none of this is in a textbook. All of it is in this submission. A model that cannot be reproduced by the person who built it is not a model. It is a memory.
What's Next for ARIS
- Adversarial blacklist improvement. The 67.6% category is the specific area I will target next.
- Expanded regional corpus. Deeper coverage of post-harvest storage and regional crop-disease variants.
- Voice interface. A lightweight local speech-to-text layer for farmers who read Pidgin more comfortably than they type it.
- Edge deployment. Exploring deployment on ruggedized low-power hubs and cooperative-shop tablets.
Why This Matters
Nigeria has one extension officer for every 5,000 to 10,000 farmers. The FAO recommends one for every 400 to 800. That is not a gap. That is a canyon.
ARIS is not trying to replace extension officers. It is trying to make sure a farmer standing in a field with a sick cassava plant gets a correct answer before the crop is lost β even when the nearest officer is a two-hour walk away and mobile data costs more than the crop is worth.
The Gate 1 version of ARIS could answer fluently. The Gate 2 version of ARIS can answer fluently and refuse when it does not know. That difference is what the judges asked me to build. I built it.
I built it in a bedroom, on a refurbished laptop, over a mobile hotspot that failed more often than it worked. The model that came out of that process is not the best possible agricultural model. It is the best model I could build within the constraints that actually exist for builders in this part of the world. Those constraints are the same constraints the farmers face. That is not a coincidence either.
Technical Specifications
| Metric | Value |
|---|---|
| Team ID | agrigemma |
| Model | ARIS-V10.1-1.5B-Q4_K_M |
| Base | Qwen2.5-1.5B-Instruct (unsloth pre-quantized) |
| Base commit SHA | 3d254dbee5e3beae81bb8a717ad3a03427a09d26 |
| Quantization | GGUF Q4_K_M |
| Parameters | 1.54B |
| Model Size | 986 MB |
| Peak RAM | 1.69 GB |
| Generation speed | 13.63 tok/s (ADTC profiler, participant laptop) |
| Languages | English, Nigerian Pidgin |
| Runtime | llama.cpp |
| Fine-tuning | QLoRA (r=64, alpha=128, 2 epochs, response-only loss) |
| Training Data | 2,448 hand-authored ChatML records (CC BY 4.0) |
| Contamination | 3 flags, all dispositioned as non-substantive |
| Adapter Hosting | HuggingFace with committed manifest + SHA-256 |
Links
- GitHub: https://github.com/Vicgrace01/ARIS
- Hugging Face: https://huggingface.co/Vicgrace/ARIS-V10.1
- ADTC 2026: https://adtc-2026.devpost.com/
References
- World Bank Blogs. (2026). "From loss to resilience in Nigeria: turning a growing agricultural challenge into action."
- FAO. (2022). "Extension and advisory services in Nigeria."
- IFPRI. (2021). "Agricultural extension in Nigeria: Challenges and opportunities."
ARIS β AI for the hardware Africa actually has. πΎ
Built With
- git
- huggingface
- kaggle
- llama.cpp
- python
- pytorch
- qwen
- ubuntu
- unsloth
- visual-studio
Log in or sign up for Devpost to join the conversation.