Inspiration

Africa has no shortage of people who want to learn programming, but access to powerful computers, dedicated GPUs, and reliable internet can still be a practical barrier.

We wanted to explore a different approach: what if a useful coding assistant could run locally on an ordinary, resource-constrained computer?

Africasai was built around that idea.

Rather than simply choosing the largest model we could run, we treated efficiency as part of the engineering problem. We wanted to find a model that provided useful programming assistance while remaining practical for CPU-based, offline inference on a computer with approximately 8 GB of RAM.

The principle was simple:

Build for the machine people actually have, not the machine we wish they had.

What it does

Africasai is an offline-first coding assistant designed to support programming education.

It can help learners:

  • understand programming concepts;
  • explain and compare data structures;
  • generate and improve code;
  • reason about programming problems;
  • learn without depending on continuous internet connectivity.

The current prototype uses Qwen2.5-1.5B-Instruct, quantized to GGUF Q4_K_M and running locally through llama.cpp.

We deliberately evaluated the system under CPU-only constraints rather than assuming access to a powerful GPU.

In our development profiling, the final configuration achieved 76% accuracy on a 50-sample ARC-Easy evaluation, with approximately 1.85 GB peak RSS, 1.77 GB steady-state RSS, and 8.08 tokens/second generation speed using two CPU threads.

How we built it

We started by benchmarking small local language models rather than immediately committing to a particular architecture.

We evaluated:

  • SmolLM2-135M-Instruct;
  • Qwen2.5-1.5B-Instruct;
  • Qwen2.5-3B-Instruct;
  • and experimental Ternary-Bonsai models.

All of these experiments were conducted with local inference through llama.cpp or compatible llama.cpp builds.

We compared the models using the same general criteria:

  • generation throughput;
  • first-token latency;
  • memory consumption;
  • CPU utilization;
  • thermal behaviour;
  • reproducibility;
  • and benchmark accuracy where available.

Quantization was particularly important. Instead of treating parameter count as the primary measure of usefulness, we looked for a practical quality/performance point that would allow the assistant to remain responsive on constrained hardware.

The final candidate was Qwen2.5-1.5B-Instruct Q4_K_M.

It was not the largest model we tested. It was the model that gave us the most convincing balance between capability, memory usage, and inference speed for our target environment.

Challenges we ran into

One of our biggest lessons was that a larger model is not automatically a better model for a constrained environment.

Our Qwen2.5-3B Q4_K_M experiment reached approximately 3.84 tokens/second, compared with 8.08 tokens/second for the 1.5B model. It also consumed approximately 3.48 GB peak RSS, compared with 1.85 GB for the 1.5B configuration.

Although the 3B model remained within our approximate 8 GB memory target, its substantially slower generation speed and higher CPU utilization made it less attractive for our intended use case.

We also experimented with Ternary-Bonsai models. Their extremely low-bit representations made them interesting from a memory-efficiency perspective, but the 8B Q2_0 model was far too slow for our target environment. Different runtime configurations produced different results, including approximately 0.3 tokens/second with the mainline configuration and approximately 2.8 tokens/second with a specialized llama.cpp fork.

These experiments reinforced an important point: reducing model size or increasing parameter count in isolation does not tell us whether a model will actually be useful on the hardware available to the user.

Another challenge was reproducible evaluation. Our development environment was a Codespace with limited CPU resources, so we explicitly controlled the number of CPU threads used during inference. We therefore treat these measurements as development benchmarks, rather than claiming that they represent the final ADTC evaluation hardware.

Accomplishments we're proud of

We're proud that Africasai became an actual engineering experiment rather than simply an attempt to package an existing language model as an AI solution.

We established a reproducible workflow for downloading, running, and profiling local models, then compared candidates using measurable performance and quality criteria.

Our selected configuration, Qwen2.5-1.5B-Instruct Q4_K_M, achieved in our development evaluation:

  • 76% ARC-Easy accuracy on a 50-sample evaluation;
  • 8.08 tokens/second generation;
  • 13.08 seconds first-token latency;
  • 1.85 GB peak RSS;
  • 1.77 GB steady-state RSS;
  • 77.8% CPU utilization at p99;
  • no observed thermal throttling.

Most importantly, the model remained comfortably within our approximate 8 GB memory target.

The result is not intended to claim that a 1.5B model can replace larger coding models in every situation. Instead, it demonstrates that a relatively small quantized model can provide useful language-model capability while remaining practical for local CPU inference.

We're particularly proud of the engineering principle behind the project:

AI should be designed around the hardware people actually have.

What we learned

We learned that deploying AI in constrained environments is a different engineering problem from simply maximizing benchmark scores.

Model selection, quantization, context length, CPU utilization, memory consumption, and inference speed are interconnected.

A model that looks impressive in isolation may become impractical when inference takes too long or consumes most of the available memory.

We also learned the value of measuring everything.

Our experiments showed that increasing from 1.5B to 3B parameters came with a substantial inference-speed penalty on our constrained CPU environment. The larger model remained feasible from a memory perspective, but the additional capability had to be weighed against significantly slower generation.

Most importantly, we learned that accessibility can itself be an engineering objective.

Making AI smaller, cheaper, and more locally deployable can be just as meaningful as making it more capable.

What's next for Africasai

This prototype is only the beginning.

Our next step is to move beyond a general coding assistant toward a more complete African educational AI system—one designed around the realities of students learning and building technology locally.

We want to explore:

  • offline local retrieval-augmented generation for educational materials;
  • support for English and an African language;
  • deterministic tools for tasks where exact computation is preferable to language-model generation;
  • improved context and memory management;
  • larger and more coding-specific evaluation sets;
  • deployment on genuinely low-cost consumer hardware;
  • and broader evaluation across African educational use cases.

The long-term vision is simple:

AI that doesn't require Africa to first acquire the infrastructure of somewhere else.

Africasai is our attempt to approach that problem from the other direction: take capable AI and engineer it down to the realities of the people who need it.

Built With

Share this project:

Updates

Submission history