Inspiration
What it does## Overview
African Offline Coding Assistant is an offline AI coding assistant designed for learners, beginner programmers, developers, and technical users who may not have continuous access to cloud-based AI services.
The project uses Qwen2.5-Coder-1.5B-Instruct in GGUF Q3_K_M format with llama.cpp for local CPU inference. The goal is to provide practical coding explanations, debugging guidance, and code-generation support without requiring cloud inference during use.
Why We Built It
Many users work with limited connectivity, constrained hardware, or environments where cloud AI access can be unreliable or costly. We wanted to demonstrate that useful coding assistance can run locally on ordinary consumer hardware.
How We Built It
We evaluated compact language models with a focus on the trade-off between coding quality, inference speed, and memory usage.
We initially tested a much smaller model, SmolLM2-135M-Instruct, because of its very low resource requirements. Direct coding tests showed that its coding reliability was not strong enough for the intended use.
We therefore selected Qwen2.5-Coder-1.5B-Instruct with GGUF Q3_K_M quantization and integrated it with llama.cpp for CPU-only inference.
The final submission contains:
metadata.jsondownload_model.shREPORT.md- the required GitHub submission structure
- a publicly downloadable GGUF model referenced by the runtime metadata
The model weights are intentionally excluded from Git, while the download script retrieves them from a public source.
Challenges
The main challenge was balancing coding quality, speed, memory use, and offline operation.
We also had to adapt the submission and profiling workflow to a Windows development environment, including configuring the llama.cpp tools required by the ADTC profiler.
Another important challenge was selecting a model that could provide useful coding responses while remaining within the competition's hardware constraints.
Final Results
The latest ADTC participant profile reported:
- 23.5 tokens/sec generation
- 1,212.79 MB peak RSS
- 1,158.95 MB steady-state RSS
- CPU-only inference
- 62.5% CPU p99
- No thermal throttling
- Parameter consistency check passed
- Measured on
participant_laptop
The final model remained well below the competition's 8 GB RAM requirement.
Test Prompts
The submission contains exactly two domain-specific prompts covering:
- Nigerian phone-number validation using Python.
- Debugging and correcting an off-by-one error in Python.
The organizers also provide additional hidden prompts during evaluation to test generalization.
What We Learned
The project demonstrated that model selection is a practical engineering trade-off. The smallest model was faster and lighter, but coding quality was not reliable enough. The final Qwen model provided stronger coding behavior while remaining within the resource constraint.
Our goal is practical local coding assistance without cloud inference.
Log in or sign up for Devpost to join the conversation.