Inspiration
Picture this: It's 11pm in Lagos. The IT officer at a mid-sized fintech company gets a notification — 847 failed SSH login attempts in 12 minutes from an overseas IP. He has no SOC team. No Splunk. No CrowdStrike. Just a laptop, a free IDS tool, and a log file he doesn't have time to read. He dismisses the alert. He's seen hundreds like it this week. Three hours later, an attacker is inside. This scenario is not hypothetical. It plays out daily across Nigerian fintechs, hospitals, NGOs, and public sector organisations, institutions that handle sensitive data but cannot afford enterprise security tools built for organisations with stable cloud connectivity, dedicated analysts, and five-figure monthly SaaS budgets. My undergraduate thesis studies this exact problem: alert fatigue in under-resourced security teams. The research is clear: the volume of alerts is not the problem. The lack of tools that work under African infrastructure constraints is. TactOS is our answer. An offline AI triage assistant that runs on the laptop already on that IT officer's desk, no internet required, no subscription fee, no cloud dependency. He pastes the alert. TactOS tells him what it means, what to do, and who to call. In ten seconds. Even during load-shedding. Built with my teammate Jameel Adediran.
What it does
TactOS is a fully offline, on-device cybersecurity triage assistant for African SMEs. An analyst pastes an IDS alert or incident description into the terminal and receives an immediate structured triage brief, recommended immediate actions, and an incident response memo — all generated locally in seconds with no internet connection required.
How we built it
We built TactOS using Qwen2.5-1.5B-Instruct quantized to GGUF Q4_K_M, running through llama-server from llama.cpp. A pure Python TF-IDF retrieval engine — zero external dependencies — pulls relevant context from a local knowledge base covering MITRE ATT&CK Initial Access techniques and an SME incident response playbook with Nigerian regulatory guidance (ngCERT, NDPC). The LLM is queried via HTTP on localhost using Python's built-in urllib, keeping the entire stack offline and dependency-free.
Challenges we ran into
- Dependency hell under a Nigerian internet connection.
Every large pip install —
llama-cpp-python,chromadb,faiss-cpu— either timed out mid-download or hit dependency conflicts. The zero-dependency TF-IDF RAG was not an academic design choice. It was a survival decision. I built the entire retrieval engine from scratch using only Python's standard library:math,re,collections. It works. - llama-cli interactive mode.
The newer llama.cpp builds auto-detect chat templates and enters conversation mode regardless of flags. After several failed attempts with
--no-cnvand--single-turn(which killed stdout entirely), I switched tollama-serverand queried it viaurllib.request— Python's built-in HTTP client. Clean output, no hanging, no banner noise. - WSL2 limitations.
Working entirely in WSL2 on Windows means thermal sensor readings return
null— the profiler cannot read CPU temperature through the virtualisation layer. The organiser evaluation on real Linux hardware will populate this correctly. Building and testing in a virtualised environment added friction at every step. ## Accomplishments that we're proud of - Built a fully functional offline cybersecurity triage assistant from scratch in under two weeks
- Achieved 11.41 tokens/sec on CPU-only hardware with no GPU
- Peak RAM of only 1.82 GB — well within the 8 GB constraint
- Zero external Python dependencies — runs entirely on the standard library
- Grounded in real academic research on alert fatigue in African cybersecurity contexts
The scoring formula for this competition is:
$$S_{total} = 0.50 \cdot S_{acc} + 0.30 \cdot S_{perf} + 0.20 \cdot S_{eff} - P_{thermal}$$
My benchmarks on an Intel Core i5-10210U (CPU-only, no GPU):
Metric Value
Generation speed 11.41 tokens/sec
Peak RAM 1,820 MB
Thermal throttling None
Which projects to:
$$S_{perf} = 100 \times \frac{11.41}{15.0} \approx 76.1$$
$$S_{eff} = 100 \times \frac{7.0 - 1.82}{7.0} \approx 74.0$$
## What we learned
Quantization is not just a performance trick — it is the difference between a model that fits in 8 GB and one that doesn't. Q4_K_M hit a sweet spot: 1.82 GB peak RAM, negligible quality loss on structured generation tasks.
Small models (1.5B parameters) are surprisingly capable at domain-specific structured output when the prompt is well-engineered and grounded with RAG context.
The
llama-serverHTTP interface is far more reliable for programmatic use thanllama-cli— treat it like a local microservice, not a CLI tool. Offline-first is not a constraint. For most of Africa, it is the baseline. ## What's next for TactOS Expand the knowledge base with MITRE ATT&CK lateral movement, privilege escalation, and exfiltration techniques Add a lightweight web UI (single HTML file, no framework) for non-terminal users Nigerian threat intelligence layer: common attack patterns targeting Nigerian fintech and telco infrastructure Hausa and Yoruba prompt support for broader.
Log in or sign up for Devpost to join the conversation.