We will be undergoing planned maintenance on Oct 7th 6:00AM UTC / Oct 7th 2:00AM ET

Inspiration

Setting up phone-protection software correctly can be difficult. A technician may encounter missing information, an unexpected phone state, or an installation error and not know whether it is safe to continue.

We built VowLock Setup Companion to give technicians clear, step-by-step help without depending on cloud AI or constant internet access. It is designed to run locally on the affordable laptops commonly used by small phone shops across Africa.

What it does

VowLock Setup Companion is an offline AI assistant for phone technicians.

It explains the current setup state in simple language, identifies missing evidence, helps diagnose known problems, and recommends whether the technician should continue, wait, retry a known step, or stop.

The AI does not directly control the phone or make security-sensitive decisions. A separate rule-based system verifies the evidence and controls every important action. This prevents the model from inventing commands or continuing when required information is missing.

The current ADTC prototype uses synthetic phone-setup records so we can test the assistant safely without experimenting on customer phones.

How we built it

We tested four publicly available small language models using the same eleven setup scenarios. Qwen3 0.6B provided the best combination of speed, memory efficiency, and reliable responses.

We converted the model to GGUF and selected Q4_K_M quantization. The final model is approximately 397 MB and runs entirely offline through llama.cpp on a CPU.

The system includes:

  • A Python-based deterministic decision engine
  • JSON schemas that restrict what the AI may return
  • A small local Qwen3 model for explanations and guidance
  • Independent validation of every model response
  • A public, credential-free model download script
  • SHA-256 and file-size verification
  • Automated tests and ADTC profiler integration

After downloading the model, inference requires no internet connection or external API.

Challenges we ran into

Small models sometimes produced incomplete JSON or exceeded their response limit. We addressed this with strict output schemas, limited repair attempts, and safe failure behaviour.

During quantization, the profiler exposed a duplicated tied-embedding tensor. We corrected the conversion process and verified the resulting model using parameter counts and file hashes.

In our sealed 24-case evaluation, 23 cases passed. One response remained incomplete after reaching the token limit, so the system rejected it instead of guessing or continuing unsafely.

We also discovered that unstable connections could cause a large model download to restart. We improved the downloader so interrupted transfers resume from the saved byte position.

Accomplishments that we're proud of

  • A 397 MB model that runs locally on an ordinary laptop
  • Eleven out of eleven development scenarios passed
  • Twenty-three out of twenty-four sealed scenarios passed
  • No unsafe continuation when required evidence was missing
  • All 37 repository tests pass
  • Public model download works without credentials
  • Interrupted downloads resume correctly
  • The ADTC profiler produces "measured_on": "participant_laptop"
  • A full physical Ubuntu profiler run measured 29.68 generation tokens per second and 746.57 MB peak memory
  • The physical run completed 50 ARC-Easy samples at 54% normalized accuracy
  • The repeat physical run peaked at 76°C without thermal throttling

The Ubuntu test laptop uses an older 7th-generation Intel Core i5, so we present these as real-hardware compatibility measurements rather than performance on the newer ADTC reference CPU.

We are especially proud that the model remains helpful without being given unchecked authority over a customer's phone.

What we learned

We learned that AI safety cannot depend only on prompting the model to “be careful.” The safer architecture is to separate responsibilities: deterministic software makes consequential decisions, while the language model explains those decisions to the technician.

We also learned that a smaller model can be genuinely useful when its task is focused, its inputs are verified, and its output is independently checked.

Finally, reproducibility matters as much as model quality. Public weights, checksums, offline operation, transparent failures, and honest benchmark boundaries are essential for building trustworthy local AI.

What's next for VowLock Setup Companion

Next, we plan to:

  • Test on the official ADTC laptop specification
  • Conduct usability tests with real phone technicians
  • Add safe, read-only device diagnostics
  • Improve recovery from incomplete model responses
  • Support more installation and troubleshooting workflows
  • Explore local-language and Nigerian Pidgin guidance
  • Integrate the companion into the VowLock vendor application

Physical customer-phone installation will only be tested later through a separately approved safety process.

Built With

Share this project:

Updates

Submission history