Inspiration

Vaccine wastage after cold-chain failures is a daily problem across African health facilities. A fridge loses power overnight, and in the morning a nurse stares at 200 doses of pentavalent worth more than her monthly salary. The WHO protocol exists, but it requires cross-referencing three different monitoring instruments, checking freeze sensitivity by product, and consulting manufacturer stability data, all under pressure, often without internet or a pharmacist on site. Most facilities just discard everything to be safe, or worse, use stock that should have been quarantined. Both outcomes cost lives. I wanted to see if a model small enough to run on the same laptop that holds the stock register could walk someone through that decision correctly, right there, during the outage.

What it does

Tigany helps vaccine-store nurses, pharmacists, and cold-chain technicians decide what to do with vaccine stock after a temperature excursion. It interprets three types of cold-chain evidence (vaccine vial monitor stage, freeze indicator status, and continuous temperature logger records), classifies each product by freeze sensitivity, and produces a structured disposition recommendation. The final decision stays with an authorized supervisor. The model works offline, which is exactly when it's needed most: during the power failure itself.

How we built it

  • Started with Qwen2.5-1.5B-Instruct as the base model for its strong size-to-quality ratio
  • Identified three specific factual failures through a baseline audit (VVM misidentification, harmful freeze advice, no structured disposition)
  • Built a 68-example training dataset sourced exclusively from WHO, UNICEF, PAHO, and GAVI public guidance
  • Fine-tuned with LoRA (rank 16, 4 epochs) on a free Kaggle T4 GPU in 12 minutes
  • Merged the adapter and quantized to GGUF Q4_K_M using llama.cpp
  • Validated on the same c6i.xlarge (4 vCPU, 8 GB) that mirrors the ADTC Standard Laptop

Challenges we ran into

  • The base model confidently invented facts about vaccine monitoring equipment (called VVM a "programmable module"), which a system prompt alone could not fix
  • No existing NLP dataset covers vaccine cold-chain disposition, so training data had to be hand-crafted from authoritative guidance documents
  • No GPU access on AWS (quota denied), forcing us to use free Kaggle compute
  • Scoring only evaluates the raw GGUF through llama.cpp, not an app or RAG system, so every improvement had to live in the weights

What we learned

  • A 1.5B model can hold domain-specific procedural knowledge after a targeted fine-tune of just 68 examples
  • Staying small (1.5B Q4_K_M) wins 50% of the score automatically through throughput and memory efficiency
  • The real risk in medical AI is not wrong format but wrong facts stated confidently. Auditing actual model outputs before submission revealed errors that benchmark scores (0.61 MMLU) completely missed
  • Fine-tuning for knowledge injection is fast and cheap; the bottleneck is finding authoritative training data, not compute

Built With

  • amazon-web-services
  • ec2
  • gguf
  • kaggle
  • llama.cpp
  • lora
  • qwen2.5-1.5b-instruct
  • transformers)-llama.cpp-(inference-runtime
  • trl
Share this project:

Updates

Submission history