MEDLLM: Medical AI That Runs Entirely Offline

Inspiration

In many Nigerian clinics, access to medical knowledge is not the main problem. Access to it at the right moment is.

Useful guidance already exists in the Nigeria Standard Treatment Guidelines, the NCDC guidelines, the Nigeria Essential Medicines List, and WHO IMCI materials. The trouble is that this information is spread across long PDFs and separate websites. During a consultation, a clinician may not have the time or internet connection to search through all of them.

Cloud-based language models do not fully solve that problem. Connectivity can be unreliable, devices are often shared, and patient information should not have to leave the clinic just to produce a useful answer.

That led us to a practical question:

Can a 4-billion-parameter model provide clear and safe medical information, entirely offline, on an ordinary laptop?

We were not trying to build a model that could talk endlessly about anything. We wanted a focused assistant that would give the action first, ask questions only when an important detail was missing, and never present itself as a replacement for a clinician.

That idea became MEDLLM. It is built on Qwen3.5-4B-Base (Qwen3_5ForConditionalGeneration, 32 layers) and designed to run locally through LM Studio or llama.cpp on consumer hardware.

What MEDLLM does

MEDLLM follows three response patterns. We built these patterns into the training data instead of relying on a prompt to enforce them later.

Response style Share of data Expected behavior
Direct reply 75% Give the answer first. Patient education replies are usually 20 to 150 words, with a hard limit of 180 words.
Clarify, then answer 15% Ask one to three focused questions when a missing fact could change the urgency or advice, then give the action first. The full exchange is limited to 240 words.
Answer with a brief rationale 10% Give the answer first, followed by one to three decisive reasons. The reply stays within 180 words and does not expose hidden chain-of-thought.

Every SFT2 conversation uses the same system instruction:

Give clear and concise medical information. Ask only necessary questions. Give urgent action first when needed.

For triage, the model uses four labels: red, yellow, green, and black. We tuned this behavior to keep recall high for red cases without turning every situation into an emergency. For multiple-choice and guideline questions, the model returns the answer itself, not just a letter.

The purpose is simple: make the response useful under time pressure.

How we built it

We developed MEDLLM in three stages. First, we taught the base model more medicine. Next, we taught it how we wanted it to respond. Finally, we converted it into a format that could run on a laptop.

1. Continual pre-training: building medical knowledge

We continually pre-trained Qwen/Qwen3.5-4B-Base on 717 million curated tokens across 587 054 sources. The corpus combined medical articles and African health guidance with replay data from fineweb-edu, finemath, and codeparrot. That replay data was important because we wanted to add medical knowledge without sacrificing the model's general reasoning ability.

Training ran on a Modal L40S using Unsloth. We used LoRA with r=64 and alpha=128, applied to:

q_proj, k_proj, v_proj, o_proj,
gate_proj, up_proj, down_proj

Other settings included bf16, a maximum sequence length of 4,096 tokens, and packing enabled.

The continual pre-training objective was the standard causal language-model loss:

L_CPT = -(1 / N) * Σ[i=1 to N] log pθ(x_i | x_<i)

Here, N is the number of tokens, x_i is the current token, and x_<i represents all tokens that came before it.

We monitored forgetting with GSM8K, MMLU, and MedQA probes through a custom callback. We also created a proportional evaluation split across general, medical, and reasoning data, with about 120 evaluation rows in total. Early stopping used eval_all_loss rather than a single-domain score.

This stage produced Laptopllm/medLLM_v1_cpt_16bit. In zero-shot evaluation with lm-eval 0.4.12, it scored 81.25% on MMLU Professional Medicine and 88.89% on College Biology.

2. Supervised fine-tuning: teaching the model how to respond

SFT2 did not start again from the original base model. It started from the merged 16-bit CPT checkpoint:

SFT2 starting weights = merged CPT weights

We assembled 30,000 conversations in a single JSONL schema:

id, messages, interaction_style, answer_policy,
evidence, source, family_id, split

The dataset contains 27,000 training examples, 1,500 development examples, and 1,500 internal test examples. Its sources include:

  • 6,070 rows from official Nigerian and African guidance
  • AfriMed-QA and AfriHealth-QA
  • DDXPlus
  • question-only examples from ChatDoctor
  • 8,460 controlled, authored cases that fill gaps in emergency, referral, and maternal-health coverage

Each conversation was formatted with the model's chat template:

tokenizer.apply_chat_template(
    convo,
    add_generation_prompt=False,
)

We trained with assistant-only loss. In other words, the system and user tokens gave the model context, but only assistant tokens contributed to the loss:

L_SFT = -Σ[t in A] log pθ(y_t | y_<t, x) / Σ[t in A] 1

A = assistant tokens only

For clarify_then_answer conversations, both the clarification question and the final response are included in A. Packing was disabled, so conversations were padded rather than joined together. This prevents the end of one patient's conversation from becoming context for another.

For SFT2, we used a fresh LoRA adapter with r=32 and alpha=64. The learning rate followed a cosine schedule:

η_t = η_min + 0.5(η_max - η_min) * [1 + cos(πt / T)]

The schedule used the following values:

  • Maximum learning rate: 5 × 10^-5
  • Warmup ratio: 0.05
  • Training duration: 3 epochs

The effective batch size was 32:

8 examples per batch × 4 gradient accumulation steps = 32

We also used adamw_8bit, max_seq_length=1024, and:

train_on_responses_only(
    "<|im_start|>user\n",
    "<|im_start|>assistant\n",
)

Triage-repair pilot

After SFT2, we ran a small repair stage to improve triage behavior without disturbing the rest of the model.

The repair set contained 1,800 triage examples and 2,700 replay examples, giving us 4,500 training rows. A separate 500-row development set was balanced equally across the four triage labels. Counterfactual families stayed together so closely related cases could not leak across splits.

We trained for one conservative epoch with r=16, alpha=32, and a learning rate of 1 × 10^-5. With an effective batch size of 32, one epoch is approximately:

4,500 training rows / 32 examples per step = 140.625, or about 141 steps

We saved checkpoints at steps 35, 70, 105, and 140, corresponding to roughly 25%, 50%, 75%, and 100% of the run. We kept all four for manual gate evaluation rather than assuming the final checkpoint would be the best one.

3. Edge delivery: getting the model onto a laptop

Training a useful model was only part of the job. We also had to make the final checkpoint work reliably in LM Studio.

Unsloth's save_pretrained_gguf path crashed on our Qwen3.5 fine-tune because of VLM auto-detection and the assertion opt_num_mtp_layers != 0. We worked around that failure with the following export path:

save_pretrained_merged (16-bit)
    -> patch config.json
    -> set mtp_num_hidden_layers = 0
    -> convert_hf_to_gguf.py --outtype f16
    -> llama-quantize q4_k_m

We used the upstream ggml-org/llama.cpp conversion tools. The MTP head is used for throughput, not answer quality. Setting it to zero matched the tensors that were actually present, blk.0 through blk.31, and removed LM Studio's error for the missing blk.32.attn_norm.weight tensor.

The merged 16-bit model was published as Laptopllm/MEDLLM_V1_SFT_2.0, with the Q4_K_M GGUF build in Laptopllm/MEDLLM_V1_SFT_3.0_GGUF.

A few useful dataset calculations

The response-style targets translate into the following counts across 30,000 conversations:

30,000 × 75% = 22,500 direct replies
30,000 × 15% =  4,500 clarify-then-answer conversations
30,000 × 10% =  3,000 answers with a brief rationale

The split also checks out:

27,000 training + 1,500 development + 1,500 internal test = 30,000

For the balanced triage development set:

500 examples / 4 triage labels = 125 examples per label

These calculations are simple, but making them explicit helped us catch drift while combining data from several sources.

Challenges we ran into

Cost of fine-tuning and GPU access.

We had no local GPU capable of training a 4B model. Access meant renting hourly L40S instances via Modal, where a full CPT epoch is 73,367 steps (717M tokens at 4096 ctx, r64 α128 LoRA) and every resumed step after a preemption costs wall-clock and credits. An early 2e-4 LR pilot diverged and wasted a paid run, forcing a conservative 2e-5 cosine restart. This directly pushed us to parameter-efficient LoRA over full fine-tuning (CPT r64 → SFT r32 → pending repair r16), 4-bit adapter training with VolumeCommitCallback to survive preemptions, and capped schedules single epoch CPT and a pending triage repair of only ~141 steps. Choosing Qwen3.5-4B over 7–8B and Q4_K_M (~2.7 GB) over Q5/Q8 was a cost decision as much as a hardware one: larger weights would have multiplied GPU hours, VRAM, download size, and peak RSS risk on the 8 GB budget laptop. Even the evaluation was compute-limited the matched 10,422 Questions llama.cpp run and the Medical MMLU comparison (Qwen4B 0-shot 0.810 vs 5-shot larger models) were batched to avoid repeated paid evals. In short, scarce and expensive GPU access forced a smaller, shorter, LoRA-first stack that trains and quantizes affordably and still fits the laptop.

Qwen3.5 GGUF conversion was not ready out of the box

Two Unsloth issues, combined with the MTP assertion, meant we had to take ownership of the export path. We used upstream llama.cpp, patched conversion/qwen.py, and then had to convert the model first to F16 weights using llama.cpp which was then converted to Q4_K_M GGUF format, which to out the Unsloth Conversion to gguf issue.

The datasets did not naturally fit one schema

Optional fields such as evidence, clinical_domain, and triage which was not available in all the different dataset sources caused DatasetGenerationCastError because Hugging Face datasets inferred different schemas across shards. Rebuilding the dataset with line with our custom schema solved the casting problem.

Medical learning had to be balanced against forgetting

During CPT, a learning rate of 2 × 10^-5 barely changed the loss. During SFT, however, 1 × 10^-4 risked erasing the gains we had already made. We settled on 5 × 10^-5, evaluated every 200 steps, and used EarlyStopping(patience=4) so the stopping rule had enough evaluations to be meaningful.

Evaluation needed the same discipline as training

We fixed the evaluation environment to lm-eval==0.4.12, for the MMLU evaluations kept the same zero-shot tasks, and split related examples. We also set enable_thinking=False, which produces an empty <think></think> scaffold, to prevent the model from spending its output budget on uncontrolled reasoning. We also make a custom evaluation suite with around 10k questions focused on African and Nigerian medical context, using llama.cpp as the engine for the model in order to get the model's results

Accomplishments that we're proud of

The first result we care about is that the model gained medical knowledge without becoming unnecessarily large. After CPT, it retained roughly 74% to 89% accuracy across six MMLU medical subsets. The evaluation ran reproducibly on a Tesla T4 in 765 seconds using bf16.

We are also proud that the response behavior is measurable. The 30,000-row dataset maintains the 75/15/10 style mix within one percentage point. It contains no <think> or Complex_CoT leakage. The median direct reply is no more than 70 words, and direct replies account for at least 65% of the token share. These rules are validated by unified-medical-sft.schema.json.

Safety is treated as a loop, not a claim made after training. A repair checkpoint must meet all of the following gates:

  • red recall of at least 85%
  • improved recall for yellow, green, and black cases
  • a positive change in macro F1
  • no more than a 1% drop in average general performance
  • no individual benchmark falling by more than 2%

We evaluate each candidate through both ChatML and a plain System:/User:/Assistant: harness, using 87 protected triage cases and 725 general multiple-choice questions.

Most importantly, MEDLLM fits on a laptop. Its weights are private, portable, and auditable. It can run in LM Studio on a mid-range machine without sending patient questions to a cloud service.

What we learned

The largest lessons were not about chasing a bigger model.

Unsloth and TRL made training much faster, but Qwen3.5's combination of hybrid linear_attention, full_attention, and MTP meant we still needed to understand and control the export process and learn how the model itself works.

Assistant-only training also mattered more than we expected. Without train_on_responses_only, the model learns to reproduce the prompt as well as the response. For an assistant, that is the wrong behavior.

Packing turned out to be useful for CPT and a poor fit for patient dialogues. The right setting depends on the data, not on a universal training recipe.

The triage pilot reinforced another lesson: one careful repair epoch at 1 × 10^-5 can be more useful than three aggressive epochs.

Finally, the hardest part of the dataset was not its size. It was building reliable source adapters for direct_answer, question_only, structure_only, evidence_grounding, and authored data while keeping each family_id intact. Without that discipline, evaluation leakage can look like progress.

What's next for MEDLLM

  1. Select and merge the triage-repair checkpoint. We will compare ckpt-35, ckpt-70, ckpt-105, and ckpt-140 against the gated scoreboard, merge the winner with merge_and_unload, and quantize it as a single Q4_K_M release.
  2. Add retrieval over Nigerian guidance. The next version will retrieve relevant pages from NCDC, NEML, and IMCI documents and show evidence.url and locator citations in the interface. The source references should come from the retrieval layer, not be memorized in the weights.
  3. Expand Nigerian English and Pidgin coverage. We plan to grow the 8,460-example authored set, preserve the English-only guarantee for triage, and add a separate language: pcm evaluation track.
  4. Run a prospective clinic pilot. We want to compare MEDLLM with paper IMCI chart booklets in a real clinical workflow, measuring on-device latency, triage concordance, and clinician trust.
  5. A Group of Models: We want to work on a group of models around this sizes and parameters to provide choices for people to choose from when working with the product.

MEDLLM is not meant to replace a clinician. It is meant to make clear, relevant medical information available when time and connectivity are limited, using the laptop that is already in the room.

Built With

Share this project:

Updates

Submission history