Inspiration

Many first-contact clinics operate with intermittent connectivity, limited power, and older laptops. OpenTriage-SLM explores whether a compact local model can support trained staff without making the model the safety authority.

What it does

OpenTriage-SLM is an offline medical-triage research prototype. Structured vital signs and observed danger signs first pass through deterministic RED/YELLOW/GREEN safety gates. A small local GGUF model can then explain the result using local protocol context, but it cannot downgrade or replace the deterministic result. Missing core data produces UNASSESSED rather than a guess.

This is not a clinically validated medical device and must not be used for autonomous diagnosis, prescribing, treatment, or emergency decisions. A qualified clinician remains responsible for assessment and action.

How we built it

We fine-tuned Qwen2.5-1.5B-Instruct, exported a Q4_K_M GGUF, and run it locally with llama.cpp. The Python runtime combines structured intake models, versioned offline protocol retrieval, deterministic safety rules, grammar-constrained explanation output, and rejection of malformed or colour-changing model responses. The public model artifact is hosted on Hugging Face and the repository contains reproducible download, validation, tests, and benchmark scripts.

Constraints and challenges

The target is an 8 GB-class CPU-only laptop with unreliable connectivity. On an Intel i7-6600U development laptop, the measured fast profiler run reached 13.53 generated tokens/s, 1,698.24 MiB peak RSS, and 78 C peak temperature with no detected throttling. The fine-tuning source is medical-exam QA rather than a clinician-validated triage corpus, so the deterministic safety layer remains load-bearing and the evidence boundaries are explicit.

Accomplishments

We delivered an end-to-end offline workflow, a pinned 986 MB GGUF artifact, two declared English healthcare prompts, deterministic RED and GREEN fixtures, offline protocol provenance, 24 automated tests, OPSEC checks, and metadata compatible with the official submission structure.

What we learned

A small language model can make a constrained offline interface more understandable, but it should not own a high-stakes acuity decision. Separating the safety-critical rule engine from the explanatory model made the system easier to test, audit, and run on commodity hardware.

What's next

Before any real-world pilot, we want clinician review of the criteria and wording, a complete official accuracy run, and native-speaker plus clinician review of any local-language interface. We will continue improving reproducibility and measurement rather than presenting this research prototype as clinical validation.

Built With

Share this project:

Updates

Submission history