Inspiration
Tender and procurement documents can be difficult and time-consuming to analyse, especially for small contractors who may not have dedicated procurement teams or expensive software.
The idea behind TenderGPT Edge came from a simple question:
Can useful AI-powered tender intelligence run entirely on an ordinary laptop, without sending sensitive procurement documents to the cloud?
For African businesses, this matters. Connectivity, cloud costs, data privacy, and access to high-end computing resources can all be constraints. We wanted to explore whether a smaller local language model could be combined with deterministic document processing and evidence retrieval to create something genuinely useful on commodity hardware.
Rather than building another cloud chatbot, we focused on an offline-first system designed around a real business workflow: understanding procurement documents and helping a contractor decide what deserves attention before submitting a bid.
What it does
TenderGPT Edge analyses procurement and tender documents locally.
The system can:
- inspect and extract text from PDF documents;
- identify important tender information such as dates, locations, duration, sector and workforce requirements;
- retrieve relevant evidence from the source document;
- provide page-aware evidence for answers;
- calculate and display tender readiness information;
- identify procurement risks and missing information;
- provide bid-oriented recommendations;
- interact with a local GGUF language model through
llama.cpp; - operate without requiring cloud inference during normal AI use;
- export tender analysis and reports.
The important distinction is that the language model is not expected to do everything.
TenderGPT Edge combines deterministic extraction and evidence retrieval with local AI reasoning. This gives the system a structured foundation before the language model is asked to interpret information.
How we built it
TenderGPT Edge was built as an offline-first desktop application using Python and PySide6.
The document intelligence pipeline uses PDF inspection and extraction, with OCR available for documents where normal text extraction is insufficient.
The architecture separates the major responsibilities:
- Document layer — PDF inspection, text extraction and chunking.
- Knowledge layer — deterministic evidence retrieval and source references.
- Reasoning layer — procurement analysis and structured decision logic.
- AI layer — a local Qwen2.5 1.5B Instruct GGUF model running through
llama.cpp. - Desktop layer — a PySide6 interface for interacting with the system.
- Evaluation layer — deterministic retrieval benchmarks and the official ADTC profiler.
We deliberately avoided making the language model the source of truth for document facts. Evidence retrieved from the tender remains the foundation for the system's answers.
The local model is quantized as GGUF Q4_K_M so that it can be evaluated on constrained laptop hardware.
Challenges we ran into
The biggest challenge was not simply getting an LLM to answer questions.
It was building a reliable system around a small local model.
We had to deal with:
- large and inconsistent procurement PDFs;
- scanned documents requiring OCR;
- unreliable or incomplete table extraction;
- preventing invalid extracted data from contaminating downstream calculations;
- deterministic evidence retrieval;
- local model memory constraints;
- llama.cpp integration;
- responsive desktop UI behaviour;
- separating presentation logic from business logic;
- designing the system so that it remains useful when the model is unavailable.
We also learned that an AI application should never pretend that an operation succeeded when it did not. Explicit validation, diagnostics and evidence are more valuable than a convincing-looking answer.
Hardware & Performance Notes
All performance numbers in this submission (submission.json) were measured on an older Intel Core i3-5005U (2.0 GHz) laptop with only 2.8 GB of usable RAM.
This machine is significantly below the official ADTC Standard Laptop profile (Intel Core i5 10th–12th gen or AMD Ryzen 5 with 8 GB RAM).
On the weaker hardware we observed:
- ~8.7 tokens/sec generation speed
- ~25 seconds first-token latency
We expect substantially better throughput and lower latency when the same model and application are run on the official ADTC reference hardware. Memory usage remained efficient (~1.8 GB peak), which aligns well with the efficiency scoring criteria.
The model (Qwen2.5-1.5B-Instruct Q4_K_M) and application were deliberately designed to run on constrained commodity laptops, which is the core goal of this challenge.
Accomplishments that we're proud of
We are particularly proud that TenderGPT Edge evolved from a document-processing concept into a working offline intelligence system.
The application has a complete desktop workflow covering document inspection, extraction, evidence search, AI assistance, tender intelligence, readiness analysis, recommendations and reporting.
We also built the system around a small local model rather than assuming access to a powerful cloud GPU.
Another important accomplishment was integrating evidence retrieval with local AI. The system can ground its reasoning in information extracted from the actual procurement document rather than relying solely on the language model's general knowledge.
Most importantly, we are now testing the system against the official Africa Deep Tech Challenge profiler rather than relying on assumptions about performance.
What we learned
We learned that deploying AI on constrained hardware is a systems engineering problem.
Model selection and quantization matter, but so do document preprocessing, retrieval, context management, memory usage, CPU configuration and application architecture.
We also learned that a smaller model can become much more useful when it is surrounded by good deterministic systems.
For procurement intelligence, the question is not simply:
"Can the model generate an answer?"
It is:
"Can the system produce a useful answer, show where that answer came from, and do it reliably on the hardware available to the user?"
That principle has shaped the architecture of TenderGPT Edge.
What's next for TenderGPT Edge
The immediate priority is rigorous evaluation.
We are using the official ADTC profiler to establish real measurements for inference performance, memory usage and throughput on the target laptop profile.
We will use those results to identify bottlenecks and optimise the system without compromising reliability.
Longer term, TenderGPT Edge could expand beyond basic tender analysis into a broader offline procurement intelligence platform, including deeper compliance checking, opportunity matching, historical tender analysis, bid preparation assistance and additional African procurement workflows.
The core principle will remain the same:
Useful AI should not require a data centre.
Built With
- application
- artificial-intelligence
- desktop
- gguf
- information-retrieval
- llama.cpp
- local-ai
- natural-language-processing
- ocr
- offline-ai-large-language-models
- pdf-processing
- procurement
- pyside6
- python
- qwen2.5
- retrieval-augmented-generation
- sqlite
- tender-intelligence
Log in or sign up for Devpost to join the conversation.