Inspiration

Every government tender ends with the same slow step. A procurement officer opens a bidder's PAN card, GST certificate, Udyam certificate, OEM authorization letter and local content declaration, and compares each one against the tender conditions by hand. It is repetitive, easy to get wrong at the end of a long day, and hard to audit afterwards, because the reasoning lives in the officer's head.

AI looked like the obvious fix, but we kept coming back to one worry: in public procurement, an AI that quietly accepts or rejects a bidder is unacceptable. A bidder's livelihood can depend on that decision, and "the model said so" is not an answer anyone can defend.

So we asked a different question: what is the most useful thing AI can do for an officer while never making the decision itself?

What it does

The platform takes a tender and a bidder's documents and gives the Procurement Officer a verified, explainable recommendation.

  1. Upload the tender. The system extracts the requirements a bidder must meet (PAN, GST, Udyam, OEM authorization, local content, no debarment).
  2. Upload the bidder's documents. Fields are extracted from each PDF, with a confidence figure, and the officer can correct any field before checking.
  3. Rules decide. Fixed program rules mark each requirement PASS, FAIL or REVIEW and calculate the compliance score and risk level. The same documents always give the same result.
  4. Documents are cross-checked against each other. The PAN inside the GSTIN, the company name and the address must agree across documents.
  5. The officer gets a report: for every check, the evidence found, the source document, and the rule it was checked against, cited from a knowledge base. A short AI-written summary sits on top, and the officer records the final decision.

What makes it different:

  • AI reads and explains, plain code decides. The AI never changes a status or a score. If the AI service is down, the report still completes with a template summary.
  • Confidence is measured, not guessed. Our own code checks that each extracted value has a valid format and really appears in the PDF text. A low score turns a PASS into REVIEW.
  • It never says "reject". Its strongest recommendation is "review", because only the officer can disqualify a bidder.

Our demo case: ABC Technologies has valid PAN, GST, Udyam and OEM documents, but declares 42% local content against the tender's 50%. The system scores it 80/100, risk MEDIUM, flags the failed requirement with the exact numbers and the Make in India rule behind it, and recommends officer review.

This is an academic prototype. All documents are synthetic, and registry and debarment checks use simulated data. The design is ready for authorized API integration, but we do not claim live government connections.

How we built it

We split the project three ways around a shared API contract, so each of us could build against mock data and wire things together later.

  • AI and rule search (Prema Subhasri R): tender and bidder extraction, the rule knowledge base, and the explanation layer.
  • Backend and rule engine (Bhumika S): the FastAPI routes, the scoring engine and the database models.
  • Frontend (Nithyashree J): the React screens for the tender, the bidder's documents, the results and the officer's decision.

The pipeline:

  1. PDF to text. PyMuPDF reads the uploaded tender and bidder documents.
  2. Extraction. Gemini turns the text into structured JSON: the tender's requirements, and the fields from all five bidder documents in a single batched call. Our own code then cleans and checks every value before it is used.
  3. Verification. Each extracted value must have a valid format (PAN, GSTIN and Udyam patterns, real dates) and must actually appear in the PDF text. The share of values that pass becomes the document's confidence, which is measured and not self-reported by the model.
  4. Decision. The rule engine, plain Python, marks each requirement and calculates the score and risk level.
  5. Evidence and citations. A knowledge base of 20 compliance rules sits in ChromaDB with sentence-transformer embeddings. When the check is known, the rule is looked up by its exact id. Otherwise it is found by meaning.
  6. Cross-checks. The PAN inside the GSTIN, the company name and the address are compared across documents.
  7. Summary. One AI call writes a short summary from the verified facts only. If the AI service is unavailable, a template summary is used and the report still completes.

Stack: React and Vite, FastAPI, SQLModel and SQLite, PyMuPDF, sentence-transformers (all-MiniLM-L6-v2), ChromaDB, the Gemini API, and a disk cache for AI answers.

Challenges we ran into

  • Free-tier limits. The AI service allowed only about 20 requests a day per model, and it sometimes returned "overloaded" errors. We learned to tell the two apart. Overloads get retried with growing waits, while a daily limit stops immediately instead of burning more quota. We also cut bidder extraction from five AI calls to one.
  • A cache that kept missing. Answers are cached by the exact text of the prompt, and the upload form sent the files in a different order from our tests, so the same documents looked new every time. Sorting the files before building the prompt fixed it.
  • Three branches, overlapping files. Our work touched some of the same files, and the first merge attempts produced conflicts. We agreed on who owns which folder, merged in a fixed order, and wrote the connection steps down so each of us could follow them.
  • Different environments. One laptop used Python 3.14, which several libraries didn't support yet, and another used conda, where packages were installed in a different place from where the server ran. We fixed these with a documented setup guide.
  • Search that ranks imperfectly. Our rule search sometimes ranked a related rule above the best one. We stopped relying on search for checks we already knew about and now look those rules up by exact id.
  • Trusting AI output. A first version of the engine only checked that a value was non-empty, so a made-up PAN would have passed. Adding format-and-presence verification on top of the AI's output closed that gap.

Accomplishments that we're proud of

  • A working end-to-end flow: from uploading a tender and a bidder's documents to a scored report with evidence, citations and an officer decision screen, running on real services and not mock data.
  • A design we can defend. The AI never decides. Fixed rules do, so the same documents give the same result every time, and a judge can ask "what if the AI is wrong?" and get a real answer.
  • A system that never says "reject". Its strongest recommendation is "review", because only the Procurement Officer can disqualify a bidder.
  • Cross-document checks that really catch things. On a deliberately flawed sample bidder, the blacklist check and all three cross-checks (PAN in GSTIN, company name, address) fired.
  • Reliability under free-tier limits: caching, retries, and a fallback summary mean a quota problem doesn't break the report.
  • A team that learned to ship together: two of us had never used Git branches and pull requests before this project.

What we learned

  • Make the model do the reading and the code do the judging. Splitting those jobs made the system more reliable, easier to explain, and easier to test than asking the model for a verdict.
  • Verify before you trust. Checking AI output against the source text and against known formats catches errors that look perfectly confident.
  • Agree on the data shape first. A shared API contract let three people build in parallel, and it made the final integration a matter of small fixes.
  • Free tiers shape the architecture. Batching, caching and graceful fallbacks aren't polish. They decide whether a demo survives.
  • Team workflow matters as much as code. Clear folder ownership, small pull requests and a fixed merge order saved us from most of the pain.

What's next for GeM Bid Compliance Verifier

  • Authorized registry integration. Connect the existing API-ready design to GSTN, Udyam, MCA and debarment sources, with proper authorization, in place of the simulated data.
  • Scanned and handwritten documents. Add OCR so the system handles real-world uploads and not just digital PDFs.
  • A verified rule library. Review every rule against official sources and have procurement experts sign off, replacing our simplified "prototype policy" entries.
  • Persistent records. Store uploads and officer decisions in the database with a full audit trail, instead of in memory.
  • Login and roles for officers and reviewers.
  • Multi-bidder comparison, so an officer can review all bidders for a tender side by side.
  • Measured accuracy. Evaluate precision and recall on anonymized real tenders, and publish the time saved per bid.
  • More languages, since many tenders and documents aren't in English.

Built With

Share this project:

Updates

Submission history