Inspiration

Utilities receive thousands of complaints across phone, email, and web, handled by separate intake teams and legacy systems. Northwind's backlog showed the cost: most open cases were past their target, and the old way of assigning priority didn't reliably separate a billing question from a real emergency. We wanted to help staff work the right case first, and to do it without sending customer data to a cloud AI provider, because a utility can't compromise on that.

What it does

Trev is a fully on-prem complaint triage and case-assist system.

  • Classifies every complaint locally. Laya, a small decision model that runs on GPU, reads each complaint's intake template and text. It then decides the domain, the team, and the urgency on its own instead of inheriting Northwind's historic priority. Rule-based flags can raise urgency but never lower it. When Laya isn't confident, it marks the case for a human to check.
  • Ranks the worklist by real urgency. Staff see open complaints sorted by priority and days overdue, with the classification attached. New complaints appear on the list the moment they're entered.
  • Two AI buttons per case, with no chatbot. Get context runs read-only tools that pull the complaint, the account's history, patterns for that category and region, and the region's meter data, then returns action items. Generate draft uses the same tools to draft a reply for staff to review. Each button can write only one kind of output, and no action is ever taken on a real account. Every run is logged for audit.
  • Finds root causes, not just symptoms. Dashboards track how the backlog moves over time, break it down along any dimension, and compare estimated-meter-read rates against billing complaints by region. That shows where complaints are actually coming from.

How we built it

  • Backend: FastAPI and SQLAlchemy on Postgres (Tiger Data), with SQL views powering the reporting APIs. Every list endpoint shares the same pagination, sorting, and filtering.
  • Classification: Laya answers typed questions (choice, score, and yes/no probability) in a single forward pass on CPU. Our taxonomy, confidence thresholds, and flag rules sit on top of it. The model loads once at startup, and answers are cached so duplicate complaints are cheap.
  • Agents: LangChain tools bound to a single complaint, calling a local Gemma 4 model through Ollama. Each endpoint exposes exactly one output tool, so the model structurally cannot produce the wrong kind of output.
  • Data: Northwind's complaint backlog, meter data, and account history, seeded into Postgres. The backlog is classified by a resumable background job.

Challenges we ran into

  • Deciding how much to trust the old data. Northwind's historic priorities weren't a good signal, so we had Laya decide urgency from scratch. Backlog rows from the CSV had no customer text, so we fall back to the recorded category when Laya is unsure.
  • Keeping deadline risk from turning everything into P1. Most of the backlog was already past target, so we show deadline risk as a marker instead of letting it raise urgency.
  • Getting tool calling to work locally. Gemma 3 has no tool-calling support in Ollama, so we moved to Gemma 4.
  • CPU speed. Laya takes about 5 seconds per complaint, so we added preloading, caching, and a background job for the backlog.
  • Environment issues, such as macOS shipping Python 3.9 while our code needed 3.10 or newer.

Accomplishments that we're proud of

  • The whole pipeline, from classification to agents to data, runs on local hardware, and a single flag turns the AI agent off completely for confidentiality.
  • The agent can't go off the rails: it's read-only, each button has exactly one output, and every run is audited.
  • Triage that corrects Northwind's own prioritization instead of copying its mistakes.
  • A root-cause view that links billing complaints to estimated meter reads, a problem you can fix upstream.

What we learned

  • A small specialized model like Laya can handle classification well on CPU, so you don't need a large LLM for every step.
  • Limiting what an agent is allowed to do is a more reliable safeguard than prompting it to behave.
  • For triage, the quality of the historic labels matters as much as the model, and sometimes the right move is not to trust them.
  • Designing for how staff actually work (a worklist and two buttons) beat building another chatbot.

What's next for Northwind Triage Desk

  • Write domain-specific agent instructions and tools for billing, metering, and outages, working with Northwind staff.
  • Track which drafts staff edit or reject and feed that back into prompts and thresholds.
  • Move classification to GPU or batch it so large backlogs clear faster.
  • Integrate directly with Northwind's intake systems instead of relying on CSV imports.
  • Pilot with one intake team and measure time-to-resolution and missed-emergency rates against the old process.

Built With

Share this project:

Updates

Submission history