Case study · in production

A platform we built, broke, measured, and fixed.

A complete self-hosted agentic AI stack for a Canadian accounting & professional-services firm — retrieval over a 300,000-chunk corpus, deterministic financial computation over a live ledger, document generation and per-client access control — running the firm's real client work today, on one GPU, with no external API dependency.

The stack

Seventeen containers. One machine. Nothing leaves.

An NVFP4-quantised model served on vLLM, a durable Postgres job queue, a Qdrant vector store, sandboxed code execution, ingestion workers, an admin console and an automated regression suite — all orchestrated under one Docker Compose stack on a single box. It's the current system (a 96 GB GPU), but the same architecture has run on cards as small as 8 GB across several years of builds — right-sized to the workload, not to a fixed rig, and never to a cloud bill.

  • Retrieval. Hybrid semantic + lexical search with a two-stage rerank-and-MMR layer, a measured relevance floor, and a mandatory per-client visibility filter on every query.
  • Correctness. Every figure computed deterministically in pandas over Postgres — never by the model. Free-form analysis requires two independent generations to agree on the digits, or the system visibly refuses.
  • Data integrity. Extraction gated by reconciliation against the source document's own printed totals.
  • Trust. Citations verified in code — reachability, in-range source numbers, embedding freshness — never by asking the model.
  • Access. Fail-closed per-client authorisation from one table, re-checked live at job-execution time.
ops@stack ~ docker compose ps
$ docker compose ps --format "table {{.Name}}\t{{.Status}}" NAME STATUS agent Up (healthy) vllm Up · NVFP4 27B vllm2 Up · polish model qdrant Up · 302,240 pts postgres Up (healthy) ingestion-worker Up ingest-api Up :8020 stack-worker Up · podcast/video exec-worker Up · sandboxed ops-console Up :8200 ollama-chat Up · embeddings searxng Up … 17 services · 0 external API calls

Representative output. Service names generic; no client data.

assistant · grounded answer
What's the GST treatment of exported consulting services to a US client?
Exported services to a non-resident are generally zero-rated under the Excise Tax Act, Sch. VI, Part V, s. 7 — GST charged at 0%, with input tax credits still recoverable, subject to the place-of-supply and non-resident conditions.
◦ Excise Tax Act — Sch. VI, Part V, s.7  ✓ citation verified
◦ 169 chunks in scope · relevance 0.71 · floor 0.45
And what are our top three client AR balances?
⟂ scope: general  That's client-ledger data. This conversation isn't scoped to a client, so I won't pull it — switch to a client workspace and I'll answer from the ledger.

Anonymized exchange. Verified citation + fail-closed scope.

Grounding

It cites what it can prove — and refuses what it can't reach.

Every citation is checked in code: the source is reachable, the quoted numbers are in range, and the indexed copy is still fresh. Off-topic questions don't get a confident near-miss — a measured relevance floor sits in the gap between noise and signal, so the system says “I don't have that” instead of citing the nearest unrelated document.

And a general chat can never be answered with client data by accident. Access is scoped up front and re-checked at execution time; “found nothing”, “denied” and “not in scope” are three different answers, never one.

Correctness

The number is computed, not generated.

Letting the model compute a figure over a real ledger gives you a different wrong number every run. This is the failure that ends engagements — not an error, not a crash, just three plausible, well-formatted, different numbers, none of them right.

So anything presented as a figure is computed deterministically in code; the model only chooses which computation and phrases the result. Where free-form analysis is unavoidable, two independent passes must agree on the digits or the system refuses. And a qualified professional stays the signer — the system's job is to make a figure defensible and show its work, not to replace the judgment that's accountable for it.

“total billable WIP this engagement” · same query ×3
Run 1 — model-written SQL$1,375.00
Run 2 — model-written SQL$8,975.00
Run 3 — model-written SQL$8,975.00
Deterministic (pandas / Postgres)$22,753.00

Temperature 0. Same prompt. A different ad-hoc filter invented each run — the figure a partner would have signed.

Anonymized reproduction of a real incident.

The six expensive ones

What actually goes wrong — ranked by what it costs.

These aren't accidents unique to us — they're the failure modes every retrieval agent hits, and each produced output indistinguishable from success, which is exactly why a demo can't surface them. The difference is that we found and closed ours ourselves: the dangerous ones — the scoping slip, the data near-miss — were caught in our own testing and internal verification, before they ever reached a client. That's what separates a practice that has operated one of these from one that is about to build its first. Published here, candidly, because a vendor who won't show you their failure list is a vendor who hasn't found it yet.

01

Retrieval silently returns a fraction of your corpus.

A vector-DB filter meant to mean “public” matched only records where a field was present and null — not records where it was absent. Two ingestion paths never set it. Search still worked; it just came from a quarter of the library.

Cost: 226,777 chunks — the two most valuable corpora — invisible to every query for weeks. Found by auditing a payload dump, not by any test.
02

When the answer isn't there, the model cites the nearest noise.

Semantic search always returns something. Those near-misses get presented as sources with the same confidence as a real citation. The fix is measurement, not a vibe: off-topic queries cluster in a narrow band; the floor belongs in the gap, not inside the noise.

Measured: off-topic noise 0.38–0.42 · real relevance 0.60+ · a floor at 0.40 sits inside the noise Symptom: a chat about a video game cited farming-income tax PDFs
03

A router that sorts every question into one bucket can't handle real questions.

Classify-then-dispatch structurally can't serve a multi-part question, and it drags ordinary questions out of the grounded path into a handler that can't cope. A bounded tool loop fixes a whole class at once — the hard part is the preconditions that stop it doing something expensive on a coincidence.

Real incident: a general chat about public quantum-computing stocks was answered with the firm's client roster — because one keyword force-routed a client-data tool.
04

Letting the model compute a number gives you a different wrong one each run.

Model-generated analytics over a real ledger, at temperature zero, invented a different filter every run. In a regulated context that is the worst output a system can produce — plausible, formatted, and wrong. (See the panel above.)

Same question, three runs: 1,375.00 / 8,975.00 / 8,975.00 · correct: 22,753.00
05

Authorization checked at request time is not authorization.

Long-running AI work becomes a queued job, authorized at T0 and executed at T1 — and permissions move in between. Every job carries its requester and re-checks live at execute time; a revoked grant stops it mid-flight and records why. Nothing is held for later delivery.

Adjacent lesson: the permission check itself must be immune to load. Ours once timed out at 15s under a CPU-starved API, failed closed, and locked the authorized user out mid-engagement.
06

Every silent limit added to fix a timeout changes what the system can find.

Under pressure someone caps something to make a symptom disappear. It's invisible in the output and silently redefines the product. The rule: no cap, floor, budget or timeout without an explicit decision — and when approved, it's named in the output wherever it affected the result.

One silent rewrite (a full-range scan quietly narrowed to the top 1,000 ports to stop a timeout): three of six vulnerable services lived outside that range and were never found.

Method

Two audits, not one.

When a bug is found, the shape audit sweeps the codebase for every other instance of the same pattern and writes the regression test before calling it done. But the shape audit only finds copies of a bug — it doesn't find why the bug was writable.

So the second audit — the reason — asks: what did this code have to guess, and did the system already know the answer somewhere upstream? In the client-roster incident, the clever fix (a better keyword guess) covered one tool. The real fix was one line: a flag three functions upstream already recorded whether the chat was client-scoped, and nothing consulted it. Consuming that known state as a precondition covered five tools.

A fix that makes an inference more accurate is a mitigation. Known state consumed as a precondition is a fix. Teams that only run the first audit ship the same bug class, quarterly, forever.

Want the full technical walkthrough?

The figures here are anonymized for client confidentiality — but they're not asking for blind trust. A live walkthrough of the running system, the architecture, the retrieval pipeline, decision trees and the operational runbook are available on request, under NDA. See it work, then decide.