Proposal · In-house agentic AI capability

You are about to spend a year
learning what I already paid for.

I built a production agentic AI platform for a Canadian accounting and professional-services firm — retrieval, deterministic financial computation, document generation, per-client access control — entirely on one machine, with zero external API calls. It runs the firm's real client work today. I want to help your consultancy build the same capability without repeating the eleven months of failures that made mine correct.

Pouyesh Sabouri Delta, British Columbia 22 years · TELUS → Government of Alberta → Long View Systems → Coast Capital Savings → Fortinet
302,240Indexed
KB chunks
96 GBSingle GPU,
no cloud
262kToken context
in production
29Deterministic
finance recipes
207Automated
test files
0Bytes leaving
the building

The problem with the current market

Everyone can demo an agent. Almost nobody can operate one.

A retrieval-augmented agent takes about two weeks to demo and about a year to make trustworthy. The gap between those two numbers is where consultancies lose margin, lose client confidence, and quietly shelve the practice.

The demo works because you tested it on questions you already knew the answers to. Production breaks because of a specific, finite set of failure modes that all look like the system working — a confident answer, a clean citation, a plausible number. Nothing throws. Nothing alerts. The client finds it.

Every serious failure I hit produced output that was indistinguishable from success. That is the whole problem, and it is why the demo tells you nothing.

I have hit those failures, diagnosed them against real data, measured the fix, and shipped it. Below are the six that cost the most — ranked by what they cost, not by when they happened. If your practice lead cannot tell you what each of these is and how they would catch it, your first build will find them for you, in front of a client.

The six expensive ones

What actually goes wrong, and what it costs.

01

Your retrieval is silently returning a fraction of your corpus.

A vector-database filter intended to mean "public content" matched only records where a field was present and null — not records where the field was absent entirely. Two ingestion paths never set the field at all.

Search still worked. Results still came back. They just came from a quarter of the library, and the system answered "we don't have that" about documents sitting right there.

Cost: 226,777 chunks — the two most valuable corpora in the system — invisible to every query, for weeks. Found by auditing a payload dump, not by any test.
02

When the answer isn't in your corpus, the model cites the nearest noise.

Ask an off-topic question and semantic search returns something — it always returns something. Those near-misses then get presented as sources, with the same formatting and the same confidence as a real citation.

The fix is not a vibe. It is measurement: off-topic queries cluster in a narrow similarity band, genuinely relevant content sits well above it, and the floor belongs in the gap — not inside the noise plateau, where it catches some junk and misses the rest.

Measured: off-topic noise 0.38–0.42 · real relevance 0.60+ · a floor at 0.40 sits inside the noise (catches 0.396, misses 0.416) Symptom before the fix: a chat about a video game cited farming-income tax PDFs
03

A router that classifies every question into one bucket cannot handle real questions.

The standard architecture — classify the utterance, dispatch to a handler — structurally cannot serve a multi-part question, and it drags ordinary questions out of the grounded path into a specialist handler that then can't cope.

Replacing it with a bounded tool loop, where the model chooses which tools to call and with what query, mid-turn, fixed a whole class of garbled answers at once. The hard part is not the loop. It is the preconditions that stop the loop doing something expensive on a coincidence.

Real incident: a general chat about public quantum-computing stocks was answered with the firm's client roster — because one keyword force-routed a client-data tool.
04

Letting the model compute a number gives you a different wrong number every run.

This is the one that ends engagements. Model-generated analytics code over a real ledger, run three times at temperature zero, invented a different ad-hoc filter each time.

Not an error. Not a crash. Three plausible, well-formatted, different numbers — and none of them right. In a regulated or financial context that is the worst output a system can produce.

Same question, three runs: 1,375.00 / 8,975.00 / 8,975.00 Correct answer: 22,753.00

The architecture that fixes it: deterministic computation in code for anything that will be presented as a figure, the model restricted to choosing which computation and phrasing the result — plus, where free-form analysis is unavoidable, two independent generations that must agree on the digits or the system visibly refuses to answer.

05

Authorization checked at request time is not authorization.

Long-running AI work becomes a queued job. A job is authorized at T0 and executes at T1, and permissions move in between. Almost every reference architecture checks once.

Every job in my system carries its requester and re-checks live at execute time; a revoked grant stops the job mid-flight, discards partial work, and records why. Nothing is held for later delivery — holding a result the requester may no longer be entitled to see is a liability with no upside.

Adjacent lesson: the permission check itself must be immune to load. Ours once called an HTTP API that got CPU-starved by another workload, timed out at 15s, failed closed — and locked the authorized user out mid-engagement.
06

Every silent limit added to fix a timeout changes what the system can find.

Under delivery pressure, someone caps something to make a symptom go away. It is invisible in the output. It silently redefines the product.

The rule I now work to: no cap, threshold, budget or timeout without an explicit decision — and when one is approved, it must be named in the output wherever it affected the result. A scan that covered 1,000 of 65,535 ports has to say so.

Cost of one silent rewrite (a full-range scan quietly narrowed to the top 1,000 to stop a timeout): three of six vulnerable services on the target lived outside that range and were never found.

Method

Two audits, not one.

The habit I would most want to install in your team is not a technology choice. It is how a fix gets closed.

First audit — the shape. When a bug is found, sweep the codebase for every other instance of the same pattern: the same return convention, the same missing guard, the same caller treating a sentinel as permission. Check the sibling enum values, not just the one in the report. Write the regression test before calling it done.

Second audit — the reason. The shape audit finds copies of a bug. It does not find why the bug was writable. So ask: what did this code have to guess, and did the system already know the answer somewhere upstream?

In the client-roster incident above, the shape audit was done properly — all 212 keywords reviewed, the failure rate measured, three layers of improvement shipped. All three were a better guess. The actual fix was one line: a flag three functions upstream already recorded whether the conversation was scoped to a client, and nothing consulted it. Consuming that known state as a precondition covered five tools where the clever fix covered one.

A fix that makes an inference more accurate is a mitigation. Known state consumed as a precondition is a fix.

Teams that only ever run the first audit ship the same bug class, quarterly, forever — and the client is the one who notices.

Engagement

What the first ninety days look like.

1
Weeks 1–3 · Evidence before architecture

Pick one workflow and build the scoring corpus first.

One real workflow with a real owner, and — before any model work — a corpus of questions with known correct answers, including the ones the system should refuse. Without this you cannot tell a tuning win from a lucky sample, and every later decision becomes taste.

2
Weeks 4–8 · The spine

Retrieval and access control, built together.

Hybrid retrieval — vector search cannot find an identifier, because an identifier has no meaning to embed — with a measured relevance floor, a mandatory visibility filter on every query, and a fail-closed permission model whose failures say which failure occurred. "Found nothing", "denied", and "nothing ingested" must never render as the same output.

3
Weeks 9–12 · Trust

Deterministic computation and code-level verification.

Anything presented as a figure computed in code, not by the model. Citations verified programmatically — never by asking the model whether it complied. Extraction gated by reconciliation against the source document's own totals. Then ship to a real user group and watch what they actually ask, which is never what the pilot asked.

Two weeks

Readiness assessment

Your intended architecture reviewed against these six failure modes, with a written finding and a remediation order. Useful whether or not I build anything.

Twelve weeks

Build and hand over

The ninety days above, delivered with your engineers in the code, not watching. You keep the system, the tests, and the reasoning behind every threshold in it.

Ongoing

Fractional practice lead

Design authority and code review across your client engagements, so the second build is cheaper than the first and the fifth is a product.

Any shape

Fully on-premises

Everything I have built runs with no external API dependency. For clients with data-residency, privilege or regulatory constraints, that is not a preference — it is the entry requirement.

Candour

What I am not.

  • Not a research lab. I don't train foundation models. I make existing ones safe to put in front of someone's books.
  • Not a slideware practice. I write the code, run it in production, and get paged when it's wrong.
  • Not willing to ship an agent that states figures it can't show its work for. If that is the deliverable, I am the wrong person and I will say so in week one rather than month six.
  • Not selling a platform. There's no licence here. You end up owning what gets built.

What I do bring that a pure ML hire usually doesn't:

  • Twenty-two years accountable for production. TELUS, the Government of Alberta's Justice & Solicitor General network, managed services for municipal government and hospitality clients at Long View Systems, then a credit union's entire network, telephony and banking infrastructure stack. I have been the person on the call when it's down.
  • Consulting delivery, not just engineering. Presales, RFP response, design documentation, client working groups, SLA escalation — at Long View and now at Fortinet. I can sit in front of your client and be the reason they sign.
  • Automation as a management discipline. At Coast Capital I led teams of 8–12 and pushed Python, Django, Ansible and Jenkins into the operational core — including automating my own role away: a program that created and removed networks, VPNs, interfaces and firewall rules across the organisation. I have run the "get the team writing code" transition that most consultancies are about to attempt.
  • Security architecture. Fortinet, Cisco and Palo Alto at certified-specialist depth, with delivery against PCI and HIPAA. Agentic systems execute code and reach data — the isolation, egress, identity and audit conversation is the entire risk conversation, and it is one I have been having for two decades.

Every figure in this document is a real measurement from a production system, not an illustration. Client names, data and identifying detail are deliberately excluded; the platform described is referred to only by its domain. A full technical walkthrough — architecture, retrieval pipeline, decision trees and operational runbook — is available on request.