Proposal · In-house agentic AI capability
I built a production agentic AI platform for a Canadian accounting and professional-services firm — retrieval, deterministic financial computation, document generation, per-client access control — entirely on one machine, with zero external API calls. It runs the firm's real client work today. I want to help your consultancy build the same capability without repeating the eleven months of failures that made mine correct.
The problem with the current market
A retrieval-augmented agent takes about two weeks to demo and about a year to make trustworthy. The gap between those two numbers is where consultancies lose margin, lose client confidence, and quietly shelve the practice.
The demo works because you tested it on questions you already knew the answers to. Production breaks because of a specific, finite set of failure modes that all look like the system working — a confident answer, a clean citation, a plausible number. Nothing throws. Nothing alerts. The client finds it.
Every serious failure I hit produced output that was indistinguishable from success. That is the whole problem, and it is why the demo tells you nothing.
I have hit those failures, diagnosed them against real data, measured the fix, and shipped it. Below are the six that cost the most — ranked by what they cost, not by when they happened. If your practice lead cannot tell you what each of these is and how they would catch it, your first build will find them for you, in front of a client.
The six expensive ones
A vector-database filter intended to mean "public content" matched only records where a field was present and null — not records where the field was absent entirely. Two ingestion paths never set the field at all.
Search still worked. Results still came back. They just came from a quarter of the library, and the system answered "we don't have that" about documents sitting right there.
Cost: 226,777 chunks — the two most valuable corpora in the system — invisible to every query, for weeks. Found by auditing a payload dump, not by any test.Ask an off-topic question and semantic search returns something — it always returns something. Those near-misses then get presented as sources, with the same formatting and the same confidence as a real citation.
The fix is not a vibe. It is measurement: off-topic queries cluster in a narrow similarity band, genuinely relevant content sits well above it, and the floor belongs in the gap — not inside the noise plateau, where it catches some junk and misses the rest.
Measured: off-topic noise 0.38–0.42 · real relevance 0.60+ · a floor at 0.40 sits inside the noise (catches 0.396, misses 0.416) Symptom before the fix: a chat about a video game cited farming-income tax PDFsThe standard architecture — classify the utterance, dispatch to a handler — structurally cannot serve a multi-part question, and it drags ordinary questions out of the grounded path into a specialist handler that then can't cope.
Replacing it with a bounded tool loop, where the model chooses which tools to call and with what query, mid-turn, fixed a whole class of garbled answers at once. The hard part is not the loop. It is the preconditions that stop the loop doing something expensive on a coincidence.
Real incident: a general chat about public quantum-computing stocks was answered with the firm's client roster — because one keyword force-routed a client-data tool.This is the one that ends engagements. Model-generated analytics code over a real ledger, run three times at temperature zero, invented a different ad-hoc filter each time.
Not an error. Not a crash. Three plausible, well-formatted, different numbers — and none of them right. In a regulated or financial context that is the worst output a system can produce.
Same question, three runs: 1,375.00 / 8,975.00 / 8,975.00 Correct answer: 22,753.00The architecture that fixes it: deterministic computation in code for anything that will be presented as a figure, the model restricted to choosing which computation and phrasing the result — plus, where free-form analysis is unavoidable, two independent generations that must agree on the digits or the system visibly refuses to answer.
Long-running AI work becomes a queued job. A job is authorized at T0 and executes at T1, and permissions move in between. Almost every reference architecture checks once.
Every job in my system carries its requester and re-checks live at execute time; a revoked grant stops the job mid-flight, discards partial work, and records why. Nothing is held for later delivery — holding a result the requester may no longer be entitled to see is a liability with no upside.
Adjacent lesson: the permission check itself must be immune to load. Ours once called an HTTP API that got CPU-starved by another workload, timed out at 15s, failed closed — and locked the authorized user out mid-engagement.Under delivery pressure, someone caps something to make a symptom go away. It is invisible in the output. It silently redefines the product.
The rule I now work to: no cap, threshold, budget or timeout without an explicit decision — and when one is approved, it must be named in the output wherever it affected the result. A scan that covered 1,000 of 65,535 ports has to say so.
Cost of one silent rewrite (a full-range scan quietly narrowed to the top 1,000 to stop a timeout): three of six vulnerable services on the target lived outside that range and were never found.Method
The habit I would most want to install in your team is not a technology choice. It is how a fix gets closed.
First audit — the shape. When a bug is found, sweep the codebase for every other instance of the same pattern: the same return convention, the same missing guard, the same caller treating a sentinel as permission. Check the sibling enum values, not just the one in the report. Write the regression test before calling it done.
Second audit — the reason. The shape audit finds copies of a bug. It does not find why the bug was writable. So ask: what did this code have to guess, and did the system already know the answer somewhere upstream?
In the client-roster incident above, the shape audit was done properly — all 212 keywords reviewed, the failure rate measured, three layers of improvement shipped. All three were a better guess. The actual fix was one line: a flag three functions upstream already recorded whether the conversation was scoped to a client, and nothing consulted it. Consuming that known state as a precondition covered five tools where the clever fix covered one.
A fix that makes an inference more accurate is a mitigation. Known state consumed as a precondition is a fix.
Teams that only ever run the first audit ship the same bug class, quarterly, forever — and the client is the one who notices.
Engagement
One real workflow with a real owner, and — before any model work — a corpus of questions with known correct answers, including the ones the system should refuse. Without this you cannot tell a tuning win from a lucky sample, and every later decision becomes taste.
Hybrid retrieval — vector search cannot find an identifier, because an identifier has no meaning to embed — with a measured relevance floor, a mandatory visibility filter on every query, and a fail-closed permission model whose failures say which failure occurred. "Found nothing", "denied", and "nothing ingested" must never render as the same output.
Anything presented as a figure computed in code, not by the model. Citations verified programmatically — never by asking the model whether it complied. Extraction gated by reconciliation against the source document's own totals. Then ship to a real user group and watch what they actually ask, which is never what the pilot asked.
Your intended architecture reviewed against these six failure modes, with a written finding and a remediation order. Useful whether or not I build anything.
The ninety days above, delivered with your engineers in the code, not watching. You keep the system, the tests, and the reasoning behind every threshold in it.
Design authority and code review across your client engagements, so the second build is cheaper than the first and the fifth is a product.
Everything I have built runs with no external API dependency. For clients with data-residency, privilege or regulatory constraints, that is not a preference — it is the entry requirement.
Candour
What I do bring that a pure ML hire usually doesn't:
Every figure in this document is a real measurement from a production system, not an illustration. Client names, data and identifying detail are deliberately excluded; the platform described is referred to only by its domain. A full technical walkthrough — architecture, retrieval pipeline, decision trees and operational runbook — is available on request.