The Assurance Engine: Teaching a Robot to Do My Compliance Homework (Without Leaking the Answers)
Somewhere in the last couple of years I became the sort of person who gets genuinely excited about a well-behaved checkpoint file. This is not a personality trait I asked for. It arrived, uninvited, after one too many evenings spent doing the least glamorous job in security: sitting with a small mountain of evidence in one window and a control framework in the other, working out by hand whether document A satisfied requirement B or merely gestured politely in its general direction. Do that across a few hundred controls, each one written by a committee apparently paid by the ambiguity, and something in you quietly gives way. Usually the part that used to enjoy weekends.
This is the story of a side project Iâve been tinkering with, on and off, for the best part of a year: a tool Iâll call the Assurance Engine, because the real name is boring and also none of your business. It takes a big pile of evidence documents (policies, configs, screenshots, the sort of thing an auditor requests and you dread producing), points a language model at each requirement in a security control framework, and drafts an assessment of whether the evidence actually supports the claim. Think of it as a very diligent, slightly literal-minded graduate who has read every document you gave it and never gets bored, resentful, or distracted by their phone.
It sounds like a small thing. It wasnât. It turned out to be an excellent way to spend a year accidentally relearning why all the âboringâ engineering practices exist in the first place.
the bit where I have to defend the entire premise
Letâs get the obvious objection out of the way, because someone always raises it at this point in the conversation (thanks, Phil): âCouldnât you just chuck the PDFs into a general-purpose AI chat tool and ask it questions?â
Yes. Obviously. You could also cut your own hair with kitchen scissors. Itâs technically possible and someone, somewhere, has definitely done it successfully. The point isnât whether the capability exists; every large model going can read a PDF and have an opinion about it. The point is what happens the other 99% of the time, when a stressed person is doing this at 4pm on a Friday with forty documents and no time to be careful.
What actually happens, in my experience, looks like this: someone emails a document to their personal account so they can open it at home. They upload it to whichever AI tool is free that week. They download the response, tidy it up in Word, and send it over chat to someone else for review. Multiply that by every control in a decent-sized framework, and youâve got sensitive material fanned out across four devices, three cloud accounts, and at least one browser extension named something like âPDF Wizard Proâ that nobody has ever properly vetted. Nobody did anything malicious. Nobody even did anything against policy, necessarily. Itâs just what happens when a useful capability isnât wrapped in a process.
Thatâs the actual product, if Iâm honest with myself. Itâs not âAI reads documentsâ; any of a dozen tools do that today. Itâs âthere is exactly one door in, one door out, and everything that happens to your evidence in between is logged, scoped, and cleaned up afterwards.â Unglamorous. Extremely necessary. Nobody puts âconsistent, boring, well-audited data handlingâ on a conference slide, but itâs the whole reason the tool is worth building rather than just telling everyone to be more careful with ChatGPT.
what itâs actually good at (hint: not thinking)
I want to be careful here, because it would be very easy, and very dishonest, to write the next bit as âand then AI replaced the assessor, the end.â It didnât, and if it had, Iâd be worried about the assessor.
What the tool is really good at is the mundane part: reading forty pages of evidence to find the one paragraph thatâs relevant to a specific control, cross-referencing it against the requirement, and drafting a first pass at âis this adequate, weak, or completely absent.â Thatâs real, thankless work: the kind that eats a junior assessorâs entire week and leaves them cross-eyed by Thursday. Handing that first pass to a human reviewer, who then spends their time on judgement calls rather than page-flipping, is a proper step up in how the work gets done.
What it is not good at, and what I had to build proper engineering around rather than just trusting the model to notice, is knowing when it has failed. Early on I had a demo (the sort you really donât want to go wrong in front of an audience) where the assessment âcompleted successfullyâ and produced a report that was quietly, comprehensively empty. Every single evidence lookup had failed upstream (an infrastructure hiccup, unrelated to the model itself), and the batch processor, in a moment of what I can only call misplaced optimism, caught the error, wrote a placeholder answer of âprocessing errorâ into every single result, and reported the whole run as a success. Nobody lied. The code just wasnât built to distinguish between âthe AI concluded thereâs no evidenceâ and ânothing actually got asked.â I had, in effect, taught a very confident intern to say âlooks fine to me!â regardless of whether theyâd read anything at all.
That one stung, and it taught me the real lesson of this whole project: the AI is the bit thatâs allowed to be uncertain. The plumbing around it is not allowed to be uncertain. If zero requests succeed, the system must fail loudly, not produce a beautifully formatted document full of nothing.
the year I learned to love a checkpoint
The single most expensive lesson came from an assessment covering just over a thousand controls, running on infrastructure that treated the containerâs local disk as, essentially, a sandcastle at high tide. Progress files, checkpoints, half-finished results: all of it lived in storage that could vanish the instant the underlying instance got recycled. One did. Roughly forty minutes and a properly uncomfortable amount of API spend disappeared with it, and the retry, because of course there was a retry, didnât resume from where it left off. It started again from batch zero, quite happily paying for the same answers twice, because the checkpoint had never been written in a form that survived out-of-order completions.
Fixing it properly meant admitting that âcall an AI model in a loopâ is a distributed systems problem wearing a trench coat. It needed a checkpoint keyed by which specific batch had finished (not just a running count, because concurrent batches finish out of order and a naive counter will happily skip an unfinished one and re-run a finished one). It needed durable storage that didnât evaporate when a container did. It needed a lease so that if a job died mid-run, its replacement could tell the difference between âthe previous attempt is still alive, back offâ and âthe previous attempt is dead, take over now.â Get that wrong and you either get two workers cheerfully re-doing the same expensive work in parallel, or a run that sits at zero forever because everyoneâs too polite to take the lease off a corpse. Digital manners, it turns out, are a remarkably efficient way to bankrupt yourself.
None of that is exciting. None of it will ever appear in a demo. All of it is the difference between a toy and a tool youâd trust with real money and real client data.
logs remember everything, whether you asked them to or not
The security lessons were, almost without exception, not about the AI at all. They were about everything wrapped around it, which is exactly where youâd expect the mistakes to hide.
The API key, for instance, started life being passed to a subprocess as a plain command-line argument, which on most systems means itâs sitting in the process list for any other process on the box to read. That got moved to a short-lived temporary file, written just before use and deleted in a finally block immediately after, regardless of whether the run succeeded or blew up. A small change, obvious in hindsight. The sort of thing you only notice once youâve specifically gone looking for âwhere does this secret exist as plain text, and for how long.â
Error logging had its own version of the same problem: a failed API call is exactly the moment you most want a verbose log message, and exactly the moment youâre most likely to log something you shouldnât: a full response body, an oversized payload, a chunk of evidence that got attached to the request. Every error path ended up with a hard character cap and explicit truncation, because âlog everything for debuggingâ and âdonât accidentally persist someoneâs sensitive document into a log aggregatorâ are in permanent tension, and the log always wins if youâre not paying attention.
And then thereâs the boring-but-vital stuff: uploaded evidence lives in per-run directories with randomised names, gets deleted the moment the run finishes (success or failure), and result files get signed with an HMAC the moment theyâre written, so a tampered report doesnât quietly get served back to someone as if it were the original, the paperwork equivalent of swapping the exam answers after theyâve already been marked. Isolation between one clientâs assessment and anotherâs isnât a policy document; itâs a directory boundary, a unique ID, and a cleanup step that runs even when everything else goes wrong. Thatâs the actual âAI securityâ work, and approximately none of it involves the AI.
the tour through models, prompts, and an increasingly opinionated editor
Underneath all of this, the tooling I was using to build it kept changing under my feet, which is its own small comedy. I started stitching prompts together by hand, testing them against whatever model was current at the time, long before anything resembling an âagentâ existed in my editor. Then the assistants got better at holding context across a whole file. Then across a whole project. Then they started proposing edits, running my tests, and occasionally getting slightly too enthusiastic about refactoring something I hadnât asked it to touch (a habit Iâve had to explicitly tell it to knock off, repeatedly, in writing).
The shift that actually mattered wasnât âthe model got smarter,â although it did. It was the tooling around the model catching up to actually understanding a codebase rather than a single pasted snippet, the same theme as the whole project, really. The clever bit was never in short supply. The workflow around it always was.
so, was it worth it
Yes, unreservedly, though not for the reason I expected when I started. I thought I was building something to make AI answer audit questions. What I actually built was a small case study in how much unglamorous engineering it takes to make an exciting capability boring, reliable, and safe enough to hand to someone who isnât going to read the source code before they trust it.
The AI was never really the hard part. The hard part was making sure that when it got something wrong, everyone would notice, and that when it got something right, nobodyâs data had gone on an unplanned tour of four devices to get there.