There is a particular kind of silence that falls over a governance meeting when someone asks a simple question: Can we show the regulator exactly how that decision was made?
Not in theory. Not in principle. Not by referencing the vendor's documentation or the model card filed at procurement. Can you show — with evidence, in sequence, traceable to a specific output on a specific date — how your live AI system arrived at a consequential decision that has already affected a real person?
For most regulated organisations, the honest answer is no. And the regulator is already asking.
The Audit Request That Exposed the Gap
The scenario is no longer hypothetical. Across financial services, insurance, healthcare, and the public sector, regulatory bodies are beginning to exercise their powers in ways that assume AI accountability frameworks are already in place. In many cases, they are not.
Consider what an audit request actually demands. It is not a forward-looking question about what your AI policy says you will do. It is a backward-looking demand: show us what your system did, when it did it, what data it used, what weight it gave to which factors, and how that produced the output that affected this individual. Show us the chain of custody for that decision.
Organisations that deployed AI systems — sometimes years ago, sometimes under different leadership, sometimes through a third-party vendor whose contractual obligations were never fully interrogated — are discovering that they cannot reconstruct this chain. The system is live. The decisions have been made. The audit trail does not exist.
This is not a failure of intent. Most organisations that deployed AI did so with genuine governance ambitions. They had ethics principles. They ran bias assessments. They engaged legal counsel. What they did not build — or did not insist their vendors build — was the infrastructure of retrospective accountability: the logging architecture, the versioning records, the explainability outputs, the human oversight documentation that would allow anyone, months or years later, to answer the regulator's question.
The gap between what organisations believed they had and what they can actually produce under scrutiny is where the accountability crisis lives.
Why Live AI Systems Create Retrospective Accountability Failures
There is a structural reason why live AI systems are particularly prone to retrospective accountability failures, and it has nothing to do with bad faith. It has to do with the difference between how AI systems are built and what accountability frameworks actually require.
AI systems — particularly those using machine learning — do not make decisions the way a human decision-maker does, and they do not record their reasoning the way a human decision-maker would. A credit officer who declines a loan application can, in principle, articulate the factors that informed that judgement. A gradient boosting model that produces the same output cannot narrate its own reasoning. Explainability tools can approximate it — SHAP values, LIME outputs, attention weights — but these are post-hoc analytical constructs, not records of the decision process itself.
This matters enormously for regulated organisations, because accountability frameworks — whether under GDPR's Article 22, the FCA's Consumer Duty, or the EU AI Act's requirements for high-risk systems — do not ask for approximations. They ask for explanations that are meaningful, specific, and capable of being challenged.
Live AI systems create retrospective accountability failures in at least three ways. First, they are often updated continuously, meaning that the model that made a decision six months ago may no longer exist in the form that made it — and if model versioning was not rigorously maintained, that version may be gone entirely. Second, the input data that produced a specific output is rarely preserved in a form that links it to that specific output. Logs exist, but they are not always structured to support retrospective reconstruction. Third, the human oversight layer — where it exists — is often undocumented. A human reviewer may have approved or overridden a model output, but if that review was not recorded with the right granularity, it cannot be evidenced.
The result is that organisations are operating systems whose decisions are, in a meaningful sense, already lost to history — even as those decisions continue to generate regulatory and legal exposure.
The Explainability Illusion: What Regulated Organisations Actually Have on File
One of the most consequential gaps in current AI governance practice is the confusion between having an explainability capability and having explainability evidence. These are not the same thing, and regulated organisations are discovering the difference under pressure.
Most organisations that have taken AI governance seriously have, at minimum, conducted some form of explainability assessment at the point of deployment. They have reviewed model outputs, examined feature importance, and satisfied themselves that the system's behaviour is broadly interpretable. Some have gone further, commissioning third-party audits or engaging with explainability toolkits. This is not nothing. But it is not an audit trail.
What regulators and courts increasingly want is not evidence that a system was explainable in principle at the point of deployment. They want evidence that a specific decision, made at a specific time, can be explained in terms that are meaningful to the individual affected by it. That requires something very different: it requires that explainability outputs were generated and preserved at the point of each consequential decision, or that the system's state at that point can be reconstructed with sufficient fidelity to generate those outputs retrospectively.
Very few organisations have this. What they have on file is typically a combination of: the original model documentation, a pre-deployment bias assessment, a data protection impact assessment that addresses the system in general rather than individual decisions, and perhaps some internal explainability analysis that was never designed to be reproduced at scale or on demand.
This is the explainability illusion: the comfortable belief that because the system was assessed as explainable, it can be explained. The assessment happened once. The decisions are happening continuously. The gap between the two is where accountability breaks down.
Where AI Governance Advisory Breaks Down After Deployment
The market for AI governance advisory has grown significantly over the past three years, but it has grown in a particular direction: toward pre-deployment. Ethics frameworks, bias assessments, model cards, impact assessments, responsible AI principles — the majority of advisory effort, and the majority of what passes for AI governance advisory in the market, is oriented toward the point of deployment or before it.
This is understandable. Pre-deployment is where the decisions that shape a system's behaviour are made. It is where the leverage is highest. It is also, frankly, where advisory work is most comfortable — because it deals in recommendations rather than evidence, in frameworks rather than specific decisions, in what should happen rather than what did.
But regulated organisations do not only face forward-looking governance challenges. They face backward-looking ones. And the AI governance advisory market has not kept pace with the retrospective accountability demands that live AI systems now create.
The breakdown manifests in several specific ways. Governance teams that have engaged advisory support — sometimes extensively — find that their advisors are not equipped to help them respond to a specific regulatory inquiry about a specific decision. The frameworks are there. The policies are there. The evidence is not. When the regulator asks for the audit trail, the governance team discovers that their advisory relationship was built for a different kind of problem.
There is also a seniority problem. AI governance advisory in many organisations is delivered at a level that lacks the authority to drive systemic change in how decisions are logged, how models are versioned, how human oversight is documented. The advice is technically sound but organisationally ineffective — it does not reach the people who control the engineering decisions, the vendor relationships, the data architecture. The result is governance that exists on paper but does not exist in the systems.
This is where AI governance advisory breaks down: not at the level of principles, but at the level of implementation authority and retrospective accountability infrastructure.
What a Defensible Audit Trail Actually Requires
Building a defensible audit trail for live AI systems is not a theoretical exercise. It requires specific, practical infrastructure — and it requires that infrastructure to be built into the systems themselves, not bolted on after the fact.
A defensible audit trail for a regulated AI system has at least five components.
Decision-level logging. Every consequential output must be logged at the point of generation, with sufficient context to reconstruct the decision: the input data used, the model version that produced it, the timestamp, and — where applicable — the human review action taken. This logging must be structured, preserved, and accessible. It is not the same as general system logging.
Model versioning and provenance. The model that produced a decision must be identifiable and, where necessary, reproducible. This means maintaining a version history with sufficient granularity to identify which model was live at which point in time, and what changes were made between versions. For systems that update continuously, this is a non-trivial engineering requirement.
Explainability output preservation. Where explainability tools are used, their outputs must be generated and preserved at the point of decision — or the system must be capable of reproducing them with sufficient fidelity for specific historical decisions. This requires a deliberate decision about which explainability methodology is used, and a commitment to preserving its outputs as part of the audit trail.
Human oversight documentation. Where human reviewers are part of the decision process — whether as approvers, overriders, or escalation points — their actions must be documented in a way that links them to specific model outputs. A general record that human oversight exists is not sufficient. The specific human action taken in relation to a specific decision must be traceable.
Data lineage. The data that was used to train the model, and the data that was used as input to a specific decision, must be traceable. This is particularly challenging for organisations that use third-party data, that have complex data pipelines, or that have updated their training data over time. But it is a requirement that regulators will apply, and organisations must be able to meet it.
Building this infrastructure in systems that are already live is harder than building it at the point of deployment — but it is not impossible. It requires a systematic audit of existing logging and versioning practices, identification of the gaps between what is currently captured and what a defensible audit trail requires, and a prioritised programme of remediation. It also requires someone with sufficient seniority and technical authority to drive that programme to completion.
The Senior Advisory Deficit and Who Is Left Holding the Risk
The central governance problem facing regulated organisations with live AI systems is not a lack of awareness that accountability is required. It is a lack of senior advisory capacity — people with the technical depth to understand what a defensible audit trail requires, the regulatory knowledge to understand what regulators will demand, and the organisational authority to drive the changes needed.
This senior advisory deficit is structural. The AI governance function in most regulated organisations sits below the level at which it can mandate systemic change. It can produce policies. It can commission assessments. It can write frameworks. What it cannot always do is compel engineering teams to change their logging architecture, compel procurement to renegotiate vendor contracts, or compel the board to treat retrospective accountability as a risk that requires immediate investment.
The result is that risk accumulates at the governance level without being resolved. Governance teams know there is a problem. They know the audit trail is insufficient. They may not know exactly what a defensible audit trail requires, or they may know but lack the authority to mandate it. They are holding risk they cannot quantify, because quantifying it would require a level of technical analysis that is itself beyond the capacity of most governance functions.
This is not a failure of individual governance professionals. It is a structural failure in how organisations have resourced AI governance — treating it as a compliance and policy function rather than as a senior advisory function with real authority over technical and commercial decisions.
The organisations that are best positioned to navigate retrospective accountability demands are those that have senior AI governance advisory capacity operating at the level of the CRO, the General Counsel, or the board — people who can translate technical accountability requirements into concrete engineering and commercial decisions, and who have the authority to ensure those decisions are implemented. The EU AI Act's requirements for high-risk AI systems provide a useful benchmark for the level of logging, human oversight, and transparency documentation that regulators across jurisdictions are increasingly likely to expect.
For organisations that do not have this capacity internally, the option is external senior advisory support — not frameworks and principles, but people who can audit the existing audit trail, identify the gaps, and drive a remediation programme with the seniority and specificity the problem requires.
The risk that remains unaddressed does not sit with the AI system. It sits with the organisation, with the individuals who signed off on deployment, and — in the absence of a defensible audit trail — with the governance team that is left to explain decisions that the organisation cannot actually explain.
The regulator is already asking. The question is whether the organisation is in a position to answer — and whether it has the senior advisory support to get there before the next request arrives.
Navitec AI provides senior AI governance advisory support to regulated organisations navigating live deployment accountability, regulatory inquiry, and audit trail remediation. If your organisation is facing retrospective accountability challenges, we work with governance, legal, and technology functions to build the evidentiary infrastructure that defensible AI accountability requires.