Government can outsource the system. It cannot outsource the accountability.
Every AI system a government agency buys arrives with a transfer of capability and no transfer of responsibility. There is a test any agency can apply to any system it is considering, and it is three questions long.

Every AI system a government agency buys arrives with a transfer of capability and no transfer of responsibility. The vendor builds it, hosts it, patches it. The official still holds the delegation. The minister still answers the question. When a determination is challenged three years later, no procurement instrument ever written moves that burden onto the supplier.
Which means there is a test any agency can apply to any system it is considering, and it is three questions long.
An agent accessed a record, drafted advice, or made a determination on an official's behalf.
Question three is the one an agency will actually be asked — by an auditor, a tribunal, a committee, or the person whose payment was stopped. It is also the only one of the three that cannot be added later.
The evidence, briefly
Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027, and attributes the failures to escalating costs, unclear business value and inadequate risk controls rather than to any limitation of the technology. Governance, not capability, is what breaks these programmes.
McKinsey and Deloitte reach the same conclusion from different samples: roughly a third of organisations have governance maturity adequate for the agents they are already running in McKinsey's 2026 research, and 21% report a mature agentic governance model in Deloitte's. IBM's 2026 breach study adds the operational consequence — 68% of organisations have no governance capable of managing AI or detecting unsanctioned use.
Three independent bodies, one finding: deployment is several years ahead of the ability to account for it. That is the whole of the evidence needed here, and it is not seriously disputed.
40%+
Gartner — agentic AI projects cancelled by end of 2027
~1/3
McKinsey — orgs with governance maturity for the agents they run
21%
Deloitte — report a mature agentic governance model
68%
IBM — no governance capable of managing AI or detecting unsanctioned use
Why procurement misses it
Look at what a government AI evaluation actually tests. Security posture. Data residency and sovereignty. Privacy impact. Accessibility. Functional fit against requirements. Pricing, support, exit provisions. All necessary, all well developed, all refined over decades of buying software.
Now look for the criterion that asks whether the system can reconstruct one of its own decisions. In most tender documentation it is absent — and where something adjacent appears, it asks for audit logging, which is a different thing wearing similar clothes.
If an organisation cannot produce auditable evidence that a sensitive AI output was generated from an authenticated request, using authorised data and under enforced controls, it does not yet have auditable AI governance.
ISACA, 2026 guidance on AI audit trails
Their architectural instruction is blunter still — build for traceability at runtime, rather than attempting to reconstruct trust after the fact.
A log cannot meet that standard, and the reason is structural rather than careless. It is written after the event, by the system that performed it, and it records behaviour rather than entitlement. It will tell you an agent accessed a record. It will not reliably tell you which policy permitted the access, whether that policy was in force at that moment, whether the grant had been revoked an hour earlier, or which model version produced the reasoning.
Recent research on agent-runtime evidence gives the failure a name — replay divergence — and a test: does a stored decision record bind enough properties to re-derive that decision under substituted conditions. Most do not. When a model version moves underneath a system, the log stays identical and the reproduction fails.
That is why question three is worded the way it is. Not "can you show me what happened." Reproduce it. Exactly.
Why the timing is the real problem
None of this would matter much if it were fixable later. It is not.
Procurement fixes architecture. A three-to-five year contract determines, on the day it is signed, what evidence the system will be capable of producing for its entire life. Evidence capability is not a feature that can be switched on in a later release — it is a property of how the execution path was built, and retrofitting it means rebuilding the thing.
So the sequence runs: an agency evaluates on security and functionality, signs, deploys, operates for two or three years, and is then asked to reproduce a decision. At that point the options are to explain that it cannot be done, or to commission a replacement. Neither is a good day. Contractual remedy is no help either — an agency can pursue a vendor and still not produce the evidence, still not un-make the determination, still not answer the committee.
Procurement is the last moment at which this is cheap. It is a paragraph in an evaluation matrix now, and a rebuild later.
What would actually have to be true
If the requirement is that a decision can be reconstructed and defended years afterwards, that imposes specific architectural properties. These are written to be lifted directly into evaluation criteria.
Admission before action, in the execution path
The control has to sit where the action passes through it and be able to refuse, not observe and alert. Anything evaluating after the fact is a report, not a control.
Fail-closed when the governance layer itself fails
Policy evaluating and denying is handled well nearly everywhere. The harder case is the policy engine unreachable, malformed, or timing out — the condition present during an incident. If the system proceeds and logs a warning, there is no control. Ask for the answer in writing.
Policy expressed as versioned code, not documents
A policy in a PDF is a statement of intent. It cannot be enforced automatically, and it cannot be proven to have been in force at a particular moment. "What was the rule at 14:32 on the eleventh" needs a precise, retrievable answer.
Authority declared and validated before deployment
Every risk-bearing capability — opening a connection, holding a credential, running an agent, routing to a model — declared explicitly, checked automatically in the build pipeline. Capability that exists but was never declared is where incidents come from.
Decision records that re-derive rather than describe
The record must bind inputs, model version, policy version, evaluation result and reasoning trace — enough to reproduce, not merely recount. It should tell the consumer, as a typed field, whether exact reproduction is available for this record or only partial evidence.
A cryptographic answer to “what code was running”
Pinned versions and verifiable manifests, so the software state at the moment of a decision is a fact rather than a deployment note. Without this, question three has no floor to stand on.
Portability across the estate
Controls that stop at one vendor's boundary govern one vendor's share of the problem. No agency runs a single-ecosystem estate, and governance configuration that does not transfer when a second provider is added is not governance — it is a dependency.
Seven criteria. Any vendor can be asked all seven in a tender, and the answers separate quickly.
The Australian timing
Two obligations converge this December: the remaining requirements under the Commonwealth's policy for the responsible use of AI in government, and the automated decision-making transparency provisions of the Privacy and Other Legislation Amendment Act 2024. Every agency already has a designated accountable official for AI.
It is worth noting how this propagates. Assurance expectations for AI in government are set centrally and cascade outward — which means a procurement pattern established by one agency becomes, in practice, the pattern several others inherit. The three questions above are not a departmental matter. Answered once, at the point where a whole-of-government approach is being defined, they set what every agency downstream is able to evidence.
Meanwhile, contracts being signed this quarter will determine what those accountable officials can produce in 2029.
Organisations that have not designed their accountability model by the end of 2026 risk finding it designed for them instead — by an audit finding, by a regulatory requirement, or by a visible failure.
Deloitte
Designing it deliberately means writing it into what you buy.
This is the first of two pieces. It sets out what has to be true. The second piece looks at what meeting these criteria requires architecturally, and what it looks like when a system has been built for them from the beginning.
Gus Quiroga
Founder, KYROGA AI. Works with agencies deploying AI agents into decisions that have to withstand scrutiny — from architecture through to the evidence trail.
Talk to us about agent authority →Sources
- Gartner — Predicts over 40% of agentic AI projects will be canceled by end of 2027
- McKinsey — State of AI trust in 2026: shifting to the agentic era
- Deloitte — State of AI in the Enterprise, 2026
- Deloitte — AI agents are scaling faster than their guardrails
- IBM — Cost of a Data Breach Report 2026
- ISACA — The AI audit trail: from AI policy to AI proof
- DEMM-Bench — a cross-regime benchmark for agent-runtime governance-evidence sufficiency
- Australian Government — Policy for the responsible use of AI in government
- Australian Government, Department of Finance — National framework for the assurance of artificial intelligence in government
