Perspective · 01 of 02Government AI Procurement

Government can outsource the system. It cannot outsource the accountability.

Every AI system a government agency buys arrives with a transfer of capability and no transfer of responsibility. There is a test any agency can apply to any system it is considering, and it is three questions long.

personGus Quiroga
schedule8 min read
calendar_today30 July 2026
Three-question test for government AI procurement — Government can outsource the system, it cannot outsource the accountability

Every AI system a government agency buys arrives with a transfer of capability and no transfer of responsibility. The vendor builds it, hosts it, patches it. The official still holds the delegation. The minister still answers the question. When a determination is challenged three years later, no procurement instrument ever written moves that burden onto the supplier.

Which means there is a test any agency can apply to any system it is considering, and it is three questions long.

An agent accessed a record, drafted advice, or made a determination on an official's behalf.

Fig. 01 — The three-question testEach question thins the field
01 — Was it permitted to?most vendors answer
02 — Under which policy, in force at that moment?the field thins
03 — Can you reproduce that decision exactly, two years from now?very few remain

Question three is the one an agency will actually be asked — by an auditor, a tribunal, a committee, or the person whose payment was stopped. It is also the only one of the three that cannot be added later.

Ask the first and a good part of the market can answer. Ask the third and very few vendors are left in the room.

The evidence, briefly

Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027, and attributes the failures to escalating costs, unclear business value and inadequate risk controls rather than to any limitation of the technology. Governance, not capability, is what breaks these programmes.

McKinsey and Deloitte reach the same conclusion from different samples: roughly a third of organisations have governance maturity adequate for the agents they are already running in McKinsey's 2026 research, and 21% report a mature agentic governance model in Deloitte's. IBM's 2026 breach study adds the operational consequence — 68% of organisations have no governance capable of managing AI or detecting unsanctioned use.

Three independent bodies, one finding: deployment is several years ahead of the ability to account for it. That is the whole of the evidence needed here, and it is not seriously disputed.

40%+

Gartner — agentic AI projects cancelled by end of 2027

~1/3

McKinsey — orgs with governance maturity for the agents they run

21%

Deloitte — report a mature agentic governance model

68%

IBM — no governance capable of managing AI or detecting unsanctioned use

Why procurement misses it

Look at what a government AI evaluation actually tests. Security posture. Data residency and sovereignty. Privacy impact. Accessibility. Functional fit against requirements. Pricing, support, exit provisions. All necessary, all well developed, all refined over decades of buying software.

Now look for the criterion that asks whether the system can reconstruct one of its own decisions. In most tender documentation it is absent — and where something adjacent appears, it asks for audit logging, which is a different thing wearing similar clothes.

If an organisation cannot produce auditable evidence that a sensitive AI output was generated from an authenticated request, using authorised data and under enforced controls, it does not yet have auditable AI governance.

ISACA, 2026 guidance on AI audit trails

Their architectural instruction is blunter still — build for traceability at runtime, rather than attempting to reconstruct trust after the fact.

A log cannot meet that standard, and the reason is structural rather than careless. It is written after the event, by the system that performed it, and it records behaviour rather than entitlement. It will tell you an agent accessed a record. It will not reliably tell you which policy permitted the access, whether that policy was in force at that moment, whether the grant had been revoked an hour earlier, or which model version produced the reasoning.

Recent research on agent-runtime evidence gives the failure a name — replay divergence — and a test: does a stored decision record bind enough properties to re-derive that decision under substituted conditions. Most do not. When a model version moves underneath a system, the log stays identical and the reproduction fails.

That is why question three is worded the way it is. Not "can you show me what happened." Reproduce it. Exactly.

Why the timing is the real problem

None of this would matter much if it were fixable later. It is not.

Procurement fixes architecture. A three-to-five year contract determines, on the day it is signed, what evidence the system will be capable of producing for its entire life. Evidence capability is not a feature that can be switched on in a later release — it is a property of how the execution path was built, and retrofitting it means rebuilding the thing.

So the sequence runs: an agency evaluates on security and functionality, signs, deploys, operates for two or three years, and is then asked to reproduce a decision. At that point the options are to explain that it cannot be done, or to commission a replacement. Neither is a good day. Contractual remedy is no help either — an agency can pursue a vendor and still not produce the evidence, still not un-make the determination, still not answer the committee.

Procurement is the last moment at which this is cheap. It is a paragraph in an evaluation matrix now, and a rebuild later.

What would actually have to be true

If the requirement is that a decision can be reconstructed and defended years afterwards, that imposes specific architectural properties. These are written to be lifted directly into evaluation criteria.

01

Admission before action, in the execution path

The control has to sit where the action passes through it and be able to refuse, not observe and alert. Anything evaluating after the fact is a report, not a control.

02

Fail-closed when the governance layer itself fails

Policy evaluating and denying is handled well nearly everywhere. The harder case is the policy engine unreachable, malformed, or timing out — the condition present during an incident. If the system proceeds and logs a warning, there is no control. Ask for the answer in writing.

03

Policy expressed as versioned code, not documents

A policy in a PDF is a statement of intent. It cannot be enforced automatically, and it cannot be proven to have been in force at a particular moment. "What was the rule at 14:32 on the eleventh" needs a precise, retrievable answer.

04

Authority declared and validated before deployment

Every risk-bearing capability — opening a connection, holding a credential, running an agent, routing to a model — declared explicitly, checked automatically in the build pipeline. Capability that exists but was never declared is where incidents come from.

05

Decision records that re-derive rather than describe

The record must bind inputs, model version, policy version, evaluation result and reasoning trace — enough to reproduce, not merely recount. It should tell the consumer, as a typed field, whether exact reproduction is available for this record or only partial evidence.

06

A cryptographic answer to “what code was running”

Pinned versions and verifiable manifests, so the software state at the moment of a decision is a fact rather than a deployment note. Without this, question three has no floor to stand on.

07

Portability across the estate

Controls that stop at one vendor's boundary govern one vendor's share of the problem. No agency runs a single-ecosystem estate, and governance configuration that does not transfer when a second provider is added is not governance — it is a dependency.

Seven criteria. Any vendor can be asked all seven in a tender, and the answers separate quickly.

The Australian timing

Two obligations converge this December: the remaining requirements under the Commonwealth's policy for the responsible use of AI in government, and the automated decision-making transparency provisions of the Privacy and Other Legislation Amendment Act 2024. Every agency already has a designated accountable official for AI.

It is worth noting how this propagates. Assurance expectations for AI in government are set centrally and cascade outward — which means a procurement pattern established by one agency becomes, in practice, the pattern several others inherit. The three questions above are not a departmental matter. Answered once, at the point where a whole-of-government approach is being defined, they set what every agency downstream is able to evidence.

Meanwhile, contracts being signed this quarter will determine what those accountable officials can produce in 2029.

Organisations that have not designed their accountability model by the end of 2026 risk finding it designed for them instead — by an audit finding, by a regulatory requirement, or by a visible failure.

Deloitte

Designing it deliberately means writing it into what you buy.

This is the first of two pieces. It sets out what has to be true. The second piece looks at what meeting these criteria requires architecturally, and what it looks like when a system has been built for them from the beginning.

GQ

Gus Quiroga

Founder, KYROGA AI. Works with agencies deploying AI agents into decisions that have to withstand scrutiny — from architecture through to the evidence trail.

Talk to us about agent authority →
Share this article