Procurement leaders are no longer being asked whether their teams are experimenting with artificial intelligence. They are being asked when those experiments will produce measurable value.
That creates a difficult decision. A CPO can authorise more pilots, extend licences and encourage wider use, creating visible momentum. But the same function may still have fragmented data, inconsistent processes, unclear accountability, weak human-review controls and no reliable baseline against which to measure benefit.
The result is a procurement organisation that appears active with AI without being ready to scale it.
This is not a theoretical distinction. The Hackett Group reported in its 2025 Procurement Key Issues Study that approximately 49% of procurement teams had piloted generative AI use cases, while only 4% reported large-scale deployment. Its 2026 study found deployment progressing rapidly and AI-enabled technology entering procurement’s top three priorities. The direction is clear; the organisational capability required to scale is less evenly developed.
The central argument of this article is therefore simple:
AI activity is not the same as procurement AI readiness.
KOR principle
Readiness is demonstrated when a function can select, govern, integrate, adopt, measure and improve AI in a repeatable way without creating disproportionate operational, commercial or regulatory risk. Pilot volume, licence count and informal usage are weak proxies for that capability.
For most CPOs, the immediate decision is not whether AI has potential. It is whether the organisation has built enough capability to move from isolated experimentation to controlled scale.
How this framework was developed
KOR reviewed procurement-specific AI maturity and readiness models, broader enterprise AI frameworks, responsible-AI guidance and assurance approaches. The recurring strengths were clear: strategy, data, technology, governance, skills and value all matter.
The recurring weaknesses were equally important. Many models use self-reporting without requiring evidence. Culture is often acknowledged but not measured operationally. Third-party AI assurance is frequently treated separately from procurement transformation. Pilot activity can also inflate an overall score even when value and control remain unproven.
The Procurement AI Readiness–Realisation Matrix and Evidence Confidence Model presented here are KOR’s proposed synthesis of those findings. They are practical decision frameworks, not claimed industry standards or statistically validated benchmarks.
What procurement AI readiness actually means
The market often uses adoption, readiness, maturity and realisation as though they describe the same condition. They do not.
| Term | What it describes | Procurement example |
|---|---|---|
| AI adoption | Whether people or systems are using AI | A category manager uses a general-purpose assistant for market research |
| AI readiness | Whether the function has the foundations to deploy and scale AI appropriately | Data ownership, governance, workflow integration, skills and oversight are in place |
| AI maturity | How consistently those capabilities have become repeatable ways of working | Several use cases are governed, monitored and improved through a common operating model |
| AI realisation | Whether AI is producing adopted, measurable outcomes | Contract-review cycle time falls without an unacceptable increase in errors or risk |
A procurement function can be advanced in one dimension and weak in another. It may have enthusiastic users but unreliable data. It may have a comprehensive policy but no embedded workflows. It may own sophisticated tools that few people use. A useful assessment must expose those contradictions rather than averaging them into a reassuring score.
Why pilots can give procurement leaders false confidence
Pilots are useful. They help teams test demand, technical feasibility and user response before committing to wider investment. The problem begins when the existence of a pilot is treated as evidence of readiness.
Tool presence is mistaken for capability
The existence of an AI-enabled platform proves that software is available. It does not prove the procurement function can use it safely or consistently.
Consider a contract-analysis tool. A demonstration may show rapid clause extraction and summarisation. Operational readiness requires more: reliable document ingestion, agreed clause standards, permissions, version control, escalation routes, human-review rules, auditability and clarity about which outputs can inform a decision.
Without those elements, the tool remains an interesting capability around the edge of the operating model.
Self-reporting disguises material differences
Questionnaires are efficient, but respondents interpret maturity differently.
One person may rate data governance as established because a policy exists. Another may reserve that rating for a function with named owners, lineage, quality controls and an active remediation plan. Both answers appear equivalent in a simple score even though the underlying capability is materially different.
The same applies to adoption. Training attendance is easy to report. Regular use of approved tools, confidence in reviewing outputs, managerial reinforcement and changed working practices are harder to demonstrate.
KOR’s view is that a readiness score should state not only the assessed level, but also the confidence that can reasonably be placed in it.
Culture is treated as a communications task
Culture is frequently acknowledged and then reduced to training or change communications.
That is insufficient. Adoption depends on whether people understand where AI is useful, trust the controls around it, feel able to challenge an output and see managers using the tools with discipline. It also depends on whether roles, measures and incentives support the new process.
A team can be technically enabled and behaviourally unready. That is not a problem to solve after deployment. It is part of the scale decision itself.
Third-party AI is assessed outside the transformation
Procurement has two roles in the AI transition. It is a user of AI and a buyer of AI-enabled products and services.
Those roles are often governed separately. They should not be.
The UK Government’s AI Management Essentials tool asks organisations to maintain an inventory of the AI systems they use and to request relevant documentation from third-party providers. That is directly relevant to procurement. A function cannot claim mature AI governance if it controls employee use internally while buying software whose model providers, data practices, subprocessors, update mechanisms or assurance evidence remain poorly understood.
Activity is reported before value is verified
A successful demonstration is not a business case.
Procurement teams often report hours saved, faster analysis or improved insight without defining a baseline, measuring adoption or accounting for the full operating cost of the use case. Data preparation, integration, change support, governance and human review all consume resources.
The relevant question is not whether the system can produce an output. It is whether the organisation can produce a better outcome repeatedly at an acceptable cost and level of risk.
The Procurement AI Readiness–Realisation Matrix
A conventional maturity curve implies one direction of travel. In practice, procurement functions can be highly active without being well prepared, or well prepared before they have produced significant live value.
KOR therefore separates two questions:
- Readiness: How strong are the organisational foundations required to scale AI?
- Realisation: To what extent is AI producing adopted, measurable procurement outcomes?
Together, the two axes create four useful positions.
| Position | Observable evidence | Principal risk | Appropriate CPO decision |
|---|---|---|---|
| Unprepared experimentation | Multiple pilots, informal use, limited telemetry, inconsistent controls and anecdotal benefits | Wider deployment magnifies weak processes and unmanaged risk | Contain the portfolio, establish ownership and substantiate capability before scaling |
| Foundation building | Data, architecture, policies, skills and ownership are being strengthened; live value remains limited | The organisation may over-engineer foundations without testing real use cases | Continue controlled tests while resolving the dependencies that block scale |
| Controlled scaling | Selected use cases are integrated into workflows, governed and producing evidence of adoption and value | Expansion may outpace oversight or operational support | Scale selectively against explicit stage gates and monitoring |
| Embedded value | AI is part of repeatable operations; outcomes, controls and portfolio decisions are reviewed continuously | Complacency, control drift or continued investment in weak use cases | Optimise, retire underperforming applications and reassess as technology and risk change |
The important point is that foundation building is not necessarily less mature than uncontrolled experimentation. A procurement function running fewer pilots may be better positioned to create sustainable value if it is building coherent data, governance and adoption capability.
A worked example: should a contract-intelligence pilot be scaled?
Consider a hypothetical multinational procurement function piloting AI-assisted contract review.
The tool is popular with several category teams. Users report that summaries are faster to produce, and the programme sponsor wants to roll it out across five countries.
The initial evidence sounds encouraging. A readiness assessment, however, finds that:
- Contracts are held across three repositories with inconsistent metadata.
- Clause standards differ by country and category.
- Users copy outputs into email and presentations, so review activity is not logged.
- No baseline exists for review time, accuracy or escalation rates.
- Human review is expected but not defined for higher-risk clauses.
- The supplier’s terms do not yet provide enough clarity on data retention, model providers or material model changes.
- Reported time savings are estimates rather than system or finance evidence.
This is not a failed pilot. It is an unprepared scale decision.
The correct response is not to abandon the tool or approve a five-country rollout. It is to convert the pilot into a controlled learning programme:
- Select a defined contract population and risk profile.
- Establish a baseline for cycle time, accuracy and escalation.
- Standardise the required human-review process.
- Resolve document access, metadata and clause-library issues.
- Obtain and assess supplier documentation and contractual protections.
- Capture usage, exceptions and corrections through telemetry.
- Agree the evidence required for a further scale decision.
At that point, procurement can distinguish user enthusiasm from operational readiness and measured value.
Eight capabilities a procurement AI readiness assessment should examine
The objective is not to create the highest possible average score. It is to identify dependencies, contradictions and decision gates.
1. Strategic intent and operating model
The first question is not which tool to buy. It is what procurement outcome the organisation is trying to improve.
Readiness requires a clear link between AI activity and outcomes such as cycle time, compliance, risk detection, decision quality, working-capital performance or capacity released for higher-value work. Executive sponsorship matters, but so do ownership and decision rights.
Procurement, technology, legal, security, finance and data teams need clarity on who can approve a use case, who owns it after deployment, who monitors it and who is accountable when outputs are wrong.
A long list of ideas is not a strategy. A credible roadmap prioritises a limited portfolio against value, feasibility, risk and organisational capacity.
Decision implication: Do not fund a use case whose business owner, decision rights or intended outcome remain unclear.
2. Data foundations and interoperability
AI amplifies the quality and structure of the information it receives. Procurement information is often distributed across spend systems, contracts, supplier records, ERP platforms, email, shared drives and local spreadsheets.
The assessment should examine data quality, ownership, taxonomy consistency, lineage, access and integration. It should also examine unstructured information because valuable procurement use cases frequently depend on contracts, specifications, supplier correspondence and market documents rather than clean tables.
The practical question is not whether data exists. It is how much manual work is required before it becomes usable, trustworthy and appropriately governed.
Current procurement research continues to identify data quality as a leading barrier to scaling AI. That makes data readiness a portfolio constraint, not merely a technical workstream.
Decision implication: Scale only when the use case has a reliable, governed information supply—not when the organisation claims its data is generally “good enough”.
3. Technology, tooling and orchestration
Procurement teams increasingly access AI through source-to-pay platforms, contract tools, point solutions and general-purpose assistants. The readiness question is how those capabilities fit together.
Useful indicators include workflow integration, API availability, identity and access controls, approved experimentation environments, administration, model or configuration visibility, logging and monitoring. A collection of disconnected tools can increase workload if users repeatedly move information between systems or reconcile conflicting outputs.
The goal is not a perfectly unified platform. It is an architecture in which AI can operate within controlled workflows and produce traceable outputs.
Decision implication: Distinguish a usable feature from an operational service that the function can support and govern.
4. Governance, risk and human oversight
Governance is often reduced to the existence of an AI policy. A policy is necessary, but it is not an operating control.
The NIST AI Risk Management Framework organises AI risk work around Govern, Map, Measure and Manage, with governance intended to operate across the lifecycle. ISO/IEC 42001 similarly treats AI governance as a management system that must be established, maintained and continually improved.
For procurement, readiness requires use-case approval, risk classification, data-protection and security controls, defined human review, logging, incident management and accountability. The level of control should reflect the decision being supported.
Using AI to summarise market reports does not require the same oversight as using it to recommend supplier exclusion, interpret contractual risk or influence an award decision.
Decision implication: Define what the system may recommend, what a human must verify and which decisions it must never make autonomously.
5. Use cases and process redesign
AI should be applied to a defined operational problem, not layered onto an inefficient process because the technology is available.
The assessment should examine how use cases are identified, prioritised, tested and retired. It should establish whether the underlying workflow has been redesigned, what information enters the process, where human judgement is required and how exceptions are handled.
A sourcing assistant that drafts documents more quickly may still create little value if approvals, stakeholder input and evaluation remain fragmented. Conversely, a narrow intake-triage use case may create substantial value if it removes a genuine bottleneck and integrates cleanly into the operating model.
Decision implication: Approve AI against an end-to-end process outcome, not a feature demonstration.
6. People, skills, adoption and trust
Training completion is not adoption.
A useful assessment examines role-specific AI literacy, regular usage of approved tools, confidence in reviewing outputs, quality of challenge, managerial role-modelling and the availability of support. It should also examine resistance, concerns about job change and whether people understand how responsibilities are shifting.
Adoption must be observed at the point of work. A tool can have many registered users and little meaningful use. It can also show high usage because employees are compensating for a poor process rather than improving it.
The GEP and University of Virginia supply-chain readiness research concluded that the differentiator among successful scalers was operational discipline rather than superior technology, with workforce readiness remaining a material gap even among stronger performers.
Decision implication: Treat adoption telemetry, user behaviour and role redesign as evidence—not as post-implementation communications metrics.
7. Vendor ecosystem and third-party AI control
AI procurement requires more than conventional software due diligence.
The assessment should examine the proposed operating model, data flows, model providers, subprocessors, security, audit rights, performance commitments, change notification, intellectual-property exposure, exit arrangements and the supplier’s own governance evidence.
The organisation should also decide what it needs to know when the product changes after contract award. A service may retain the same commercial name while the underlying model, training approach, hosting arrangement or subprocessor chain changes.
The UK Government’s AI assurance guidance explicitly addresses organisations procuring AI systems and the assurance mechanisms available to understand the risks of systems being bought.
Decision implication: Treat supplier AI controls as part of procurement AI readiness, not a separate legal checklist performed after solution selection.
8. Value measurement and scale economics
Every use case needs a baseline and a defined mechanism for validating value.
The relevant measures may include cycle time, cost, compliance, leakage, quality, risk detection, stakeholder experience or capacity released. The measure should match the operational problem.
Benefits must also be adoption-adjusted and assessed against the full cost of operation. Licence costs are only one component. Integration, data remediation, controls, support, human review and failed experiments all affect the scale economics.
Finance involvement improves credibility where benefits influence investment decisions or corporate reporting. It also prevents the programme from accumulating loosely evidenced claims that cannot survive scrutiny.
Decision implication: Set scale, pause and stop criteria before the pilot begins—not after enthusiasm has formed around it.
Evidence confidence: how much should leaders trust the score?
Two procurement functions may receive the same readiness rating while having very different evidence behind it.
KOR therefore proposes assigning an evidence-confidence level alongside each capability score.
| Confidence level | Evidence available | Appropriate use |
|---|---|---|
| Provisional | Self-report or stakeholder assertion only | Early diagnostic and interview planning |
| Substantiated | Self-report supported by policies, process artefacts, contracts, architecture or other documents | Roadmap development and internal prioritisation |
| Verified | Documentary evidence plus operational telemetry, system records, observed controls or finance-validated outcomes | Material investment, scale and benchmarking decisions |
This is not an argument against interviews or surveys. They remain useful for breadth and context. The point is that the assessment must be transparent about what has actually been tested.
An executive saying “adoption is strong” is useful evidence of perception. It is not equivalent to usage telemetry, observed workflow behaviour and output-quality data.
A maturity score is only as credible as the evidence behind it.
Evidence confidence model
How an evidence-based assessment should work
A robust procurement AI readiness assessment should triangulate four evidence sources.
Survey responses
Surveys provide breadth and identify where perceptions differ across leadership, practitioners and enabling functions. They should inform the interview plan rather than determine the final score alone.
Stakeholder interviews
Interviews explain the operating context, reveal conflicting interpretations and identify constraints that may not appear in formal documentation.
Documents and artefacts
Relevant evidence includes strategies, policies, process maps, data-quality reports, architecture diagrams, use-case registers, model or system documentation, supplier contracts, training materials, risk assessments and benefits cases.
Operational and financial evidence
Telemetry, workflow records, quality samples, audit logs, adoption data, exception rates, realised savings and finance validation provide the strongest evidence that capability exists in practice.
The assessment should also record evidence gaps. An absence of evidence is not automatically evidence of low maturity, but it should reduce confidence in the conclusion and affect the scale decision.
What a useful assessment should produce
An attractive radar chart is not enough.
The output should give a CPO a defensible set of decisions:
- The organisation’s position on readiness and realisation
- Capability strengths and blocking dependencies
- Evidence-confidence ratings
- Priority use cases and use cases that should not proceed
- Risk and control gaps
- A sequenced roadmap with accountable owners
- Scale, pause and stop criteria
- Measures for reassessment
- A clear distinction between enterprise constraints and procurement-owned actions
The sequencing matters. A function may not need to remediate every data issue before starting. It does need to resolve the specific data, control and adoption constraints that determine whether a priority use case can operate reliably.
Frequently asked questions
What is procurement AI readiness?
Procurement AI readiness is the function’s ability to deploy and scale AI through reliable data, suitable technology, clear governance, redesigned processes, capable users, controlled suppliers and measurable value. It is not measured by the number of tools or pilots in use.
How is AI readiness different from AI maturity?
Readiness concerns whether the foundations for appropriate deployment exist. Maturity concerns how consistently those foundations have become repeatable operating capability across multiple use cases. A function may be building strong foundations while still having limited live AI maturity.
Does procurement need perfect data before using AI?
No. Requiring perfect enterprise data can become an excuse for inaction. The relevant standard is use-case-specific fitness: the information needed for the proposed application must be sufficiently accurate, accessible, governed and traceable for the decision it will support.
How should procurement measure AI adoption?
Measure adoption at the point of work. Useful evidence includes active usage, completion of relevant workflows, output corrections, exception handling, user confidence, process compliance and whether the tool changes the intended operational outcome. Licence allocation and training attendance are not sufficient.
What should a CPO require before scaling an AI pilot?
At minimum: a defined outcome and baseline, an accountable owner, fit-for-purpose data, workflow integration, risk classification, human-review rules, supplier assurance, adoption evidence, operating-cost visibility and agreed scale or stop criteria.
Conclusion
Procurement does not become AI-ready by accumulating pilots.
It becomes ready when it can make disciplined choices about where AI belongs, provide the information and operating conditions it needs, govern the resulting decisions, earn adoption from the people expected to use it and demonstrate value with evidence that can withstand challenge.
That may mean slowing down a visible programme to resolve a critical dependency. It may mean scaling a narrow use case while rejecting a more fashionable one. It may mean deciding that human judgement should remain central even when greater automation is technically possible.
Those are not signs of low ambition. They are signs that procurement is treating AI as an operating-model capability rather than a collection of software features.
The organisations most likely to create durable value will not be those that can point to the most experiments. They will be those that know which experiments deserve to become part of the way procurement works—and can prove why.