Agentic Procurement Pilots at Large Manufacturers

Forty-nine percent of procurement teams are running agentic AI pilots; four percent have reached meaningful deployment. That gap is not a technology problem. It is an execution discipline problem, and the organizations that understand that distinction are the ones building durable operational advantage while everyone else runs a second pilot.
The Art of Procurement's 2026 State of AI in Procurement and the Hackett Group's 2026 Procurement Key Issues Study describe the same bottleneck from different vantage points. Hackett finds 43% of organizations actively pursuing AI deployment, with only 12% at large-scale implementation. Neither figure suggests the technology is immature. The use cases that do reach production are already generating auditable, verifiable results. The constraint is not what the agents can do; it is what the organizations around them are prepared to support.
What Agentic Procurement Actually Does That Traditional Automation Could Not
Traditional automation executes a defined task when a human triggers it. An agentic system perceives a state, selects an action, executes it, and loops, all without waiting for instruction at each step. For manufacturers, the relevant shift is from automating individual tasks to orchestrating multi-step decisions across systems and stakeholders.
Production deployments already exist across the full source-to-pay map, and their results are specific enough to be examined.
Walmart's deployment of Pactum AI for tail-spend negotiation is the most documented case in public circulation. Agents authorized to negotiate within pre-set parameters, operating across more than 2,000 suppliers, achieved a 68% supplier agreement rate against an original target of 20%, with 3% average cost savings and an 11-day average negotiation turnaround. Those are not pilot-phase projections; they are production figures from a live, ongoing deployment.
McKinsey's February 2026 reporting describes a technology company using linked agents to integrate spend data and simulate demand scenarios, identifying 12 to 20% savings in contact center operations and 20 to 29% in BPO and financial services spend. A separate McKinsey case from the same period describes an aircraft OEM using agents to automate order execution against production planning data, cutting active inventory by 30% and contributing roughly $700 million in EBIT improvement. A luxury automotive OEM working with Deloitte and ServiceNow achieved a 400% increase in supplier onboarding speed, migrating more than 2,000 Tier-1 suppliers to a new EDI system.
The deepest production footprint sits in accounts payable. Per Hackett research cited by SAP, 21% of companies are already running agentic AI in AP in production environments; that figure is higher than any other procurement domain, and the reason why matters enormously for understanding what makes scale possible.
The common thread across every working deployment is this: agents operate within a bounded decision space, with explicit human-set parameters. Not open-ended autonomy. Walmart's agents negotiate, but they negotiate within ranges that human managers defined before the first conversation with a supplier. The McKinsey OEM's agents execute orders, but against production planning data with clear escalation logic for exceptions.
The Walmart case also reveals something instructive about supplier acceptance. Eighty-three percent of suppliers rated the system easy to use, and 75% said they preferred negotiating with the AI over a human. Those numbers reflect deliberate design investment, not an accidental outcome. The technology worked because the implementation was built with the supplier experience as a first-class consideration, not an afterthought.
Where Pilots Actually Break Down: The Execution Gaps That Block Scale
McKinsey observed in February 2026 that organizations can move from prototype to pilot in weeks and from pilot to scale in under a year — with the right foundation. However, most organizations lack the right foundation, and the pilot phase is where that gap becomes invisible rather than addressed.
The failure modes cluster into five categories, and they are interdependent. Solving one in isolation rarely moves an organization past the bottleneck.
Data readiness is where most pilots quietly die. Agentic systems depend on clean, accessible, consistently structured data across ERP, supplier master, and contract systems. Large manufacturers typically have fragmented data estates accumulated through decades of M&A and system patchwork. A pilot appears to work because a small dataset has been manually cleaned for it; that discipline does not exist at enterprise scale, and no one in the pilot phase is incentivized to surface that fact.
Governance and authorization design is the second cluster. The Walmart deployment's agents operate within pre-defined business parameters, which sounds simple until you try to replicate it. Creating those parameters requires procurement leadership to make explicit, documented decisions about what the agent is and is not authorized to do, and those decisions require cross-functional alignment across legal, finance, and supply chain. A pilot team can sidestep that alignment; a production deployment cannot.
Change management for procurement staff is consistently underestimated. The role shift from executing tasks to supervising agents is structurally significant, not cosmetic. Instead of treating it as a training exercise, organizations need to redesign how procurement roles function; those that do not tend to see agents bypassed or quietly undermined by the staff who were supposed to oversee them.
Integration complexity is the fourth cluster, and it is the one most deliberately deferred during pilots. A pilot can run alongside existing processes without formally connecting to them; scaling requires replacing or integrating those processes, which surfaces ERP integration complexity that pilot teams have every incentive to postpone.
Scope creep from stakeholder enthusiasm is the fifth failure mode, and it is the most insidious because it arrives dressed as success. A pilot that works generates internal demand for expansion before the foundational issues above are resolved; the result is a broader but shallower deployment that achieves neither the coverage nor the reliability of a true production system.
The ROI divergence this creates is quantifiable. Deloitte's 2025 Global CPO Survey found that "Digital Masters," the top quartile of procurement organizations, achieve a 3.2x return on GenAI initiatives, while "Followers" achieve 1.5x. That spread is not attributable to which technology they selected; it is attributable to how they built around it.
The 4% who reach meaningful deployment treated the pilot as a discipline test. The 49% who remain in pilot phase treated it as a technology test.
What the Manufacturers Who Scaled Did Differently in the Pilot Phase Itself
The Walmart Pactum deployment offers the clearest anatomy of a pilot designed to scale. It was scoped to a specific, high-volume, low-complexity category: tail-spend supplier negotiations where decision parameters were straightforward to define and the cost of an agent error was low. It was not a showcase of maximum capability; it was a deliberate stress test of governance and supplier acceptance in a domain where failure was recoverable.
That choice of scope was strategic, not modest. The organizations that remain stuck in pilots tend to select use cases for their strategic ambition rather than their governance tractability.
McKinsey's February 2026 reporting on a chemicals company piloting autonomous sourcing in consumables illustrates the same logic from a different angle. The category was chosen because the tender process was already well-documented. Agents were automating preparation and analysis within existing workflows, not inventing new ones. The leakage reduction came from consistency in execution, not from sophisticated AI reasoning; the pilot was designed to prove a replicable process, not to demonstrate a capability ceiling.
Across the deployments that successfully scaled, several design choices appear consistently:
A single category or use case with a defined decision boundary established before launch. Not "agentic AI for procurement broadly" but a specific, bounded problem with explicit parameters.
An authorization framework written before the pilot starts. What the agent can decide autonomously, what triggers human escalation, what constitutes an error. This document should not exist as shared understanding; it should exist as a signed document.
Procurement staff involved in designing agent parameters, not merely informed of them after the fact. This serves dual purposes: it produces better-calibrated parameters, and it functions as the most effective form of change management available, because operators who helped design the system are unlikely to undermine it.
Supplier-facing deployments built with supplier experience as a design input from day one. The 75% preference-for-AI figure from Walmart's supplier base reflects investment in clarity, simplicity, and human escalation paths, not a windfall of goodwill.
A defined handoff protocol specifying what happens when the agent encounters a situation outside its parameters, and who owns that handoff.
Siemens and SAP's movement toward unified multi-agent supply chain architectures, reported by Evolving Market Research in May 2026, represents an evolution well beyond single-use-case pilots. Both organizations built that architecture on top of years of foundational data and integration work; the multi-agent architecture was the destination reached by organizations that had already done the unglamorous work.
The pattern is consistent: organizations that scale started narrower, documented more, and involved procurement operators earlier than the organizations still piloting two years later.
The Pressure That Makes Doing This Slowly Feel Impossible — and Why That's the Trap
The Hackett Group's 2026 study projects that procurement workloads will increase 8% in 2026 while headcount and operating budgets decline. Eighty percent of procurement executives identify AI-enabled technology as the most transformational trend affecting the function over the next five years. The urgency is real, not manufactured by vendors.
The trap is precisely this: resource pressure creates a powerful incentive to rush the governance and data preparation work that makes scale possible, because that work is invisible to stakeholders who want to see the agent running. A pilot can be launched faster if the data audit is abbreviated. The authorization matrix can be simplified into a verbal agreement. The operator design loop can be skipped because it surfaces uncomfortable questions about role changes and slows the timeline.
Every one of those shortcuts produces a pilot that appears to work and a production deployment that never arrives.
Consider a procurement leader — call her Maria — who joined a mid-size industrial manufacturer in early 2025. She inherited an agentic AI pilot that had been running for eight months in strategic sourcing. The demo was polished. The internal Slack channel dedicated to it had 200 members. However, when Maria pulled the data, the agent was handling 3% of intended transaction volume, and the supplier master it depended on had not been updated since the previous fiscal year. The pilot had been optimized for the presentation, not the process. So Maria shut it down, spent four months on data remediation, and relaunched in accounts payable. Eighteen months later, that AP deployment was in full production.
Accounts payable has the highest production deployment rate of any procurement domain, at 21% of companies, precisely because the process is the most structured, the decision parameters are the most explicit, and the data is the most standardized. It was the easiest to govern, not the most strategically compelling; organizations that chose AP as their first production domain were being calibrated, not unambitious.
The organizations chasing high-visibility use cases — autonomous strategic sourcing, complex multi-category supplier negotiation — before they have validated their data and governance posture in a lower-stakes domain are optimizing for the announcement rather than the outcome.
Ivalua's 2026 analysis found that organizations scaling AI beyond pilots see 3 to 5 times higher ROI than those running isolated use cases. The compounding return comes from getting the foundation right; it does not come from moving first.
What a Production-Ready Pilot Program Looks Like in Practice
Four components appear consistently across pilots that made the transition to production. They are not novel; they are the components most frequently omitted.
Scope selection based on governance tractability, not savings potential. Choose the use case where the decision logic is already documented or can be made explicit quickly. Tail spend and accounts payable are natural first domains. Autonomous strategic sourcing is a second-phase objective, reached after the foundational work is validated. The largest theoretical savings opportunity is rarely the right starting point.
A data audit before agent deployment. Identify the three to five data sources the agent will depend on. Assess their completeness, freshness, and accessibility. Fix the gaps before launch, not after. This step consistently takes longer than expected, and it kills more pilots than any technology failure; organizations that discover their supplier master data is inconsistent after the agent is live have created a situation where the agent cannot be trusted, and the pilot quietly stalls.
A written authorization matrix. Not a shared understanding. A document specifying exactly what decisions the agent can make autonomously, what it must flag for human review, and what it cannot do. This document should require sign-off from legal, finance, and the relevant category lead before the pilot launches; if obtaining those signatures surfaces disagreements about scope, that is valuable information to have before the agent is running.
An operator design loop. Procurement staff who will supervise the agent should review and approve the agent's decision parameters before launch. Their feedback should be formally incorporated, not noted and deferred. This step is routinely skipped because it slows the timeline; the organizations that skip it tend to find that operators route around the agent in practice, which produces data suggesting the agent is underperforming when the actual problem is adoption design.
Beyond these four components, a production-ready pilot requires a specific, measurable production threshold defined before launch. Not "successful pilot," but a defined percentage of transactions handled autonomously, at a defined accuracy level, over a defined period. Pilots declared successful without testing scalability conditions are demonstrations, not pilots.
For any supplier-facing agent, a supplier communication strategy is not optional. After all, Walmart's 83% ease-of-use rating reflects a deliberate investment in context, options, and human escalation paths; that result does not occur when supplier communication is treated as a footnote.
McKinsey's framing holds: the path from pilot to production can be covered in under a year, assuming the data and governance work runs in parallel with, not after, the pilot.
How the 88% ROI Figure Changes the Case for Moving Carefully
Among companies already running agentic AI in production, 88% report measurable ROI, compared to 74% of broader GenAI adopters, per Payhawk's May 2026 analysis. That 14-point gap is not attributable to technological sophistication; it is attributable to operational discipline. The production-deployed group completed the unglamorous foundation work, and the ROI advantage accrues to that completion, not to early adoption per se.
The adoption trajectory makes the competitive context clear. Generative AI adoption in procurement moved from 50% to 94% between 2023 and 2024, per AI at Wharton research cited in SupplyChainBrain in April 2026. Procurement is already the leading enterprise function for AI adoption. So the competition is not about who begins using AI; it is about who moves it into production, and the organizations building that production infrastructure now are establishing an advantage that compounds with each subsequent use case.
The question for a CPO today is not whether to pilot agentic AI. Forty-nine percent already are; the question is whether the current pilot is designed to become a production system. The answer is visible long before the pilot concludes — in whether the data audit exists before the first agent goes live, whether the authorization matrix has signatures on it, and whether the procurement operators who will supervise the system helped design its parameters. Those are the questions that separate the 4% from the 49%.


