Summary
- Enterprises abandon 46% of AI proofs-of-concept before they reach production. The reasons behind failed enterprise legal AI deployment are almost always organizational, not technical.
- Pilots succeed because they are built to: a curated set of documents, the team who piloted it, and a fraction of real volume. Daily use changes all three at once.
- Enterprise legal AI implementation is harder than most other deployments because the underlying documents are messier than they look, the stakes are higher (privilege, discovery, personal accountability), and contract work crosses legal, procurement, and often other teams.
- Four moves separate deployments that make it to production: set a KPI baseline before the tool arrives, clean up the documents before the tool meets them, change the process so the tool sits in the path of the work, and give ownership to someone with real budget and authority.
- Workflow redesign has the strongest link to bottom-line impact —but only 21% of companies have meaningfully redesigned any workflow around AI.
Companies are dropping AI projects before anyone uses them for real work.
In a study of more than 1,000 companies, the average one abandoned 46% of its AI proofs-of-concept before they ever reached production. Nearly half, gone somewhere between the demo and the desk.
It is tempting to blame the pilot. But most pilots go well, because they are built to. For instance, with legal AI: You run it on a curated set of documents with your own team, at a fraction of real volume. That is a reasonable way to test whether the tool performs as promised.
Daily use changes all three of those at once. Documents are no longer curated. The users are evolved beyond the same people who pilot it. The volume is no longer a fraction of anything.
The deployments that make it to production (and last) do four ordinary things, in a particular order, starting before the pilot phase ends. This holds across industries, but especially true for contract work, where Legal and Procurement are accountable for the outcome.
Research points at the setup around the tool, not the tool
RAND studied why AI projects fail and found the causes were almost always organizational rather than technical. Leadership. Communication. Data quality. Whether the tool fit how people already worked. They also found AI projects fail at roughly twice the rate of comparable non-AI IT projects. That’s striking because the AI software is supposed to be better than the one it replaced.
MIT's 2025 research on generative AI in business arrived at the same conclusion from the opposite direction. What separated the companies getting a return had nothing to do with model quality and nothing to do with regulation. It came down to whether the tool learned from use, connected to the other systems in the stack, and matched how the work actually moved.
Finally, McKinsey found that only 21% of companies have meaningfully redesigned any workflow around AI. Of everything McKinsey measured, workflow redesign had the strongest link to bottom-line impact.
That gap is telling: the change most closely associated with business impact is also the one that relatively few companies have made.
None of this is an argument against buying good software. What to note is that the work sitting around the software determines whether the software matters. Ready the data. Redesign the workflow around the tool. Put someone behind it in production who can actually change things.
Why legal and procurement have it harder
Both functions are adopting faster than they are building the groundwork underneath.
Thomson Reuters found generative AI use among law firm professionals nearly doubled in a year, from 14% to 26%, while 48% still had no formal policy governing it. In procurement, the Hackett Group found contract management is one of the most common things teams are piloting, at roughly 20% pilot activity.
In other words, contract work is a named procurement use case and a core legal one at the same time, so a contract tool nobody uses becomes two functions' problem simultaneously, which usually means it becomes nobody's.
Three more things make legal AI adoption harder than many other types of deployments.
First, the underlying work is often messier than it looks. Contracts are a good example: Scanned PDFs from 2014. Templates that changed three times and were never versioned. Amendments nobody labeled. Superseded drafts sitting in the same folder as the executed agreement, with nothing in the filename to distinguish them. Getting that foundation in order often takes longer than the project plan allows.
The stakes are also higher. Privilege, discovery, and personal accountability all live in this work. People need real proof before they rely on an answer a tool produced, and that caution is professional judgment doing its job rather than resistance to change.
Lastly, the work crosses teams. A contract moves between procurement and legal, sometimes finance, sometimes a business unit, which means changing the process requires two or more functions to agree on a new one. Redesigning something inside a single team is a project. Redesigning something between teams is a negotiation.
What gets a pilot to production
There are four key moves that ensure the success of a pilot. Each one needs budget, mandate, cross-functional access, and executive backing–things the Legal Ops or Procurement leads can’t provide alone.
1. Decide what number has to change, then measure it before the tool arrives
The first key is to pick a single KPI that the deployment needs to impact. Cycle time on one contract type. Cost to process a document at full volume. How many agreements get signed without legal touching them at all. Then measure it while the old process is still running. You need a baseline to prove success, which is why this step comes first.
Without a baseline, rollouts often come across successful simply by deploying seats and having pretty dashboards. Meanwhile people keep doing the work the old way, and the KPI that matters hasn’t moved. Eight months later, leadership finds activity but no real ROI.
Cost belongs here too. Pilot pricing and full-volume pricing are different. Before the pilot ends, ask your vendor what a single document costs to process at production volume, on your document mix, rather than a clean sample.
2. Clean up the documents before the tool meets them
The second is to profile the repository. Resolve versions and amendments. Fix what the tool will be reading, now, rather later.
Skipping this will give you the same pattern. A contract review tool performs well on ninety NDAs somebody hand-picked. Then it meets eleven thousand real documents. Accuracy will slip here and when people hit enough wrong answers, they will stop trusting it. Eventually, your tool will fall out of use, even without anyone deciding sunset its use.
Do this early because trust is harder to rebuild than it is to establish. Once a tool gives lawyers the wrong output, you have to earn that trust back on the same work, in front of the same people, and they remember.
3. Change the process so the tool sits in the path of the work
Check your intake, routing, thresholds, handoffs, and more. Change them accordingly to ensure that the work runs through the tool rather than past it.
Without this step, the work is just as before, but with extra steps. A procurement team pilots contract AI to speed up supplier reviews, and in the pilot it flags risky clauses fast and well. But in production, intake is unchanged, approval thresholds are unchanged, and category managers and legal pass work back and forth exactly as they always did. The tool becomes a place some people check occasionally.
Of the evidence cited here, this step has the strongest link to impact. Yet only 21% of companies have done it and it often requires authority across teams, which is why it gets skipped.
4. Give it to someone who can actually change things
This one runs alongside the other three rather than after them.
Most advice says "name an owner" and stops there. But an owner needs three things: budget, authority, and accountability for the KPI from step one. They need authority to change how work moves between legal and procurement.
Without that authority, the role is just coordination. And coordination cannot redesign a process that two teams share.
Getting to production is where the work is
The teams that got AI to deliver real results did ordinary things in the right order. They knew which number had to move. They cleaned up the data. They changed the workflow. And they gave someone the authority to keep improving it.
If your rollout has not led to sustained use and measurable impact, look beyond the tool. The next step is operational: fix the data, process, ownership, and measurement around it. Legal and Procurement are well positioned to lead that work, with the right executive backing.
And if you are mid-rollout, establish your baseline now. Measure how the work is done — time, cost, volume, and error rate of the current process before AI changes it. That gives you something to compare against and shows whether the new workflow actually improved the work.
Comments
Share your thoughts on this article. Comments are moderated before publishing.