How to Accept an AI POC: If the Acceptance Criteria Aren't in the Contract, You Just Lit Money on Fire
In 2022, I was at a 3PL warehouse in Ontario watching a vendor run a POC for vision-based cycle counting. On site, the demo hit 98.6% accuracy — they scanned a rack, the numbers on the screen all checked out. The warehouse manager wanted to sign on the spot. The contract just said "payment due upon acceptance." Three months later the system went live, and on real racks the accuracy dropped to 84%. Boxes with reflective shrink wrap on the top levels? Half of them unrecognizable. The manager asked me what to do. I told him the contract never defined how acceptance gets measured, so in a dispute, he wouldn't get a dime back. In AI projects, eight out of ten landmines are buried in the contract.
Before the Money Leaves, Nail Down the Ruler
A POC is not a demo — it's where you set the ruler. There are three rulers, and all three must be written into a contract exhibit before signing. No exceptions.
Ruler one: effectiveness metrics. This answers "does it get it right?" Don't just write "accuracy" — define the numerator and denominator. For vision counting, something like: "500 SKUs from the test set, scanned on live racks over 5 consecutive days, recognition accuracy ≥97% (a correct hit requires both SKU and quantity to be right)." Effectiveness metrics also need an error budget: which errors are tolerable (damaged cartons) and which are zero-tolerance (high-value SKUs — one miss counts as a failure).
Ruler two: engineering metrics. This answers "does it run reliably?" Three numbers matter most: availability (≥99% uptime over 30 consecutive days, excluding scheduled maintenance windows), latency (single recognition response ≤2 seconds at the 95th percentile), and false-positive rate (≤1%; each false positive deducts points from the acceptance score). Plenty of vendors demo at lightning speed, then fall over the moment the network hiccups. If engineering metrics aren't in the contract, you'll just watch the system crash three times a week with no recourse.
Ruler three: business metrics. This answers "is it worth the money?" At the end of the POC you must produce two numbers: labor hours saved (e.g., cycle count time per person-day drops from 6 hours to 1.5, averaged over 10 counting cycles), and an ROI gate (the contract states: if annualized labor savings < 1.5× the total project cost, the customer may walk away from the purchase with no penalty). The business metric is the boss's bottom line — and your license to kill a bad project.
The Three Tricks Vendors Love
First: they only test the happy path. The demo rack has standard cartons, perfect lighting, barcodes facing out. Insist on this: the test set comes from you, sampled from SKUs actually moving in your warehouse — damaged cartons, reflective film, crooked stacks all included. The contract should say the test set is provided by the customer and the vendor may not see the specific samples in advance.
Second: test set contamination. The vendor tunes the model on your test data ahead of time, so of course it scores well. The defense: prepare two test sets. One is disclosed at signing for tuning; the other is sealed and revealed only on acceptance day. Write it in: "Acceptance uses a blind test set, sealed by the customer, that the vendor has never seen."
Third: the acceptance data doesn't match your real environment. The POC ran beautifully in Building 1, but you're deploying in Building 5 — different lighting, rack heights, SKU mix. The contract must state "acceptance environment = final deployment environment," or at minimum "acceptance is conducted in two or more customer-designated zones."

The Acceptance Runbook: Who Tests, How Long, Who Wins an Argument
Every metric in the contract follows the same five-part template: baseline (what manual labor or the current system achieves today), target value (specific to one decimal place), measurement method (who provides the data, what tooling measures it), measurement window (how many consecutive days, how many runs per day), and the pass/fail rule (the bar, the retest policy, how many failures count as a failed POC). Miss one of the five and acceptance becomes a shouting match.
Here's the runbook I typically use: testing is led by the customer with vendor support, and WMS exports are the system of record for third-party data. The measurement window runs 14 consecutive days, at least 3 shifts per day. Either party can challenge a result within 48 hours, triggering one retest; the retest result is final. If the retest still fails, the customer may terminate with no conditions and the vendor refunds any POC fees already paid. Write dispute resolution as retest-not-argument — that's the clause that matters.
One more thing from the trenches: vendors will agree to acceptance metrics verbally when pitching, then start stalling when it's time to write them into the contract exhibit. My experience is simple — whoever brings the template first holds the leverage. Prepare a blank acceptance sheet with the five-part template and send it over for them to fill in. That's going first. If you wait for the vendor to draft it, every metric that favors you will come back as "to be mutually agreed" — four words that mean nothing.
Here's the honest truth: before you sign, you have all the leverage in a POC. After you sign, all you have is hope. Tomorrow, do one thing: pull the contract for whatever AI project you're currently negotiating and check whether the acceptance exhibit has all five — baseline, target, measurement method, measurement window, pass/fail rule. If any piece is missing, send it back and make the vendor fill it in; for anything you're unsure about, defer to the contract and the vendor's official documentation. That exhibit is the only real insurance policy you'll have in the entire project. And one question worth asking yourself honestly: if this system disappeared next quarter, would your operation notice the difference? If the answer is no, the POC never needed to happen in the first place.





