Last year I visited an e-commerce warehouse in Riverside that had just installed a vision inspection system: six cameras over the conveyor, checking parcel labels for damage and misapplication. The setup ran inference in the cloud — video streams pushed to a public cloud, GPU-processed, results sent back. First week of go-live, it broke: during peak hours the results came back half a beat late, the reject air cylinders never got the signal, and a dozen problem parcels flowed straight into the sort chutes. The IT manager told me: "Nobody mentioned latency before we signed the contract."

Here's my verdict up front: for AI in a warehouse, the gap between 10 ms and 300 ms is not a user-experience issue — it's a can-it-work-at-all issue. Where you deploy inference decides whether your AI is a production asset or an expensive ornament.

The latency math

Camera captures a frame, sends it to the cloud, inference runs, result comes back: typically 100–300 ms round trip, and that's on a good day. You know what warehouse networks are like: metal racking is a natural signal blocker, Wi-Fi dead zones are everywhere, and peak-season congestion can double latency.

Edge deployment? A rugged edge box with GPU compute sitting in a cabinet near the line runs inference at 10–50 ms. Camera to box over a local network, sometimes a direct connection.

Three hard rules for when edge is mandatory:

First, closed-loop control. Reject, sortation, AMR obstacle avoidance — anything where "the AI finishes and the machine must move right now." At 2 m/s belt speed, 300 ms of latency means the parcel has traveled another 0.6 m. You've missed the divert point.

Second, anything safety-related. In mixed pedestrian-vehicle zones, running AMR vision avoidance through the cloud means handing your brakes to network jitter. My verdict: edge, no discussion.

Third, high-frequency inference. One camera at 30 fps, 16 hours a day — cloud billing by API call or GPU hour gets scary fast. An edge box is a one-time investment that gets cheaper per inference the more you run it.

What is the cloud actually good for? Training, model iteration, dashboards, cross-facility analytics. Jobs with no time pressure that need elastic compute. That's the cloud's home turf.

Bandwidth and money: the overlooked lines

Most people cost the GPU line and forget bandwidth. A 1080p camera with H.264 encoding pushes roughly 4–8 Mbps upstream. Fifty cameras: 200–400 Mbps of sustained uplink. You know what a dedicated warehouse line costs per month. Worse is reliability — one circuit flap and every camera's inference drops together.

Edge bandwidth needs are close to zero: video never leaves the building, only inference results (a few KB of JSON) and alert snapshots go up. Lose the internet and the line keeps running; worst case, alerts report late.

Here's how I break down the money. Cloud is pay-as-you-go, good for small or bursty workloads; edge is fixed investment plus depreciation, good for 24/7 continuous runs. The dividing line is easy to compute: list your annual cloud inference bill (GPU + bandwidth + data transfer) and compare it against edge box procurement, three-year depreciation, power, and maintenance. My rule of thumb: above 20 camera streams running continuous inference, edge almost always wins. Below 10 streams with low duty cycle (a few hours of QC per day), cloud is the low-hassle choice.

One trap: cloud vendors have more billing line items than you'd expect — GPU hours, storage, API calls, data egress. Get written quotes for all four before signing. If anything is uncertain, get it confirmed in writing by email. Don't trust the sales rep's "roughly."

My deployment playbook: three tiers

Don't go one-size-fits-all. I split deployments into three tiers:

Tier one, pure edge: vision sortation, reject, AMR navigation and avoidance, real-time verification for pick-to-light. These get edge boxes that keep running if the network dies. On hardware: mind heat and dust — warehouses routinely exceed 35°C in summer, and consumer GPUs won't survive. Go industrial-grade or at least wide-temperature rated.

Tier two, edge inference plus cloud training: this is the mainstream pattern. Edge runs inference; hard samples get sent back to the cloud; the cloud retrains and pushes down updated models. On upload strategy: never upload everything. Only send the samples the model was unsure about (low confidence) and you cut bandwidth by 90% plus.

Tier three, pure cloud: demand forecasting, replenishment suggestions, operations dashboards, cross-warehouse reporting. No real-time requirement, and the cloud's elastic compute and ready-made data tools are genuine advantages.

One operations detail: once you have more than a handful of edge boxes, model version management gets messy. Insist on remote OTA updates and rollback. Don't send engineers around with USB sticks. I saw a warehouse with 12 edge boxes running 4 different model versions; troubleshooting took a week.

Rollout sequence: don't blanket the whole building on day one

The biggest mistake I see with first-time AI vision deployments is going all-in across the facility. My three-step rollout:

Step one, single-point pilot. Pick your most painful line — returns QC or label inspection — with just 2–4 cameras and one edge box. Run the pilot 4–6 weeks. The goal isn't saving money yet; it's producing three tables: latency, accuracy, and false-positive rate. If false positives exceed 5%, floor staff start ignoring alerts — watch that number like a hawk.

Step two, close the model iteration loop. Once the pilot is stable, wire up the edge-inference-plus-cloud-training upload pipeline. On labeling labor: hard samples can run into the hundreds per day, and someone has to label them. My approach is to have QC staff label as they go — the system pushes low-confidence snapshots to a tablet, the QC person taps "right/wrong," and labeling costs essentially nothing.

Step three, replicate. Only expand to other lines once accuracy is stable above 98% and false positives are under 3%. By then you're buying edge boxes in bulk, which typically shaves another 15–20% off unit cost.

One more practical note: network retrofits in older buildings. Edge is forgiving of networks, not free of them — run wired connections from cameras to the edge box; don't cheap out on Wi-Fi for everything. A few thousand dollars of cabling per line is nothing next to a year of cloud bandwidth.

AI vision camera inspecting boxes over a conveyor

Here's one action you can take tomorrow: list every AI scenario running (or planned) in your warehouse and score each on three columns — latency requirement, inference frequency, camera count. Anything needing <100 ms goes to edge. Low frequency, low count goes to cloud. The middle ground gets the hybrid: edge inference plus cloud training. Once that table is done, no vendor can snow you on architecture.

Is your warehouse AI running on the edge or in the cloud? Ever been burned by latency? Tell me in the comments.