After forty-five years commissioning marine systems, I've learned one lesson that never changes: nothing goes into service until it survives a sea trial.

The builder's brochure is marketing. The spec sheet is a claim. Acceptance happens one way: you instrument the vessel, you load her, you run her in the conditions she'll actually face, and you document what breaks. Then — and only then — do you sign.

Watching companies integrate AI into their workflows right now, I keep seeing the same thing: fleets being accepted off the brochure.

The market data backs the unease. The dominant enterprise finding this year is that agentic AI is real but underdelivering in production — and the bottleneck isn't the models. It's architecture, governance, and operating discipline. Stanford researchers frame 2026 as the year AI evangelism gives way to AI evaluation. The question is no longer can it do this? It's how well, under what load, and who verified it?

That's a commissioning problem. So let me ask the questions every critical system has always had to answer.

The questions nobody asks before deployment

Who ran the sea trial?

If your evidence that a tool works is the vendor's own benchmark, you have no evidence. A vendor grading its own exam and finishing first is not validation — it's advertising. What were your acceptance criteria, in writing, before the demo started?

What load was it tested under?

Every system performs in flat water. The demo dataset is flat water. Your real workflow — the malformed inputs, the exception cases, the operational pressure — is the sea state that matters. Did anyone run the tool against your ugliest real cases before it touched production?

Who signs the acceptance?

On a vessel, a named person witnesses trials and signs. Accountability has a signature on it. In most AI deployments I've observed, adoption just happens. Someone starts using a tool, output flows downstream, and six months later nobody can say who decided this system was fit for service. If no one owns the acceptance, no one owns the failure.

How will you detect failure?

A bilge pump fails loud. An AI system fails quiet — confident, fluent, and wrong. I've written before about the sycophancy trap: a system engineered to say yes will say yes right up until the consequences arrive. If nothing in your workflow catches a plausible wrong answer before it influences a decision, you don't have monitoring. You have hope.

Where the pain actually comes from

Friction in AI integration gets blamed on the technology. In my experience the causation runs elsewhere, and it runs in a familiar pattern — because it's the same pattern as every botched repower I've ever surveyed.

Pain source one: bolting new propulsion onto an old hull.

Most organizations drop AI into workflows designed for a different power plant, then wonder why the whole structure vibrates. The emerging consensus this year says it plainly: the gains come from redesigning the process around the new capability, not fastening the capability onto the old process. A repower without a structural review isn't an upgrade. It's a defect with a delivery date.

Pain source two: skipping the load path.

Nobody maps where the AI's output actually goes. Whose decision consumes it? What downstream action depends on it being right? Unverified output entering a decision chain is cognitive FOD — foreign object debris for decisions. Small, ignorable, and catastrophic at speed.

Pain source three: no recertification interval.

Vessels get resurveyed. Rigging gets reinspected. But AI deployments get accepted once — if that — and then drift. Models update. Workflows change. The tool that passed in March is not the tool you're running in September. An acceptance without a recertification clause is a snapshot, not a safety case.

Before you sign

The remedy isn't complicated. It's the discipline no one would skip on a vessel, a crane, or a pressure system — applied to a category that hasn't earned an exemption.

Write acceptance criteria before the demo. Define what fit for service means in your operation, in measurable terms, before a vendor shows you anything. If you can't write the criteria, you're not ready to buy.

Trial on your worst water. Assemble your twenty ugliest real cases — the exceptions, the edge conditions, the files that break things. That's the trial course. Anyone can pass the vendor's course.

Put one name on the signature line. One person witnesses, one person accepts, in writing. Committees don't feel accountability; names do.

Instrument for quiet failure. Build the check that catches the confident wrong answer: sampling audits, human verification at the decision points that matter, a defined threshold that triggers review. Judgment is not a phase of the project. It's a permanent watch-standing duty.

Set the resurvey date. Every accepted AI workflow gets a recertification interval. When the model updates or the process changes, the clock resets. Encode once; judge always.

The bottom line

Engineering hasn't changed. Only the technology has. The organizations winning with AI this year aren't the ones that adopted fastest — they're the ones that commissioned properly, treating deployment as an engineering acceptance rather than a software install. The friction everyone is feeling isn't the price of innovation. It's the vibration of systems that were never trialed, never instrumented, and never signed for.

First truths over jargon: if it hasn't run a sea trial, it isn't in service. It's a liability with a login.

Captain Carl McBride is Director of Technical Services at McBride Technical Group, applying forty-five years of marine commissioning and forensic survey discipline to the evaluation of AI-integrated workflows. High Tech. High Touch. High Performance.