Something changed in how AI work is sold this year. The unit is no longer a bespoke build. It is an assembled platform: an agent runtime, a governance layer, a marketplace of prebuilt components, and people to wire all of it into whatever you already run. Firms arrive with software they have already written instead of only with engineers who will write some for you.
We are moving the same way and think it is correct. Orchestration, logging, access control, lifecycle management: these are solved problems that generalize across clients, and rebuilding them per engagement was always waste. Our own intelligence layer,bearingbridge intelligence, exists for exactly that reason. It ingests data, content, and documents from multiple sources so that work does not restart at every engagement.
What none of us has found a way to productize is the other half. So the buying question is narrower than the pitch: which parts of this can be purchased, and which parts stay yours regardless of who you hire.
What is for sale, and what is not
- Agent runtime and orchestration
- Logging, access control, observability
- Prebuilt components for common functions
- Multi-cloud, multi-model plumbing
- Deployment patterns for regulated environments
- The state of your source data
- Whether your real rules are written down
- Who is accountable for a wrong answer
- The threshold at which you stop
- The number you claim it moved
The left column is genuine engineering, and buying it is usually the right call. The right column is where AI programs actually fail. No procurement decision moves a single item from right to left.
The promise that moves the risk
The common selling line now is a governed semantic layer across your databases, document stores, and event streams, delivered without the consolidation program as a prerequisite. Agents get a structured view of the company without an eighteen-month migration first.
That is an attractive promise and a technically honest one. Anyone who has watched a data consolidation program die in month nine understands why it sells.
The risk moves, though. It does not disappear. A governed view over inconsistent data returns inconsistent answers faster, and logs every one of them. If the same customer sits in three systems under two spellings, a knowledge graph inherits that, it does not resolve it. If your product hierarchy changed in 2023 and half the documents still use the old one, an agent reads both and treats them as equally current. Governance tells you what the agent read. It cannot tell you which version was true.
Governance tells you what the agent read. It cannot tell you which version was true.
Hence the house line: context is everything, and context comes from data. That line came out of an FMCG pilot where the outputs were poor and the client assumed the model was at fault. It was not. The reference material was thin, stale in places, and self-contradicting in others. Swapping models changed nothing measurable. Fixing the source material changed the result.
An assembled platform does not exempt anyone from that work, ours included. It changes when you find out you needed it, and finding out after deployment costs more.
Four kinds of number
Platform pitches carry numbers, and four different species of number get delivered in the same confident tone. Somewhere between the announcement and slide four, “projected to” becomes “delivers.” Nobody sits down and decides to do this. It just happens somewhere in the deck, and by the time anyone notices it is in the budget.
Open a row to see what it establishes, and what to ask for.
A deployment countShipped, somewhere
It has shipped somewhere
To whom, doing what, over what period
A self-reported speedupAn impression
An impression, unaudited
Baseline, method, who measured it
A projected savingA model
A model of the future, not a result
The assumptions, and who signs for them
A measured outcomeSettles it
The only class that settles anything
Before and after, both dated
Row two deserves particular suspicion, because somebody tested that class properly.
In a randomized controlled trial, experienced developers took 19% longer to finish real tasks when allowed to use AI tools, while estimating afterward that the tools had made them 20% faster.
39 percentage points apart, in the direction that flattered them.
METR later reran it on newer tools and marked the original out of date on its own page, which is more than most people quoting it have noticed.
Among developers from the original study who took part again, the estimated speedup was 18%. Among newly recruited developers it was 4%.
Set the headline numbers aside. What remains is that the people in that first trial were wrong about their own productivity by 39 percentage points, in the direction that flattered them. They were experienced professionals working in code they knew well. Your team is not exempt from that, and neither are we.
Which is why a vendor productivity figure, ours included, cannot carry the weight buyers place on it. The version that would settle it looks different: a baseline taken before, the same measurement repeated after, same method, both dates recorded.
Sovereignty is a property, not a slogan
So much for the buying side. The other half of the conversation is where all of it is allowed to sit.
Every platform in this category now leads on control: your data, your models, your decisions, kept where you want them. For European buyers that is the right thing to lead with. Sovereignty is testable, though, and four questions do it faster than any architecture diagram.
Start with where the context layer physically lives, and under which jurisdiction. Not the models. The context itself: the knowledge graph, the embeddings, whatever logic got pulled out of your procedures. That artifact is a compressed description of how your company works, and it is worth more than any single document that went into it.
After that, who can read it. Vendor support staff during an incident is a normal answer. That becomes a problem only when nobody asked.
Then ask what leaves the building at inference time. A prompt carrying customer records out to a hosted model is a data transfer, whatever the diagram calls it.
What happens if you leave in year three. Exporting a workflow definition is easy. Exporting the context layer, the tuning, and the accumulated knowledge of how your agents behave is a different proposition. No lock-in is a claim to test, not a feature to accept, and the moment to test it is before signature while you still have leverage.
What we would run first
None of this is an argument against buying the platform. It is an argument about sequence. Four things, before that decision gets made.
Audit the context, not the model.
Take the twenty documents and three systems a first agent would actually read. Check whether they are current, whether they contradict each other, and how many near-duplicates exist. Sort by last-modified date first. When the policy document a support agent will quote was last touched in 2021 and the actual policy changed twice since, you have your answer before opening anything. This is days of work, not quarters.
Next, find out what is currently unwritten. Extracting logic from standard operating procedures works well where those procedures are accurate. Where the real rule lives in one senior person's head and the document has been wrong for years, extraction will faithfully encode the wrong rule and give it an audit trail.
Baseline the number you intend to move before anything deploys, and use the same method both times with both dates recorded. Without that you cannot tell a real saving from a projection everyone stopped questioning.
Write the kill criterion.
A quality floor, a cost ceiling, and a handover test, each with a number and a date, agreed while everybody is still optimistic. That is the only moment such a thing gets written honestly.
None of the four takes a quarter. Together they are a couple of weeks of work, and they change what you are buying rather than whether you buy.
Skipping that last step is how programs end up in the following number, which counts abandonment rather than failure. The two are not the same, and abandonment is the expensive one.
The share of companies abandoning most of their AI initiatives rose from 17% to 42% in a single year, with the average organization scrapping 46% of proofs of concept before production.
That last step is the Bearing phase of AZIMUTH, described on theMethod page, and it is the phase clients push back on hardest. A defensible stop is cheaper than a slow abandonment, and it stays cheaper when the thing being stopped is a platform program with a marketplace attached.
So buy the half that can be productized. It is usually better engineering than a team would produce in-house, and it costs less. Just be clear-eyed that the other half, your data and your written rules and the person whose name goes on the outcome, is not included in the contract and never was.
