Skip to content

AI

Intelligence on top of your data.

Rules have run your business for decades, and most of it should stay that way. Intelligence is what rules cannot do: read a messy supplier email, weigh it against three years of contracts, tell a buyer what it means.

This pillar decides whether a use case is worth doing, and which model should do it. Building it is the Tech pillar.

Reseller agreements
none
Every report
dated
// the routing recommendationThe cheapest model that clears each bar, named per task.

Judging the case

“Where does intelligence pay for itself here?”

Use cases, judged before anyone commits.

Three questions kill most candidates.

Most die on the second, which is why this pillar and theData pillar share one document set.

The questionWhat it catches
Can a rule do this?Work that does not need a model at all
Is the context data good enough?The use case that fails eight weeks in
Does anyone own the decision?Output nobody will act on

Survivors get tested on your real data before anyone funds a build. Each test reports cost per run today, projected cost at your volumes, and quality against examples your team graded.

// agreed before work starts

Every case also carries a kill criterion, agreed before work starts: a number and a date at which we stop. We write it down at the beginning, and when we hit it we report the stop as the result it is.

What you get

A scored roadmap, validated economics, and a verdict per case: fund it, park it, or solve it without AI. The third comes up more than clients expect.

Format

Three to six weeks, fixed price, scoped after a call.

The second question

Context is the whole game.

That second question kills more use cases than the other two combined. Here is why.

The models are good. The systems built on them are unreliable, and the cause is almost never the model.

  • One entity carrying four names across three systems
  • Documentation describing an architecture replaced last year
  • Retrieval returning what is similar rather than what is correct

Swap in a stronger model and you get the same wrong answer, faster, at higher cost.

// same context, stronger model

model Aanswers in seconds
model B, strongeranswers faster, costs more
the same wrong answerbecause the context under it never changed
// the model was never the bottleneck

Selection

“Which model, and how would we know?”

Model selection has five variables.

Public leaderboards settle none of it. The model builders document the problem themselves.

VariableWhy it decides
QualityWhere everyone starts, and rarely where it ends
Cost at your volumesChanges which use cases are viable at all
LatencyFour seconds slower loses a live workflow, wins a batch one
Governance fitDisqualifies more models than quality, and earlier
JurisdictionWho can compel access to your inputs. See below

Over 16% of MMLU samples are contaminated, with a non-trivial fraction exhibiting severe leakage.

Touvron et al., Llama 2 technical report (2023)

Frontier models now cluster above 88% on that same benchmark, so the spread across the top few sits inside the noise. A high score mostly proves a model can pass a test that is already on the internet.

Access, not a credential

Both sides of the map.

We operate in China and outside it. That is access, not a credential: Western and Chinese models on one task set, same week, same grading.

Western
  • OpenAI
  • Anthropic
  • Google
  • Mistral
Chinese
  • Qwen
  • DeepSeek
  • GLM
  • Kimi

Most consultancies test one ecosystem because they only know one. Results come back mixed often enough that testing one side is guessing.

Price is where the decision usually turns. Classification, extraction, routing, summarizing clean inputs: that is a large share of what production systems do all day.

A model fifteen times cheaper often scores the same on those tasks. At volume, that gap decides whether a use case survives its first budget review.

The reverse holds too. On hard judgment calls the expensive model earns its price, and routing everything cheap to save money fails in a quieter, more expensive way. Neither answer generalizes. Both are testable.

So we test. Your real tasks, your languages, your cost ceiling, graded blind by your team where quality is subjective.

What you get

Quality, cost per thousand runs, latency and governance fit per task. Plus a routing recommendation naming the cheapest model that clears each bar.

Format

Two to four weeks per round. Re-runs cost less; the task set already exists.

The seat we sit in

We do not resell any of this.

No reseller agreements. No partner tiers. No revenue share. We pay for the same APIs you would.

Benchmarking is only credible from a seat with nothing to sell. The large firms went the other way and built their AI practices on named hyperscaler and lab alliances. That is a reasonable way to run a consulting business, and a poor vantage point from which to grade a bake-off.

So when a Chinese open-weight model wins on cost and clears your rules, we say so, andnobody here loses margin when you act on it.

  • No reseller agreements.

  • No partner tiers.

  • No revenue share.

  • Same API invoices you would pay.

Jurisdiction

“Where is this data allowed to live?”

Sovereignty is not residency.

Data sitting in Frankfurt under a US-headquartered provider is EU-resident but not EU-sovereign. The CLOUD Act reaches the parent company wherever the servers happen to be. The same logic runs in the other direction with Chinese providers and Chinese data rules.

// cost and constraint climb with the rung

Level oneEU endpoint of a non-EU providerLowest
Level twoSelf-hosted open-weight model on EU infrastructureHigher
Level threeEU-domiciled provider governing the whole stackHighest

Most workloads do not need level three. Some cannot legally sit below it. We score that per workload during the benchmark rather than handing you a separate compliance exercise afterward.

// the dates most pages still get wrong

One correction worth making, because plenty of pages have it wrong: the May 2026 digital omnibus agreement moved high-risk obligations for biometrics, employment and critical infrastructure to 2 December 2027, and product-embedded systems to 2 August 2028.

Shelf life

“How long will this answer hold?”

Every benchmark has a shelf life.

Menlo Ventures put numbers on how fast the ground moves.

We estimate Anthropic now earns 40% of enterprise LLM spend, up from 24% last year and 12% in 2023. Over the same period, OpenAI lost nearly half of its enterprise share, falling to 27% from 50% in 2023.

Menlo Ventures, “2025: The State of Generative AI in the Enterprise”

Anyone who picked a permanent winner in 2023 has been wrong twice since. So we date every report and treat the answer as perishable.

// notice you get before something moves

  • General modelssix months notice
  • Specialized variantsthree months
  • Previewstwo weeks

In June 2026 OpenAI deprecated its own Evals platform and Agent Builder. Models retire, prices move, and providers change checkpoints behind an unchanged name.

The benchmark tells you the right answer now and how long now lasts. Building so a changed answer costs a config change instead of a rewrite is engineering, and it belongs to Tech.

Measurement

The measurement that keeps the rest honest.

Metric defined before work starts. Dated baseline. Same method at the agreed time. If the numbers say the pilot failed, the report says the pilot failed.

The dangerous failure is the silent one. A system passes every pre-launch check, then degrades because a provider changed something underneath or your data shifted. Infrastructure metrics stay green the whole time. Only a regression set catches it.

// what the dashboards do not show

infrastructure · green throughoutanswer quality · drifting
A regression set is the only instrument pointed at the second line.

So we keep one: fixed graded cases, re-run on schedule, dated and retained. That set is also what makes a cheaper model safe to adopt, because you can prove it holds the bar before you move the traffic.

We run this as a standalone audit on systems we did not build. Two to three weeks, before-and-after measurements, written verdict with the evidence attached.

One applied case

One applied case: AI visibility.

Dated baseline, repeatable method, results kept. Point that at a marketing question and you get this.

When a prospect asks an assistant who to work with, someone gets recommended. We measure whether it is you: which assistants mention you, in which markets, and whose content they cite when they answer.

The category is thick with unverifiable claims, so we treat it as measurement instead. It runs on our own platform, so your team can take it in-house whenever you want.

See bearingbridge intelligence

// who gets recommended

  • You
  • Competitor
  • Competitor
  • Competitor

share of assistant answers · dated · re-run on schedule

FAQ

Questions we get.

Which models do you resell or partner with?

None. No agreements, no tiers, no revenue share. We pay for the same APIs you would.

Can you work inside our compliance constraints?

Usually, and finding out is part of the benchmark. Where a model fails your residency or confidentiality rules, it is out whatever it scored.

Do you have a preferred model?

For a specific task, after testing, yes. In general, no. Anyone answering that in the abstract is selling something.

Are Chinese models safe to use?

Depends on the workload. For non-sensitive high-volume work the cost case is often overwhelming. For regulated or personal data, governance may rule them out before quality comes up. Governance gets scored alongside quality and cost, and the rules are yours rather than ours.

How long does a benchmark stay valid?

Shorter than most people assume. We date every report and recommend re-running quarterly for anything carrying real volume.

What if we have already picked a model?

Then it is a check rather than a search. Either it confirms the choice, or it finds half your traffic going to a tier it does not need.

Can you evaluate a system another vendor built?

Yes, as a standalone audit. We will say plainly when a system performs well, and equally plainly when it does not.

Do we need our data in order before you can start?

Not perfect, but visible to us. Where the data cannot support a use case yet, the document says so, prices the path, and that work moves to the Data pillar.

What does an engagement cost?

Benchmarks per round, use-case work as a fixed sprint, re-runs cheaper than first runs. Scoped on a call before anything is quoted.

Do you build what the benchmark recommends?

We can, and that sits in Tech with its own scope and price. Plenty of clients take the recommendation to their own engineers, which is one reason the benchmark is sold separately.

Talk to us

Talk to us about AI.

A conversation with the senior team about your markets, your data, and where intelligence would actually pay back. No slides, no obligation, and if the honest answer is that AI is not your next move, you will hear that too.