Skip to content

Data

No ROI without data AI can trust.

AI returns rest on data the model can trust: clean, reconciled, carrying enough context to mean something. Most disappointing AI projects failed one layer down, on data nobody had gotten into shape.

Four questions come first: what you have, what shape it is in, who may use it, and what it already says.

Audit
two to three weeks
First pipelines
weeks, not quarters
Every source in one place, reconciled into one number set your team refreshes itself.

The problem

Why clean data is hard.

Every system was built to run one job, and each disagrees with the next: the same thing named two ways, one customer as three rows, a date that means order here and delivery there. Handed those contradictions, a model does not flag them. It picks one and answers confidently on the wrong number. The hard part of AI is rarely the intelligence. It is the plumbing under it.

SYSTEM ASYSTEM BSYSTEM C
// the same record, drawn three ways, and no key that joins them

The offers

What AI changes.

The heavy cleansing, matching records that never shared a key, resolving one entity across systems that spell it differently, used to take a team a quarter. AI now does it in a fraction of the time, which is what makes clean foundations affordable at all.

What AI will not do is decide what “correct” means. That stays with the people who know the business.

Data audit & AI-readiness diagnostic.

“What do we actually have?”

We inventory your sources, then grade each against the use cases you want to fund, because “good data” means nothing in the abstract.

You get a diagnostic your CFO can read, not a schema diagram: each source graded, gaps priced, what to fix first. Where the honest finding is “your data cannot support this yet,” it says so.

// each source graded against the use case that needs it

  • SOURCE ONEfresh, owned, reconciledA
  • SOURCE TWOusable day to day, thin for AIC
  • SOURCE THREEcannot support the use case yetF

Two to three weeks.

“How do we get it in shape?”

Data cleansing, foundations & pipelines.

Two jobs: clean the data, then build the plumbing that keeps it clean. We reconcile the sources that disagree, put AI on the matching, and build the minimum architecture the priority use cases need, nothing more.

Platform sprawl looks like progress on a slide; eighteen months of infrastructure before the first use case ships is the quickest way to burn a budget. We do the opposite: pipeline enough for the first pilot, prove it, then extend as the roadmap earns it.

First pipelines ship in weeks, not quarters.

The work, in one project

Case: demand planning that could not start until the data agreed.

Sector · Distribution and supply chain, operating across several markets.

The client came to us for one thing: AI-assisted demand planning, to stop over-ordering some lines and running out of others. A clear, fundable use case.

The catch

The AI had nothing solid to learn from. Three systems described the business, and none of them agreed.

Stock lived in the ERP, warehouse movement in a separate PMS, online sales in the ecommerce platform.

  • The same product carried a different code in each one.
  • A single customer showed up as one buyer in one system and three in another.
  • What one called an order date, the next called a delivery date.

Every time someone wanted a number that spanned all three, an analyst rebuilt it by hand in a spreadsheet.

Feed that to a forecasting model and it does not learn demand. It learns the mess.

// many bearings resolved to one

ERPSKU-4471 · order date
PMSP/4471-A · scan date
ECOMMERCEprod_4471 · click date
reconciliation · 047°
Reconciled data setone product code · one customer identity · one calendar
Demand planning modelfollows once the data is trustworthy

What we did

We made the three systems tell one story. AI did the heavy matching at a scale no team could do by hand:

  • reconciling product identifiers across ERP, PMS, and ecommerce,
  • collapsing duplicate customers into one identity,
  • standardizing every date so a single unit, sold online, shipped from the warehouse, and booked in the ERP, was counted once instead of three times.

The result was one clean, reconciled data set a demand-planning model could train on, with every matching rule written down so the client’s own team owns it.

Where it left them

With the groundwork in place, and the daily spreadsheet rebuild gone: a cross-system number that used to be assembled by hand now comes from one agreed source.

That reconciliation is what the whole use case stands on, and it is the step buyers almost always underestimate. So we scoped and priced it on its own, first. The client watched their data become trustworthy before committing a cent to the model on top of it.

In our experience that order, foundation first, is what keeps an AI budget from being spent twice.

“Who owns it, and where may it live?”

Governance & sovereignty.

The section people skip until it stops a launch.

Rules for who can access data, how long you keep it, and where it may live. A European dataset that cannot feed a US-hosted model. A mainland-China subsidiary whose data cannot leave. A vendor whose terms quietly claim training rights on your inputs. Between Western and Chinese markets, this is where most global AI plans break, and we work both sides of that line.

You get a rulebook mapped to your real data flows, written for operators and reviewed with your counsel.

Two to four weeks, usually alongside the audit.

“What is it already telling us?”

Analytics & reporting.

Before you buy intelligence, extract the intelligence you already paid for. We consolidate your numbers into dashboards your team refreshes itself, on agreed definitions: one dated version of the truth, so trends are provable, not debatable. Often the quickest win of the engagement.

Two to four weeks, and your team keeps the dashboards and refreshes them without us.

// the order the work runs in, when it runs in full

Audittwo to three weeks
Cleansingthe repair the audit priced
Pipelinesweeks, not quarters
Governance and reportingtwo to four weeks, usually in parallel

The audit is the small first commitment, and it prices everything to its right. Nobody has to fund the whole line to find out whether the first stretch was worth it.

FAQ

Questions we get.

We already have a data team. What do you add?

The use-case lens and the East-West governance experience. Internal teams are asked to make data good in general; we make it good for a specific use case.

Do you replace our existing stack?

Almost never. We build on what you run, and any recommendation to replace a tool comes with the migration cost priced in.

Is our data clean enough for AI?

That is what the audit answers, source by source. “Clean enough” is never absolute; it depends on what you ask the data to do.

What does this cost?

Scope and price are set per project after the audit. The audit is the small first commitment, so you can size the rest before funding it.

How does this connect to the AI pillar?

The audit output is scored against the same use-case roadmap the AI pillar produces. The two share one document set.

What do you need from us to start, and how long does the audit take?

Read-only access to the sources in scope, whoever maintains them for an hour each, and the use case you want the data judged against. Two to four weeks is the usual range, and it is shorter when the systems are few and longer when the answer lives in spreadsheets on people’s desktops. We do not need a copy of everything: we need enough of each source to grade it honestly.

Do we need a warehouse or a lakehouse before any of this works?

Usually not, and buying one first is the most common way this budget gets burned. The first pilot needs the sources it touches to agree and to arrive on schedule, which is a pipeline question, not a platform question. If the roadmap later earns central storage, the audit gives you the case for it with the volumes and the cost attached, rather than a vendor slide.

What happens if a source comes back graded F?

You get the repair priced, and the use case it blocks named. A failing grade is never about the source in the abstract: it fails against a specific question you wanted to ask it. Sometimes the fix is weeks of reconciliation, sometimes it is one owner agreeing on one definition, and sometimes the honest recommendation is to change the use case rather than the data.

Who runs the pipelines once you leave?

Your team, and that is a design constraint from the first day rather than a handover at the end. We build with the tools you already operate, document the flows and the failure modes, and run the last pass with your people at the keyboard. If nobody internally can own it, we say so before building it, because an orphaned pipeline decays faster than the mess it replaced.

We operate in China as well as Europe. Can the data move?

Sometimes, in one direction, for some categories, and the rulebook says which. Mainland-China data that cannot leave, European data that cannot feed a US-hosted model, and a subsidiary that needs its own local set are all normal outcomes rather than blockers. We design the architecture around the line instead of asking your counsel to bless a crossing after the build.

Talk to us

Talk to us about AI.

A conversation with the senior team about your markets, your data, and where AI would actually pay back for you. No slides, no obligation, and if the honest answer is that AI is not your next move, you will hear that too.