// Datenbasis

What Data Does AI Need? Five Questions Before You Start

AI doesn't need perfect data — but it does need the right data in the right condition. Five questions to assess whether your data is sufficient for an AI use case. Before you bring anyone in.

August 4, 2026 · approx. 8 Min. read · Jan Fischer

Heatmap aus Datenbereichen und Stufen mit markierter Schwachstelle

Motiv

Contents

The question comes up in almost every first conversation, usually with a hint of embarrassment: Is our data actually good enough for AI? Behind it lies the worry that you'll need to complete a years-long data project before anything AI-related can happen. That concern is understandable — and in most cases, unfounded. It stems from a misunderstanding of what AI actually requires from data.

This article addresses that misunderstanding. It explains what "AI-ready" concretely means, why your own assessment tends to be off, and which five questions you can use to evaluate the state of your data yourself. In an afternoon, with tools you already have.

What does AI-ready actually mean for data?

Good enough for a specific use case. That's all the term means — no more, no less. AI-ready is not a general property that data either has or doesn't have. The same data landscape might be perfectly sufficient for a customer service chatbot and far too patchy for an ordering agent.

This isn't just a matter of semantics — it changes the task entirely. Gartner warns that a lack of AI-ready data puts projects at fundamental risk, while also noting that a significant share of organisations lack the necessary data foundations. The analysts make a point that often gets overlooked: whether data is AI-ready depends on its intended use. Asking broadly whether your data is good enough is a question nobody can answer. Asking whether your data is sufficient for this one specific use case gives you something you can actually test.

For retail businesses, this is good news. Nobody needs to overhaul their entire data landscape before launching a first AI use case. It's enough to bring the relevant data areas up to standard — the ones the use case actually needs. Everything else can stay as imperfect as it is.

Why isn't your own assessment enough?

People unconsciously compensate for poor data and therefore rate it as better than it is. That's the core of the problem. A buyer who has worked with the product master for years knows its quirks. She knows that lead times for supplier X are never accurate and adjusts in her head. For her, the data is fine — she can work with it. AI can't do that. It takes every value at face value.

The numbers on self-perception reflect this. In a Gartner survey, 63 percent of data management leaders say they either lack appropriate data management practices for AI or aren't sure whether they have them. In the Cisco AI Readiness Index, a global survey of nearly 8,000 executives, 80 percent of companies report gaps in preparing and cleaning their data for AI projects. Only 32 percent rate their own data readiness as high. All of these figures are based on self-reporting. Actual conditions tend to be worse, for the reason noted above: those who compensate underestimate the gap.

For the same reason, the free readiness checks circulating by the dozen right now don't help much. They collect self-assessments from multiple departments and calculate a score. Each department rates its own work — and that rating reliably comes out too favourable. A score without supporting evidence is an opinion with a decimal point. It feels precise but only measures how people feel. A picture becomes reliable only when statements are checked against evidence from the systems themselves: a data quality report, interface documentation, an operational log.

A worked example, with fictional data. A multichannel retailer describes its product master as "largely well-maintained" in a workshop. Direct measurement in the systems tells a different story: lead times are maintained for 54 percent of the 84,312 active items, minimum order quantities for 61 percent, and 8.3 percent of items are suspected duplicates. None of these figures were known to the people involved. The full findings are available in the free sample report.

What properties do the data need to have?

Five — and all of them are measurable. Completeness: the fields the AI uses to make decisions are populated. This means the decision-relevant fields, not full coverage across the board. Accuracy and uniqueness: values are correct, and every product, customer, or supplier exists only once. Duplicate records confuse any analysis. Timeliness: data is as current as the use case requires. Last night's stock levels are fine for slow movers, but not for fast-moving ranges.

Five measurable properties of AI-ready data
All five properties can be measured directly in the systems. None of them is a matter of opinion.

The fourth property is the most frequently overlooked: accessibility. AI — and especially agents — work through interfaces. Data that lives only in Excel files or in the heads of experienced staff simply doesn't exist for the system, no matter how good it is. Accessibility also means that data models are documented in a machine-readable way, so a system can understand what a field actually means.

And finally, ownership. Strictly speaking, this isn't a property of the data itself but of its environment — yet it determines everything else. Without a named person responsible for a data domain, any clean-up effort will deteriorate within months. We've seen data landscapes that were overhauled three times and slipped back to their original state three times. The cause was never the technology.

How does the data requirement depend on the use case?

Every use case requires a different combination of data domains — and only that combination needs to be in good shape. We work with a matrix for this: six retail data domains (product master data, customer, transaction, inventory, suppliers, online behaviour) and four maturity levels, from foundation to autonomy. A use case typically requires three to five cells in this matrix.

Matrix of data domains and maturity levels with cells marked per use case
A replenishment agent needs different cells than a product copy AI. Only what the use case requires gets assessed.

An example makes the difference tangible. A replenishment agent runs on product master data and inventory. Online behaviour data is irrelevant to it. An AI that writes product copy for the shop, on the other hand, needs rich product attributes and no inventory data at all. Dynamic pricing, by contrast, requires transactions, pricing terms, and current stock levels simultaneously — making it the most demanding use case of all. What this means in detail for replenishment is worked through in the article on agentic replenishment.

This logic leads to the most important conclusion in this article: the question about data can only be answered together with the question about the use case. Define the use case first, then assess. Anyone who reverses the order and cleans up data broadly upfront is very likely tidying the wrong areas.

Which five questions can you use to assess this yourself?

Take your most important AI use case and answer these five questions. You don't need an external consultant — just the people who work with the systems.

  1. Which fields does the decision the AI is meant to make actually require, and how completely are they maintained? The query takes an hour. If you don't know the number, you don't know your risk.
  2. Is there a single leading system for each data domain you need? If two systems maintain the same data in parallel, their contradictions will flow into every AI decision.
  3. How old is the data by the time the AI reads it, and is that fresh enough for the use case? Ask about the timestamp, not about the feeling.
  4. Can a system access the data through interfaces, or does it live in Excel and people's heads? What isn't accessible doesn't exist for the AI.
  5. Who is responsible for each data domain you need — by name? If the answer is a shrug, you've already found the most important gap.

Answered honestly, these five questions produce a solid preliminary picture. They don't replace a measurement backed by evidence from the systems, but they show whether the next step is worth taking and where it should focus. A practical tip: have two people answer the questions independently — one from the business side and one from IT. The gaps between their answers are often more revealing than the answers themselves.

What if the answers are sobering?

Close gaps selectively rather than overhauling everything at once. Sobering answers are not a verdict on your organisation — they're the norm. What matters is what you do next. The expensive reflex is the company-wide data quality project that tackles every domain simultaneously and gets overtaken by the next reorganisation 18 months later. The smarter route goes through the use case: measure the three to five cells you need, identify the weakest one, close that single gap, and then start.

For the first part of that journey, there is an open standard. How a reliability score from 0 to 100 is derived from evidence, which criteria apply, and how the bottleneck is identified is described on the methodology page. The assessment itself takes three weeks and is open-ended. In the best case, it confirms that your data is sufficient for the use case and you move forward with confidence. In the other case, you know exactly what to address first. Either outcome is worth more than another year of wondering whether your data is good enough.

Sources

SourceWhat it says
Gartner · Lack of AI-ready data puts AI projects at risk, Februar 2025
Cisco · AI Readiness Index 2024
prodct · Beispiel-Report ACME Inc. (fiktive Daten)
Production readiness audit

Three weeks, open outcome.

The production readiness audit measures whether your data is good enough for your use case. With evidence from your systems, not with self-assessment.

Go to the audit
Request the audit 3 weeks · open outcome