Ask Your Data This Question Before Buying AI
- Igor Alcantara
- Jun 30
- 10 min read
Updated: Jul 22

Sarah gets to the office at 8:47 on a Monday morning and opens her laptop to a beautiful report. It was generated overnight by the AI platform her company spent four months selecting and three months implementing. Fourteen pages. Clean charts. Confidence intervals. A ranked list of 23 customer accounts flagged as "high churn probability" in the next 90 days, each one color-coded by revenue impact.
She emails the customer success team before her coffee finishes brewing. "These are our priorities for Q3. Let's move."
Three months later, two of the accounts on that list renewed without anyone contacting them. Six had already churned before the report was generated. The data just hadn't been updated. Eleven were flagged because a well-meaning sales rep had marked every open deal older than 90 days as "stagnant" during a CRM cleanup that nobody had reviewed. Four were genuinely at risk. Of those four, the team reached two in time.
The AI did not fail.
It read the data perfectly. It told Sarah exactly what the data said.
That was the problem.

Around 380 BC, Plato wrote a thought experiment in The Republic that has survived intact for roughly 2,400 years without needing an update. A group of prisoners sit chained inside a cave, facing a blank wall. Behind them, a fire burns. Objects pass between the fire and the prisoners, casting shadows across the stone. The prisoners have never seen anything else. They study the shadows, name them, build entire theories about them. They are completely certain about what they see.
Then one prisoner escapes. He walks out into sunlight. He sees actual objects, three-dimensional and real. He is excited. He goes back to tell the others.
They don't believe him. Why would they? Their shadows are consistent, detailed, and familiar.
Sarah's customer success team was doing exactly what those prisoners do. Working from shadows. The shadows were just formatted as a PDF with a company logo on page one.

In 1781, Immanuel Kant published his Critique of Pure Reason and made the situation considerably worse.
Kant argued that even the prisoner who escapes the cave cannot access reality directly. The human mind does not passively receive the world as it is. It actively filters and structures everything through its own categories: space, time, causality. What we experience, he called the phenomenon, the world as it appears to us. The thing as it actually is, independent of any observer, he called the "Ding an sich": the "thing-in-itself."
His conclusion was unsettling. You never reach the Ding an sich. Not ever. The gap between what we perceive and what actually exists is built into the architecture of human cognition.
We don't see, hear or feel reality. We experience an interpretation of reality, curated by our brain.
It is like someone who experience the real world by a TV documentary. For a company, that gap is your data.
Your CRM, your ERP, your HR platform: these are the filters through which your business perceives itself. Your reports, your dashboards, your AI models are all built on what those systems captured, categorized, and stored. Not on the actual reality of your customers, your operations, your risk. On a representation of that reality. And representations, as Kant and Plato both understood in very different ways and across very different centuries, can be deeply, confidently, and quietly wrong.
You might think you know your customer because you have all of this CRM data, but you don not capture every interaction, every smile, every angry response, or every minute of silence. That hidden might be crucial, might be what makes you closer to reality, but there are out of reach of your AI.

The question nobody asks before buying AI
There is a question that almost never gets asked before a company starts an AI initiative. Not "which platform should I choose?" Not "who leads this?" Not even "do we have enough data?" Most companies have plenty of data. The amount was never the issue.
The question is this:
If I had made a major business decision based on my data six months ago, would it have been right?
Not "is our data in a modern system." Not "do we have a data team." The specific, uncomfortable question about whether what you believe to be true is actually true.
Sit with that for a moment.
According to a 2020 Gartner report, poor data quality costs organizations an average of $12.9 million per year. IBM Research estimated the total cost of bad data to the U.S. economy at over $3.1 trillion annually. These are not forecasts about what AI will cost when the data is bad. These are costs companies were already absorbing, quietly, before AI arrived and turned the volume all the way up.
Because that is what AI does. It does not hide data problems. It amplifies them, at scale, with confidence.
What lying data actually looks like
Let me give you three scenarios. I suspect at least one of them is going to sound familiar. Maybe uncomfortably so.
The CRM that remembers people who left
Your sales team has been using the same CRM for four years. Contacts get added regularly. They get updated occasionally. They almost never get removed. Former employees at key accounts are still listed as active contacts. Three major clients went through mergers, and the names in your system no longer match the parent company's current structure. Duplicate entries piled up through two migrations.
Your AI trains on this. It learns that accounts with multiple active contacts close faster, which might not be true. It identifies 30 accounts worth prioritizing based on that signal. It doesn't know that a third of those contacts haven't worked at those companies in over a year.
The AI is not wrong. It's telling you exactly what your data says. Your data is telling about a version of your customer base that no longer exists.
The ERP that speaks two languages
Two years ago, your company acquired a smaller competitor. The ERP integration happened quickly because leadership wanted clean consolidated financials for the following quarter. Product categories were mapped roughly, not carefully. "Industrial fasteners" in the old system became "hardware components" in the new one, except where someone manually entered them as "mechanical parts," which happened more than anyone remembers or wants to admit.
Your AI recommends inventory cuts on what it classifies as slow-moving hardware. Those are the same industrial fasteners that represent 18% of your revenue. Your AI just doesn't know they're the same thing. Honestly, most people wouldn't either.
The HR system nobody owns
Your HR platform was implemented three years ago. It was someone's project. That someone left. Optional fields have never been filled in consistently. Performance ratings use a five-point scale in some departments and a three-point scale in others, because two managers misread the configuration guide at setup. Nobody corrected it because nobody realized it needed correcting.
Your AI flags high flight risk employees. The top two on the list were promoted last month. The system wasn't updated.

None of these companies did anything obviously wrong. No recklessness, no shortcuts taken in bad faith. Just the quiet, accumulated, very human errors that happen when data is treated as a byproduct of work rather than an asset with its own maintenance requirements. Which, to be fair, is how almost every company treats it. At least until something expensive goes wrong. Sometimes we overcome these gaps using our intuition or a simple conversation with a colleague in the break room. AI only has the data to trust. If the data is incomplete, outdated or wrong, then we have a problem.
When those things happen, people blame the AI models, and the bad quality data remains untouched. Why fix the problem if we found the perfect guilty agent?
Why AI makes this worse, not better
I want to push back on an assumption I hear constantly.
The belief is that AI will surface data problems. That if the data is bad, the model will signal uncertainty. That the intelligence in "artificial intelligence" extends to knowing what it doesn't know.
It does not.
A well-trained model on poor data does not produce hedging, uncertain outputs. It produces confident, precise, beautifully visualized outputs that happen to describe a distorted reality. The perfect lie. The model learned from what it was given. It has no reference point outside your data. It cannot compare your CRM to the actual universe of your customers and notice the gaps. It will say "72% churn probability" with the same conviction whether the underlying data is solid or a four-year accumulation of quiet mistakes.

Back to Plato. The prisoners were not unsure about the shadows. They were certain. They had names for them. Certainty and accuracy are not the same thing. Your AI will be certain about whatever your data tells it to be certain about.
A 2021 study published in MIT Sloan Management Review found that the leading cause of AI project failure was not model selection, not talent gaps, not computational limitations. It was trust. Decision-makers who knew, somewhere in the back of their minds, that the data had issues couldn't bring themselves to act on recommendations that came out of it. All the investment, all the implementation, all the presentations to the executive team. And then the thing just sat there because someone in the room remembered the 2019 migration and what it did to the contact records.
That's the real cost. Not wrong recommendations. Paralysis. The expensive kind.
There is no successful AI strategy that did not start with a Data Quality strategy.
Data quality is not a cleaning project
When companies finally decide to address this, most frame it as a one-time event. A sprint. A cleanup weekend. Bring in some analysts, find the duplicates, fix the blanks, move on.
This is the wrong frame entirely.
Data quality is not a cleaning project. It is an ongoing operational discipline, the same way financial reporting is not something you do once and archive.
It starts with profiling: actually understanding what you have, not what you assume you have. I have yet to see a company that profiled its data carefully and found fewer problems than expected. (There is a first time for everything. So far, not it.)
Then you need rules. A documented definition of what good actually looks like. "Complete" means something specific. "Accurate" means something specific. Without rules, there is nothing to measure and no way to know if things are getting better or just differently wrong.
Then monitoring: ongoing observation for drift. Data degrades over time because systems change, people change, and processes change. A dataset that was clean six months ago may not be clean today. It probably isn't.
And finally, lineage: the ability to trace any data point from its source through every transformation to its final use in a report or model. When an AI recommendation goes wrong, you need to know why. Which data, from which system, processed how? If you can't answer that, you can only apologize for the outcome. You cannot prevent the next one.
None of this makes it into the AI strategy deck. But it is the foundation that determines whether everything built on top of it is something you can trust, or just something that looks good in a presentation.
This is where Qlik comes in
Everything above was about a problem. This part is about a specific solution, and I'd rather name it than imply it and hope you connect the dots.
Qlik Data Quality is one of the most complete answers I've seen to the problem this article describes. Not because it's magic. Nothing in data is magic, and anyone selling you magic is selling you something else. But because it is built around the right frame: data quality as a continuous practice, not a periodic event.
What it actually does: it profiles your data across sources, surfacing anomalies, missing values, inconsistencies, and patterns that shouldn't be there. It lets you define and enforce quality rules that run continuously, catching drift before it becomes damage. It maps data lineage across your pipeline, so you can trace a data point from the source system through every transformation to the report or AI model that consumed it. Data Product at its core. And because it integrates with Qlik Cloud Analytics and Qlik Talend Data Integration, quality is embedded in the pipeline from the beginning rather than checked, with some anxiety, at the end.

For a mid-size company preparing for AI, this matters for a reason that doesn't get said enough. Large enterprises have entire governance teams dedicated to this. Small companies move fast and figure it out later. Mid-size companies are caught in the middle: complex enough to have real data problems, lean enough that nobody's job description specifically includes finding and fixing them.
Qlik Data Quality gives companies the infrastructure to catch the CRM ghost, the ERP mistranslation, and the HR drift before the AI turns them into a confident and costly recommendation that someone acts on.
The question from the title of this article? Qlik is what helps you answer it without flinching.
Back to the cave
Plato's story has a second act that most summaries leave out.
The prisoner who escapes the cave and sees real objects in sunlight for the first time is not immediately grateful. His eyes hurt. He's disoriented. After a lifetime of shadows, the brightness is overwhelming. He wants to go back. Back where things were familiar, where the shapes were manageable, where nothing required him to reconsider everything he knew. Sounds a lot like The Matrix? It should.
Eventually he adjusts. He sees the world as it is. And he goes back into the cave to tell the others.
They don't believe him.
The technical part, by the way, is not the hard part. The hard part is organizational, cultural. When you start profiling your data seriously, you will find things that are uncomfortable. You will trace AI recommendations back to data that was wrong for years. You will find that some strategies, some decisions, some confidence intervals, were built on shadows. On a phenomenon that was never a very faithful representation of the Ding an sich.
The temptation is to go back. To keep using the data you have, to trust that the AI will be smart enough to compensate, to not ask the question.

Kant is useful here. You will never have perfect data. You will never fully close the gap between what your systems capture and what is actually true about your business. That gap is, in some sense, permanent. But you can narrow it. You can build the practices that make your phenomenon a much more faithful picture of reality. Close enough to act on with genuine confidence, not just confidence that looks good on a slide.
You might not ever eliminate the gap, but you have to know it.
That gap, between what your data says and what is actually happening, is precisely where AI either delivers on its promise or fails at scale, loudly, and with charts.
Before you sign that AI contract: ask your data the question. Run the profile. Look at what you find. Sit with the discomfort.
You might not like the answer. But you will be very glad you asked before the AI found it for you.
Author: Igor Alcantara
Note: This article was written by a human. AI was used for the illustrations and grammar/spelling check.




Comments