Share this on:
What You'll Learn
Nobody’s system prompt says “ignore the fact that this customer has four different email addresses on file.” It doesn’t have to. The agent just picks one. It sends the invoice and moves on. Nobody notices until the wrong person gets billed.
That’s usually as far as the story gets before someone blames the model.
How bad data hides in plain sight
The cost of bad data shows up in different ways. A pilot that stalls. Or a decision made on the wrong version of the truth. Nobody writes “duplicate customer records” into any board deck. But the write-down at the end of the year traces back to exactly that.
This plays out almost the same way across most enterprises. The AI initiative gets funded, staffed, and shipped to a pilot group. And it stalls because the pilot forces a question the business had been avoiding for, perhaps, years.
This failure mode has a name now – the Cleanup Trap. It’s the belief that an organization can pipe fragmented, inconsistent, ungoverned CRM data into an LLM or an agent and simply patch it at the retrieval layer.
Please know that it doesn’t work that way. Retrieval can find the record faster. But it can’t decide which of three conflicting records is the truth.
Why Gen AI success starts before the LLM
Gen AI succeeds or fails based on the data foundation underneath it. It’s tempting to treat a stalled AI pilot as a modeling problem, the wrong prompt, or the wrong architecture. Most of the time it’s a foundation problem that arrived earlier and got blamed on the wrong stage.
There’s a counterargument worth taking seriously here: that “garbage in, garbage out” is an outdated maxim, because modern LLMs have enough semantic depth to make sense of unstructured text that would have broken rigid systems. That’s largely true, and it’s a real advance.
But it applies to unstructured content, a sloppy paragraph, an inconsistent log file. Not to structured business identity. An LLM can infer meaning from messy prose. But it can’t infer which of three conflicting customer IDs is the real one, because that’s not a semantic question. It’s a governance question. And no amount of model capability answers it on its own.
An LLM, or an agent built on one, is only as reliable as what it’s allowed to read. If a customer exists as three slightly different records across the CRM, the billing system, and the support desk, an agent doesn’t know that; it just picks one and acts.
If a product catalog was updated in one system, a generative summary will confidently describe a product that no longer exists that way. But again, none of this is an LLM failure. It’s a data foundation issue. An issue that was never asked to hold this much weight before.
Getting the data right isn’t a prerequisite to a good AI project now. It’s the actual project, with the model as the visible part.
The hidden costs of a weak data foundation
Cost 01 — Sunk investment
Wasted pilot spend
A pilot that looks technically sound in review often collapses in production for the same reason every time – the data it was demoed on was cleaner than the data it actually has to run on. That gap doesn’t show up until real usage exposes it, by which point the budget and the goodwill behind the pilot are already spent.
Cost 02 — Compounding effort
Rework that multiplies instead of adding up
Data quality has always followed the same shape. Catching an error at the point of entry is cheap; catching it after it’s propagated into other systems costs meaningfully more, and catching it after it’s shaped a decision costs trust as well as time. An AI pipeline doesn’t change that shape. It just moves through all three stages faster, so the expensive outcome arrives sooner.
Cost 03 — Silent losses
Revenue leakage nobody notices until it's large
Duplicate customer records don’t cause an agent to crash. They cause it to act confidently on the wrong version of a customer. And by the time someone notices the pattern, it’s shown up many times over. Once a business user catches an agent being confidently wrong, trust in the system doesn’t degrade gradually; it drops all at once, and it’s expensive to earn back.
Why bad data is more expensive in the agentic era
Bad data gets more expensive with agentic AI. A person working from a flawed report usually pauses. They notice the number looks wrong, or they ask someone before acting on it. An agent doesn’t have that instinct built in. It acts on whatever it’s given, at whatever speed it’s been built to run at, across every workflow it contacts.
That’s the real shift agentic AI introduces. The cost of bad data hasn’t changed in kind. It has changed in exposure. A duplicate record in a CRM is a data quality issue. The same duplicate feeding an agent that emails, bills, or provisions on its own is now a customer-facing incident. And it happens at whatever speed the agent runs.
How to find your hidden data costs
Before greenlighting the next pilot, it’s worth answering a few unglamorous questions:
- How many versions of your top customers or products exist across your systems right now?
- When two systems disagree on a value, is there a documented, automatic answer for which one wins, or does someone resolve it manually every time?
- And if an agent acted on the wrong version of a record today, how long would it take anyone to notice?
If those answers are uncomfortable, that’s useful information. It’s cheaper to find out now than after the pilot has failed.
What a strong data foundation looks like for AI-ready enterprises
It’s less about a single tool and more about a small set of disciplines applied consistently.
- A resolved, authoritative version of core entities - customer, product, vendor - instead of one version per system.
- Visibility into where a value came from and how confident the match behind it is.
- And validation happening at the point data enters the system.
This is the discipline Informatica MDM and Salesforce Data 360 are built to enforce when they’re implemented as a genuine foundation.
- Master data mastered once and reused everywhere downstream.
- Governance that travels with the data instead of living in a separate spreadsheet.
- And a live connection into the Salesforce and Agentforce layer so agents are reading from the same governed source everyone else is.
How LumenData builds AI-ready data foundations
As a Platinum Informatica partner and Salesforce partner, LumenData helps enterprises build AI-ready data foundations through master data management, governance, and real-time data quality. We help establish governed, lineage-aware enterprise data required for AI, analytics, compliance, and operational automation. The gap between companies getting real value from AI and companies stuck in pilot purgatory is about which ones did the less exciting work first, so the AI layer on top of it is worth the investment.
Ready to find out what your AI pilots are standing on? Talk to LumenData about building the data foundation your next AI initiative can trust.
About LumenData
LumenData is a leading provider of Enterprise Data Management, Cloud and Analytics solutions and helps businesses handle data silos, discover their potential, and prepare for end-to-end digital transformation. Founded in 2008, the company is headquartered in Santa Clara, California, with locations in India.
With 150+ Technical and Functional Consultants, LumenData forms strong client partnerships to drive high-quality outcomes. Their work across multiple industries and with prestigious clients like Versant Health, Boston Consulting Group, FDA, Department of Labor, Kroger, Nissan, Autodesk, Bayer, Bausch & Lomb, Citibank, Credit Suisse, Cummins, Gilead, HP, Nintendo, PC Connection, Starbucks, University of Colorado, Weight Watchers, KAO, HealthEdge, Amylyx, Brinks, Clara Analytics, and Royal Caribbean Group, speaks to their capabilities.
For media inquiries, please contact: marketing@lumendata.com.
Ready to find out what your AI pilots are standing on?
Authors


