Four data management tasks where machine learning earns its place
Four, and they have something in common: each is a pattern problem, on high volume, where being occasionally wrong is survivable because a person checks the uncertain cases.
Each is described below in the same order — what it does, what it needs from your data, how it fails, and who confirms.
Anomaly detection for data quality monitoring
What it does. Learns what normal looks like in a field or a feed, then flags what does not fit. A revenue figure ten times the usual. A feed that arrived with a third of its usual row count. A postcode in a country you do not trade in. A date in 1900.
This is the most reliable use of machine learning in data management, because normal is genuinely a pattern and deviation from it is genuinely detectable.
What it needs from your data. History, mostly — enough of it to know what normal is, including the seasonal shapes. A model trained on three months of data will flag every December as an anomaly.
How it fails. Alert volume. Point it at an estate nobody has cleaned and it finds four thousand anomalies in week one, which is true and useless, so people stop reading the alerts by week three, and the tool gets switched off while everybody agrees it was a good idea.
Who confirms. A data steward reviews the flags, and the threshold is tuned down until the queue is a size somebody can actually work through. Start narrow — one feed, the fields that matter — rather than across the estate.
Entity resolution and duplicate detection in master data management
What it does. Decides how likely it is that two records describe the same real thing. The same supplier spelled four ways, or the same customer with two addresses and a changed surname. This is the oldest use of machine learning on data, and matching models are genuinely good at it now, including across languages and transliterations.
What it needs from your data. Labelled pairs — real examples of "these two are the same" and "these two are different" from your own records. A few thousand, produced by people who know the domain, and this is the work that no vendor demonstration includes.
How it fails. In both directions, and they are not symmetrical. Too cautious, and the duplicates stay; too aggressive, and it merges two genuinely different customers into one record, which is very hard to unpick afterwards, because the original two are now one and the history is mixed.
Who confirms. Above the threshold, a person reviews before the merge, always. That is not a temporary caution while the model beds in; a merge is irreversible in practice, and irreversible actions keep a person in the loop permanently. The survivorship rule — which field wins from which source — is written by the business first, and the model never guesses it. A data migration is where those rules get tested hardest, because a move surfaces every duplicate at once.
Data classification and metadata tagging at scale
What it does. Reads a field, a column or a document and says what kind of thing it holds. This column looks like personal data. This document contains bank details. This free-text field is being used for delivery instructions rather than notes.
For an estate nobody has catalogued, this is the fastest route to knowing what you have, and it is the task where AI in data management most clearly saves months rather than hours.
What it needs from your data. A classification scheme first, agreed by the business — three or four levels such as public, internal, confidential, regulated. The model applies a scheme; it cannot invent one, and a scheme invented by a tool is a scheme nobody agreed to.
How it fails. Quietly, on the things that matter most. A free-text field containing occasional card numbers gets classified on its majority content and marked internal, nothing looks wrong, and the exposure is discovered by an audit.
Who confirms. Anything classified as regulated or confidential is reviewed by a person before controls are applied, and the sensitive categories are sampled every month rather than signed off once. This is where classification meets data privacy and security, and where a wrong label has consequences beyond tidiness.
Schema mapping and data lineage inference
What it does. Works out how things connect. That this column in one system is the same field as that column in another. That this report is built from those three tables. That this field is populated from a feed nobody documented.
On a large estate, inferred data lineage is often the only lineage you will get, because the documentation was never written and the people who knew have left.
What it needs from your data. Access to the systems and their logs — queries, jobs, pipeline definitions. It infers from what actually runs rather than from what was documented, which is the reason it is worth having.
How it fails. Plausible and wrong, because two columns called cust_id in different systems may hold entirely different identifiers, and a model that matches on name and shape will link them confidently. The resulting lineage diagram looks authoritative and is fiction in one corner.
Who confirms. Treat inferred lineage as a draft that a person verifies for the flows that matter — regulatory reporting, board metrics, anything on a retirement list. Nobody verifies a whole estate, and nobody needs to.
The suggest-and-confirm pattern behind every AI data management tool
Read those four back and the same shape appears in all of them. The model produces a suggestion with a confidence score, and a threshold decides what happens next: above it the system acts, and below it a person looks.
That is the pattern, and three things make it work.
The threshold is a business decision. It is set by weighing the cost of a wrong action against the cost of a person's time, and it is different for every task. Merging customer records deserves a high bar and tagging a document does not, so nobody in IT should be choosing that number alone.
The queue below the threshold needs an owner. An unowned review queue is where these projects die — the same failure as an unowned exception queue in any automated process.
Confirmations are training data. Every time a person accepts or rejects a suggestion, that is a labelled example, so capture them. Teams that do have a retraining set after six months; teams that do not are back where they started when the model drifts.