From the Trenches of AI — Vol. 14
In this series we have covered AI models that escape test environments, companies that fired people for AI and quietly rehired them, and the difference between AI built on twenty years of domain knowledge versus a third-party API integration. Every article has circled back to the same truth: AI is only as good as what it is built on.
Today that principle gets its most literal illustration yet. Because there is a category of threat to AI systems that does not come through a hacked account or a clever prompt. It comes through something far more fundamental — the data the AI learned from in the first place.
It is called data poisoning. And every business deploying AI needs to understand it.
What it is
Data poisoning is exactly what it sounds like. An attacker — or in some cases, an aggrieved party with a legitimate grievance — introduces corrupted, misleading, or manipulated data into the information an AI system learns from. The AI absorbs it. And from that point forward, it behaves in ways that serve the attacker’s interests rather than yours.
The insidious part is the invisibility. A poisoned AI does not announce itself. It continues to sound confident, helpful, and intelligent — while quietly producing wrong answers, approving things it should flag, or behaving in ways nobody anticipated. Unlike a hacked system that triggers alerts, a poisoned model can operate normally for months before anyone realizes something is wrong. By then, the damage is baked in.
Think of it this way: a fraud detection model that starts approving fraudulent transactions because someone slipped mislabeled data into its training set months earlier. Or a pricing model that recommends the wrong rates because historical data was subtly altered before it was fed into the system. The AI is not malfunctioning. It has simply learned the wrong lessons — and it is applying them, confidently, every single day.
Where it came from — and why it matters that artists started it
Most people assume data poisoning originated in cybercriminal circles. The truth is more surprising — and more instructive.
When generative AI image tools exploded onto the scene in 2022 and 2023, they were trained on billions of images scraped from the internet — including the copyrighted work of professional artists, illustrators, and photographers, taken without consent, credit, or compensation. A large majority of professional artists were deeply concerned about their work being used to train AI without permission — and with good reason. Lawsuits were filed against large AI companies including OpenAI, Meta, Google, and others. But lawsuits move slowly.
So artists fought back with technology.
Researchers at the University of Chicago, led by Professor Ben Zhao, built two tools. The first, Glaze, applied an invisible “style cloak” to artwork before uploading — causing AI models to misread the artist’s style entirely. The second, Nightshade, went further: it altered image pixels in ways invisible to the human eye but devastating to AI training. A poisoned image of a dog would cause the AI to learn it as a cat. A car might become a cow. A hat might become a toaster. With as few as 50 to 300 poisoned images, an entire AI model’s behavior could be visibly distorted.
The intent was rights protection — to make training on unlicensed data expensive enough that licensing it became the more attractive option.
What the artists started as self-defense, the broader threat landscape quickly absorbed as a weapon. Once the world saw that a small number of corrupted data points could corrupt an entire AI system — invisibly, persistently, and at scale — the implications went far beyond art. They applied to any AI trained on any data that any adversary could reach.
How bad is it?
The scale of the problem is sobering. Research has consistently shown that it takes a surprisingly small number of corrupted data points to compromise an AI system in meaningful ways — far fewer than most business leaders would guess. A handful of manipulated documents, a fraction of a percent of mislabeled training data, a small number of poisoned images — any of these, introduced at the right point in the right pipeline, can shift an AI’s behavior in ways that persist long after the poisoning itself is forgotten.
What makes this particularly difficult to manage is that a poisoned model rarely fails obviously. It continues to perform well on standard tests. It continues to produce fluent, confident outputs. The deviation from correct behavior is often subtle enough to go undetected for weeks or months — until the consequences accumulate to the point where they cannot be ignored.
The threat is not only external. Insider threats — employees with access to data pipelines who subtly corrupt training data — are widely considered one of the most underestimated and difficult-to-detect risks in AI development. The attack surface today spans not just training data but the live knowledge bases AI uses to answer questions in real time, the memory AI agents carry between sessions, and the third-party tools they connect to.
How AI companies are fighting back
Defense against data poisoning is not a single solution. It is a discipline — and a layered one.
The most important principle is provenance: knowing exactly where your data comes from, who had access to it, and being able to prove it at every step. Many poisoning attacks succeed precisely because organizations source data from third parties or open repositories without verifying its integrity. Treating training data the way a courtroom treats evidence — documented, verified, traceable — is the standard Carnegie Mellon’s Software Engineering Institute now advocates formally.
Beyond provenance, effective defense requires continuous monitoring rather than one-time audits. Poisoned content is designed to look legitimate and pass standard quality checks. It hides in plain sight until triggered. The only posture that reliably catches it is ongoing vigilance across the entire data pipeline — before training, during training, and after deployment.
Why this matters for moving and storage
The moving and logistics sector is not building foundation models. But that does not mean this is someone else’s problem.
Consider the AI tools increasingly appearing in our operations — tools that learn from historical shipment data, customer records, pricing histories, claims outcomes, and operational logs. The quality and integrity of that data determines everything the AI will ever recommend, flag, estimate, or decide.
A pricing AI trained on manipulated rate data will recommend wrong prices — confidently and consistently. A claims AI trained on mislabeled outcomes will make wrong assessments. A routing AI fed corrupted data will generate inefficient plans. None of these failures will announce themselves. They will simply look like the AI doing its job — until the pattern of wrong answers becomes impossible to ignore.
Government moves — military, GSA, and Department of State
The data flowing through government household goods programs — military DP3 shipments, GSA civilian relocations, and Department of State assignments moving diplomats and foreign service officers around the world — represents some of the most sensitive, structured, and authoritative data in our industry. Personnel details, location information, shipment records, and logistics timelines all flow through the same platforms. An AI trained on corrupted government move data does not just make bad operational recommendations. It makes bad recommendations with compliance, contractual, and in some cases national security consequences. The trust placed in TSPs and technology providers to handle this data responsibly is not abstract — it is a program requirement, and it is tested continuously.
Corporate relocation
When a company moves an employee from one city to another, the data involved goes well beyond a shipment manifest. Home addresses, salary-linked allowance calculations, family information, HR records, lease and mortgage details, and exception approvals all travel through relocation platforms. Corporate clients — particularly large enterprises — are increasingly scrutinizing the data governance practices of their relocation partners and technology vendors. An AI producing subtly wrong outputs in a corporate relocation context does not just create operational problems. It creates the kind of errors that surface in client audits, procurement reviews, and contract renewals. The reputational cost of a pattern of wrong recommendations in this segment can be significant and swift.
Residential moves
It is tempting to think that residential moving data — names, addresses, moving dates, inventory lists — is lower stakes than government or corporate work. It is not. Residential moving companies collect highly personal information about thousands of families each year: where they live now, where they are going, when the home will be empty, and what they own. That data profile is exactly what identity thieves, burglars, and fraud networks are looking for. An AI that has learned from corrupted residential data — whether in pricing, scheduling, or inventory management — is not just operationally unreliable. It is a potential liability to the families who trusted the company with some of the most sensitive details of their lives.
The moving industry has always understood that data quality matters. Inaccurate inventory counts lead to disputes. Wrong addresses lead to failed deliveries. Bad weight estimates lead to binding estimate problems. What data poisoning adds to that familiar equation is intent — the possibility that the data errors are not accidental but engineered, not random but targeted, not fixable by retraining but persistent until the underlying corruption is found and removed.
The common thread across every segment — government, corporate, and residential — is this: moving companies are data businesses, whether they think of themselves that way or not. And in a world where the data AI learns from can be deliberately corrupted, treating that data with the rigor, governance, and security discipline it deserves is not just a compliance obligation. It is a business imperative.
The principle that does not change
The defense against data poisoning is not fundamentally different from what good data governance has always required. Know where your data comes from. Control who can touch it. Validate it continuously. Test your AI against known correct answers regularly.
These are not new principles. They are the principles that every well-run operation in our industry already applies to its shipment records, its claims files, and its customer information. The arrival of AI does not replace those principles. It raises the stakes for ignoring them.
The intelligence of an AI system is not separate from the data it learned from. It is inseparable from it. Corrupt the data and you corrupt the intelligence — quietly, persistently, and at a scale that makes the original corruption very difficult to undo.
In the age of AI, data governance is not an IT obligation. It is the foundation of everything your AI will ever do. And in an industry built on the trust of the people and organizations who hand us the details of their lives and livelihoods, that foundation is not optional.
“From the Trenches of AI” is an ongoing EDC® LinkedIn series exploring artificial intelligence through the lens of an industry that moves people, not just data.
About the Author
Diana Corona
Co-Founder, President & CEO — Enterprise Database Corporation (EDC®)
Diana Corona co-founded EDC® over 25 years ago and has spent her career building software purpose-built for the moving and storage industry. Under her leadership, EDC® has grown into one of the most trusted technology partners in the space — serving moving companies of all sizes across residential, commercial, military, government, international, and specialty move types. She writes on topics at the intersection of technology, operations, and the future of the moving industry.



