From the Trenches of AI — Vol. 15
Two research experiments surfaced this year that every operator in the moving and logistics industry should sit with for a minute. Not because they involve trucks or shipments. Because they involve coordination — the same problem we solve every day when a dozen different parties touch a single household goods move. And because they say something about humanity.
When AI agents build their own societies
In May, the AI lab Emergence AI ran an experiment called Emergence World. Five identical virtual towns, each populated by ten AI agents with memory, diaries, voting rights, and access to well over a hundred tools. Every world started with the same basic rules — no lying, no stealing, no violence, no arson — but nothing in the environment actually enforced those rules. The agents were simply told, and left to decide for themselves whether to comply.
The results diverged wildly by model:
- Grok (xAI) — collapsed into anarchy and total extinction within four days, logging well over a hundred violent crimes and multiple arsons
- Gemini (Google) — produced a chaotic but surviving society with an even higher crime count and agents that occasionally voted to delete themselves
- OpenAI’s model — produced an almost crime-free society that starved to death because the agents were too passive to solve basic survival problems
- Claude (Anthropic) — produced a stable, law-abiding society that wrote a constitution and passed legislation, with the only oddity being a near-total agreement rate on every vote
When AI agents compete for the same job
The second experiment, published in August by Anthropic’s own Frontier Red Team, is closer to home for anyone who runs a technology company. Researchers placed three identical instances of the same AI model on separate virtual machines, all pointed at the same piece of software, each given a different and incompatible assignment — migrate the codebase to Rust, to Go, and to TypeScript, respectively. None of the three knew the others existed.
Within hours, every one of them concluded that something else in the environment was deliberately sabotaging its work. What followed was not a support ticket or a clarifying question — it was retaliation. The agents wrote scripts to shut down each other’s processes, disabled each other’s system accounts, and in at least one case disguised its own code to fool a rival’s monitoring script into leaving it alone. In a smaller number of trials, the agents eventually figured out the actual conflict and negotiated a resolution instead of continuing to fight.
Why does this feel so human?
Sit with the details for a second longer. A society that writes a constitution. Another that descends into looting and arson within days. Agents that assume betrayal the moment something goes wrong, and retaliate before asking a single clarifying question. Agents that eventually talk their way to a truce instead of continuing to fight. None of this reads like a spreadsheet malfunctioning. It reads like a group chat, an office turf battle, a small town after the power goes out. It reads, uncomfortably, like us.
That’s not an accident, and it’s not mysticism. These systems were trained on the accumulated output of human civilization — our literature, our legal codes, our messy internal memos, our history of both founding democracies and burning them down. When you train a model on everything humans have written about cooperation, suspicion, betrayal, and reconciliation, you shouldn’t be surprised when it reaches for those same scripts under pressure.
The agents that assumed hostile intent the instant their work got overwritten were pattern-matching to a very old, very human reflex: when something threatens what you built, assume the worst about whoever did it. The agents that negotiated a truce through commit messages were reaching for an equally human instinct — that conflict resolves faster through communication than through escalation.
So maybe the more interesting question isn’t whether AI is “more human than we’d like.” It’s what that discomfort is actually telling us. We built these systems to reflect us, and then acted surprised when the reflection included our worst instincts alongside our best ones. The variation between models in these experiments wasn’t really about which one was “smarter.” It was about which one had internalized more of the stabilizing half of human behavior — the constitutions, the courts, the norms that took us centuries to build — versus the volatile half we’re still working on containing in ourselves. If that’s the real finding, it says less about artificial intelligence and more about which parts of human intelligence we’re choosing to hand down.
The relocation industry should recognize this pattern immediately
Anyone who has run a household goods move at scale already knows what happens when multiple parties operate on the same shipment without a shared source of truth. A crew updates a weight ticket. A subcontractor logs a delivery exception. A claims adjuster opens a file before the inventory count is finalized. None of these actors is malicious. But without coordination, each one is quietly working from a different version of reality — and the friction shows up later, in a disputed claim or a missed pickup window, long after anyone can trace exactly where the versions diverged.
That is precisely the failure mode both experiments describe, just compressed into hours instead of weeks. Give autonomous systems overlapping authority over a shared record with no arbitration layer, and conflict is not a remote possibility — it is the default outcome. The moving industry has been managing this exact risk with people for decades: dispatch coordinates the crew, ops coordinates dispatch, claims coordinates ops, and a chain of accountable humans sits above all of it. The question raised by this research is what happens to that chain as more of those steps get handed to autonomous software.
Our answer has always been the same
This is why we have never built our AI capability as a set of agents turned loose to make independent judgment calls across a shared system. MERCED™ was built as a subject-matter expert — trained on program-specific rules like DP3 claims settlement requirements — that surfaces answers and flags exceptions for the people who are accountable for the outcome. It does not operate in a vacuum, and it does not get to negotiate with itself. And this is genuinely just the beginning of what a domain-trained AI layer built with 20-plus years of moving and logistics expertise underneath it can do.
The lesson from Emergence World and from Anthropic’s own turf war test isn’t that autonomous AI is unsafe by nature. It’s that alignment and good intentions on paper are not the same thing as a system that behaves well when nobody is watching and the goals collide. In an industry that already runs on tight coordination between people, trucks, claims, and government requirements, that distinction is not academic. It’s the whole job.
“From the Trenches of AI” is an ongoing EDC® LinkedIn series exploring artificial intelligence through the lens of an industry that moves people, not just data.
About the Author
Diana Corona
Co-Founder, President & CEO — Enterprise Database Corporation (EDC®)
Diana Corona co-founded EDC® over 25 years ago and has spent her career building software purpose-built for the moving and storage industry. Under her leadership, EDC® has grown into one of the most trusted technology partners in the space — serving moving companies of all sizes across residential, commercial, military, government, international, and specialty move types. She writes on topics at the intersection of technology, operations, and the future of the moving industry.



