A few months ago, I decided to do a digital deep clean. Delete the apps I don't use, close the accounts I forgot I had, get down to some Marie Kondo-esque minimum viable digital footprint. I opened up my password manager to take inventory and very nearly closed my laptop and walked away from technology forever: 241 accounts. Some research puts the average person at roughly 240 accounts requiring a password and, apparently, I am nothing if not perfectly average.
What stopped me from actually deleting anything was that I didn't want one account for everything. I wanted a burner email for newsletters, a real email for colleagues, another email for friends & family, a work Slack persona and a personal one, a finance app that has no idea what my Spotify habits look like. The clutter was annoying...but the separation, somehow, I still wanted to keep.
So I started paying attention to why I'd drawn each of those lines in the first place, and a pattern showed up that I wasn't expecting.
Why We Build Walls
Most of those 241 accounts don't actually have anything to do with identity. They're my bank, my mortgage servicer, four different airlines, a dentist's patient portal I logged into exactly once. Each one is its own island because it has to be, you can't log into United with your Chase credentials, nobody built that bridge, and probably nobody should. That kind of segmentation was never a choice I made, the internet just hands you a padlock for every door you walk through.
The accounts that are most interesting to me are the small minority I built on purpose: the burner email, the internet handle with no last name attached, the work Slack persona that never talks about what I did over the weekend. Nothing forced those on me. I built each wall myself, and every one of them followed a line that was already in my head: work self vs. weekend self, real name vs. internet handle. I think this type of segmentation is a very human way of trying to avoid what researchers Alice Marwick and Danah Boyd call “Context Collapse,” which is the academic term for the specific dread you feel right after you hit "post" and remember your boss is also your Facebook friend.
In 1967, the computer scientist Melvin Conway noticed something kind of similar happening inside companies: any organization that designs a system will produce a design that mirrors its own communication structure, not necessarily the structure the problem actually called for. This phenomenon became known as Conway’s Law. A company with three separate teams tends to ship software with three separate parts, whether or not the underlying problem breaks cleanly into three. My burner email and my work Slack persona run on the same rule at a much smaller scale: the shape of the divider ends up stamped onto whatever it divides, even when nobody consciously designed it that way.
Enter the Robots
Which brings me to the thing I actually sat down to write about: AI agents. You've probably noticed the pattern already, a coding agent, a research agent, a scheduling agent, a customer-support agent, each with its own name, its own scope, its own little kingdom. The obvious question is why we don't just build one agent that does everything. And the answer to how much compartmentalization is “just right” isn’t purely academic, it actually matters for two real world reasons: cost and quality.
- Cost: Splitting a task across agents isn't free. One study out of UIUC found multi-agent systems can burn through 4 to 220 times more tokens than a single agent doing the same job.
- Quality: Extra spending doesn't reliably buy extra accuracy, since every handoff between agents is, theoretically, a fresh opportunity for one agent's mistake to get treated as fact by the next.
Part of the reason for a multi-agent architecture comes down to true engineering constraints. Context windows are finite, and distributing work across specialized agents keeps each one's context focused and relevant. Security and compliance boundaries sometimes require isolating agents in separate, independently governed environments, the same way you'd never want the intern's badge to open the CFO's office.
But I argue that there’s another element at play here, which is the same Conway’s Law above that helps explains why we humans compartmentalize our systems to begin with. The companies building those systems already think in terms of specialized teams, specialized tools, specialized ownership. Their own communication structure runs through separate product teams, separate codebases, separate Slack channels for separate concerns. So when they sit down to design something agentic in the new world, they follow the same pattern because that’s how they have always operated. Whatever divides the system stamps its own shape onto it, whether or not the task actually earns that shape. But there's no real reason to assume the mathematically optimal split of a task across language models runs along those same lines. Humans are really good at judging whether the output is right. We're not obviously the species best positioned to decide where the seams should go.
There's already a preview of what happens when an LLM gets to draw its own lines instead, and it’s pretty interesting. Not necessarily better than how humans prefer to segment, but certainly radically different. Mixture-of-experts architectures like Mixtral split a language model into dozens of specialized sub-networks, and it's tempting to picture tidy categories a person would recognize: one expert for math, one for code, one for poetry, etc. That isn't how it plays out though. When researchers checked which “expert” agents fired on which kinds of text in the Mixtral model, they found that syntax mattered far more than subject matter, and a single sentence would scatter across many different experts instead of routing to one topic specialist. Which kind of makes sense because, after all, these are language models we’re talking about. Likewise, the most efficient way to split a complicated task across multiple agents might not track job functions at all. It might be some split none of us would ever think to make, because the categories we reach for come from how people organize work, not from whatever arrangement actually gets the best result out of a set of statistical models handing work to each other.
So Which Is It?
There appear to be two separate forces pushing agents toward compartmentalization. The first is technical: context windows fill up, permissions need to be separated, some tasks really do parallelize better across isolated workers. The second is inherited: teams that have spent decades being told to decompose everything into small, single-purpose, independently owned units reach for that same decomposition by default, whether or not this particular task earns it.
In June 2025, two huge agent builders tried to settle this question for themselves. Cognition, the company behind the coding agent Devin, published a post called "Don't Build Multi-Agents," arguing that multi-agent setups are fragile because sub-agents make decisions in isolation, and a single well-engineered agent with proper context management will usually beat a fleet of specialists. A day later, Anthropic published a piece on how it built its own research system, describing how an orchestrator coordinating multiple Claude sub-agents outperformed a single agent by a wide margin on complex research tasks. Two of the top engineering teams in the world reached opposite conclusions about whether splitting agents apart was worth doing at all.
That disagreement is the clearest sign available that nobody, including the people whose job it is to know, has a settled first-principles answer for how much compartmentalization the work actually requires. If the experts building these systems needed a very public disagreement to sort constraint from habit, the rest of us are probably not doing any better when we assume more agents automatically mean better engineering.
I don't know the real ratio between the two, and I doubt anyone fully does yet. My guess is that as context windows keep growing and coordination keeps getting cheaper, the genuinely necessary slice of compartmentalization will keep shrinking, and we'll find out how much of the current agent sprawl was load-bearing and how much was just habit carried over from human ways of working. Because I can't help but wonder if, once we've learned to divide things up, we stop asking ourselves whether the thing in front of us actually needed dividing.