Two of the most expensive mistakes I watch organizations make are mirror images of each other, and both get made by smart people.
The first: two teams step on each other, releases hurt, nobody knows who owns what. So someone proposes an event-driven rewrite. Queues, topics, asynchronous everything. The theory is that if services stop calling each other directly, people will stop needing to talk. Instead the coupling moves, and now the confusion has a message broker in the middle of it. Segment's engineers wrote honestly in 2018 about going from a monolith to microservices and back again. Their real problem was a lack of tooling to test and deploy many services at once. Splitting the system multiplied it.
The second runs the other way. Delivery is slow, changes break things in surprising places, so a governance body appears. An architecture review board, a weekly sync, a new intake process. But if the real issue is a shared database and a dependency graph nobody can draw, no meeting removes a single dependency. You have taxed every change while leaving it just as dangerous.
The research is better than most people realize. Conway's 1968 paper, rejected by Harvard Business Review before Datamation ran it, only claimed that systems come to resemble the communication structures of the organizations that build them. Colfer and Baldwin reviewed 142 empirical studies and found mirroring holds most of the time, not always. Microsoft's Windows Vista study found organizational metrics predicted failure-prone binaries with around 86% precision, beating code churn, complexity and coverage. And Ruth Malan's line has stayed with me: if the architecture of the system and the architecture of the organization are at odds, the organization wins.
So before proposing either fix, I ask one question, straight from DORA's work on loosely coupled architecture.
Can this team make a large change, test it, and deploy it without permission from anyone outside the team?
If yes and things are still slow, you have a coordination problem, and no rewrite helps. If no, you have structural coupling, and no governance body helps. Companion check: imagine both teams in one room with unlimited bandwidth. If the problem evaporates, it was information flow. If it survives, it is baked into the system.
Some coordination cannot be designed away. Safety, regulation and shared physical constraints genuinely tie teams together. Then the goal is to make coordination cheap and explicit with clear contracts and versioned interfaces rather than pretending you can remove it.
Reorgs and rewrites are both expensive, and both have poor track records when aimed at the wrong target. McKinsey's numbers on reorganizations are sobering enough that I treat them as a last resort with a review date, not a first move.
Coordination problems vs architecture problems: a research overview
TL;DR
- The evidence is strong that organizational structure and system architecture mirror each other (Conway's Law), and that org-level factors predict software quality better than code-level factors, so misdiagnosing a coordination problem as an architecture problem (or vice versa) reliably wastes money. Microsoft's Windows Vista study found organizational metrics predicted failure-prone code with 86.2% precision and 84% recall, beating code churn, complexity, coverage, dependencies and pre-release bugs.
- The two classic failure modes are well documented: (a) splitting into microservices to fix what was really a team coordination or tooling problem, producing a "distributed monolith" (Segment, Amazon Prime Video), and (b) standing up governance bodies, review boards and coordination ceremonies to manage what is really structural coupling in a shared codebase or database.
- The best diagnostic test in the literature is DORA's: can a team make large changes, test, and deploy without permission from, or coordination with, other teams? If yes and there is still friction, it is a coordination or process problem; if no, it is structural coupling that no amount of meetings will fix.
Key findings
Conway's Law is real but often over-applied
Melvin Conway's 1968 Datamation paper "How Do Committees Invent?" states: "organizations which design systems (in the broad sense used here) are constrained to produce designs which are copies of the communication structures of these organizations." It was rejected by Harvard Business Review for not proving its thesis before Datamation published it in April 1968. It is a correspondence claim, not a causal one: it does not say which direction the causation runs. Fred Brooks named it "Conway's Law" in The Mythical Man-Month. [1][2]
The mirroring hypothesis has solid but not universal empirical support
- Colfer and Baldwin (2016, Industrial and Corporate Change) reviewed 142 empirical studies. An earlier version of the work reported the hypothesis supported in 69% of cases. They found mirroring "prevalent but not universal" and identified exceptions, notably in open collaborative projects where a majority of descriptive studies (56%) did not support mirroring. Of the 142 studies, 68 were descriptive (correlations between technical dependencies and organizational ties) and 74 were normative (evaluating success or failure of mirrored vs unmirrored systems). [3]
- MacCormack, Rusnak and Baldwin (2012, Research Policy) compared paired open-source and proprietary products of similar function, treating firms as tightly coupled and open-source communities as loosely coupled. The published paper states verbatim: "The magnitude of the differences is substantial, up to a factor of six, in terms of the potential for a design change in one component to propagate to others." (The earlier HBS working paper reported "up to a factor of eight"; both versions found the loosely coupled organizations produced significantly more modular designs.) [4]
Organizational metrics predict defects better than code metrics
Nagappan, Murphy and Basili (ICSE 2008, "The Influence of Organizational Structure on Software Quality") studied Windows Vista: 3,404 binaries, over 50 million lines of code, several thousand engineers, and post-release failures measured over the first six months after release. They built eight organizational metrics: Number of Engineers (NOE), Number of Ex-Engineers (NOEE), Edit Frequency (EF), Depth of Master Ownership (DMO), Percentage of Org contributing to development (PO), Level of Organizational Code Ownership (OCO), Overall Organization Ownership (OOW), and Organization Intersection Factor (OIF). Both step-wise regression and principal component analysis retained all eight (none was redundant). Using these as predictors of failure-proneness, the model achieved 86.2% precision and 84.0% recall, compared with code churn (78.6%/79.9%), code complexity (79.3%/66.0%), dependencies (74.4%/69.9%), code coverage (83.8%/54.4%), and pre-release bugs (73.8%/62.9%). This is one of the strongest single pieces of evidence that where your organization is fractured, your defects concentrate. [5][5]
The distributed monolith is the signature of a misdiagnosis
A distributed monolith is a system split into separately deployed services that remain so tightly coupled they must be tested, deployed and changed together. It typically emerges when teams decompose along existing code boundaries rather than domain boundaries, keep a shared database, or keep the same group coordinating every release. As one write-up puts it, if development teams do not align with service boundaries, "you can decompose the code, but if the same group of people still coordinates every release, the coupling hasn't changed; it's just moved to a different layer." The tell-tale signs are synchronized deployments across services, shared databases, chatty synchronous calls, and a minor feature that requires touching many services. [6]
Details
Documented cases: architecture deployed against an organizational or scaling problem
Segment (2018, "Goodbye Microservices" by Alexandra Noonan). Segment moved from a monolith to microservices, then back. The move to microservices, combined with the decision to use separate repos, led to "exploding defect rates" and "plummeting velocity," and the small team became "mired in complexity." The core problem was operational: they lacked tooling for testing and deploying microservices when bulk updates to shared libraries were needed, so a single database or shared-library change could force every service to upgrade. They consolidated back into a single repo with a single shared library version and an aggregator called Centrifuge that replaced individual per-destination queues. This was primarily a tooling and coordination problem, not something the microservice split solved; splitting made it worse by multiplying the coordination surface. [7]
Amazon Prime Video (2023). The audio/video quality monitoring team moved a serverless (AWS Step Functions and Lambda) microservices setup to a monolith and reported over 90% infrastructure cost reduction. Two hard limits forced the change: they hit account limits on Step Function state transitions and hit a scaling ceiling at around 5% of expected load, and the cost was prohibitive. Important caveat: this was one internal service, not all of Prime Video, and Adrian Cockcroft (former AWS VP) argued it was a microservice refactoring step, not a wholesale rejection of microservices. David Heinemeier Hansson used it to argue microservices and serverless were oversold. Treat the sweeping "microservices are dead" takes as opinion, not evidence; the defensible lesson is that the original architecture was chosen for the wrong problem.
The reverse failure mode: coordination structures applied to structural coupling
Architecture review boards (ARBs) and Enterprise Architecture governance. The recurring, widely reported failure mode is that ARBs become bottlenecks: when every decision routes through the board, delivery slows and teams route around it. Critics note ARBs conflict with Lean flow and Agile autonomy principles, and that up-front board approval blocks value flow and can cause rework. Where this bites hardest is when the underlying problem is structural coupling: adding a coordination layer on top of a shared database or tangled dependency graph does not remove the coupling, it just taxes every change with a meeting. The consensus fix in the practitioner literature is to make ARBs selective (tiered scope), request-driven and advisory, or to embed architects in teams rather than gate work through a board.
The Spotify model (squads, tribes, chapters, guilds). The October 2012 whitepaper "Scaling Agile @ Spotify with Tribes, Squads, Chapters & Guilds" by Henrik Kniberg and Anders Ivarsson described Spotify at approximately 250 employees across about 30 teams in three cities, and explicitly stated: "This is not a recipe. It is a snapshot of what we are doing right now." The most common failure when copying it is adopting the vocabulary without the decision-making autonomy: a "squad" that still needs multiple approvals to deploy is not autonomous. Reported failure modes include tribes growing too large, chapters and guilds becoming under-resourced afterthoughts, duplication across tribes, and cross-tribe coordination becoming harder than in a traditional structure because no mechanism existed for it. Spotify itself moved on from the model. [8]
Conway's Law, the inverse Conway maneuver, and Team Topologies
Inverse Conway maneuver. Coined by Jonny Leroy and Matt Simons of ThoughtWorks in the December 2010 Cutter IT Journal, popularized via the ThoughtWorks Technology Radar. It recommends evolving team and org structure to promote a desired architecture. The critique, discussed by Martin Fowler and James Lewis on the ThoughtWorks podcast, is that Conway's Law is a strong force that can be leveraged or fought, and that reflexively invoking "Conway's Law" for every organizational question is reductive. Gregor Hohpe, author of The Software Architect Elevator, warns explicitly against reducing all organizational aspects to this one observation. There is no strong quantitative evidence that deliberately reorganizing people to force an architecture reliably works; it is a practitioner heuristic backed by the mirroring correlation, not by controlled studies. [9][10]
Team Topologies (Skelton and Pais, 2019). Defines exactly four team types (stream-aligned, enabling, complicated-subsystem, platform) and three interaction modes (collaboration, X-as-a-Service, facilitating). The constraint is team cognitive load; the goal is to keep it within human limits. Teams are split along "fracture planes," natural seams in the system (domain or bounded context, change cadence, regulatory compliance). The diagnostic relevance: the book's test for breaking up a monolith is whether the resulting architecture supports more autonomous teams with reduced cognitive load. It explicitly says do not "solve" cognitive load by adding process; change boundaries or services instead. Ruth Malan's line, quoted throughout the field: "If the architecture of the system and the architecture of the org are at odds, the architecture of the org wins." [11]
Coordination cost research
Brooks's Law. From The Mythical Man-Month (1975): "adding manpower to a late software project makes it later." The communication-channel formula is n(n-1)/2: a team of 5 has 10 channels, a team of 10 has 45. Brooks himself called the law an "outrageous oversimplification." The three mechanisms are ramp-up time, communication overhead, and limited divisibility of tasks. Critiques and refinements exist: Di Penta, Harman et al. (2007, University College London) found the impact of Brooks's Law can be less severe in maintenance projects where there is flexibility in team assignment; Eric Raymond argued open source partly falsifies the assumptions because work is separable into parallel subtasks and only the small core group pays the full Brooksian overhead. [12]
Cross-team dependency cost. The mechanisms are queueing delay (a request enters another team's backlog with its own priorities), priority conflicts (structural, not fixable by "better communication"), batch-size inflation, and hidden coupling. Reported wait times in enterprise settings are commonly cited as 2 to 4 weeks per dependency, and dependencies often sit in sequence rather than in parallel, so a feature that should take weeks takes months. These figures come mostly from vendor and practitioner sources rather than peer-reviewed work, so treat the specific numbers as indicative rather than established. [13]
DORA / Accelerate on architecture. DORA's research consistently finds that loosely coupled architecture is one of the strongest predictors of continuous delivery and software delivery performance. The current DORA construct (dora.dev) is measured by whether teams can: make large-scale changes to the design of their systems without permission from somebody outside the team or depending on other teams; complete work without needing fine-grained communication and coordination with people outside the team; deploy and release on demand, independently of the services they depend on or that depend on them; do most testing on demand without requiring an integrated test environment; and deploy during business hours with negligible downtime. In the book Accelerate (2018), the two core survey items are "We can do most of our testing without requiring an integrated environment" and "We can and do deploy or release our application independently of other applications/services it depends on." DORA states verbatim: "The 2021 DORA report (p. 26) shows that a loosely coupled architecture is one of the strongest predictors of successful continuous delivery: elite teams who meet their reliability targets are three times more likely to have adopted such an architecture than low-performing teams." The 2022 report (p. 31) found high performers who meet reliability targets are 40% more likely to have systems based on a loosely coupled architecture (alongside 46% more likely to practice continuous delivery, 39% continuous integration, 33% version control). Accelerate found architecture was the largest contributor to continuous delivery in the 2017 analysis. [14]
DORA 2024 and 2025 on platform engineering and org structure. The 2024 report found internal developer platform users had about 8% higher individual productivity and 10% higher team performance, and organizational performance rose about 6% with a platform, but also a surprising 8% decrease in throughput and 14% decrease in change stability, so platforms are not a free win and follow a "J-curve." The 2024 report also found, verbatim: "Unstable organizational priorities cause meaningful decreases in productivity and substantial increases in burnout. This negative impact is highly resistant to mitigation and persists even in environments with strong leaders and high-quality documentation." The 2025 report ("State of AI-assisted Software Development") found, per Google Cloud's blog, "90% of organizations have adopted at least one platform" and that 76% have dedicated platform teams, and introduced the theme that "AI is an amplifier" of existing strengths and dysfunctions, with platform quality and value stream management acting as force multipliers (when platform quality is high, AI's effect on organizational performance becomes strong and positive; when it is low, that effect is negligible).
Cognitive load and Dunbar's number
Team Topologies leans on team cognitive load as the real constraint and cites Dunbar's number (around 150, with inner layers around 5, 15, 50) for trust-boundary sizing. The empirical basis for Dunbar's number is genuinely contested. Critics using larger datasets and modern statistics (Lindenfors, Wartel and Lind, and a separate Stockholm reanalysis) argue it "doesn't stand up to scrutiny," with confidence intervals so wide (one reanalysis ranged up to 520 and produced a point estimate as low as 71) as to be nearly useless. Dunbar has defended it, attributing the discrepancy to critics using the wrong regression method (least squares rather than reduced major axis). For a research-grounded post, cognitive load is a defensible constraint; the specific number 150 should be treated as a rule of thumb, not a law.
Reorganization failure rates
- McKinsey Global Survey (2014): only 23% of reorganizations met their objectives and improved performance; 44% bogged down in implementation and were never finished; a further 23% were implemented but did not meet objectives; 10% significantly impaired company performance.
- McKinsey (via HBR 2016, "Getting Reorgs Right"): more than 80% of reorgs fail to deliver the value they were supposed to in the time planned, and 10% cause real damage.
- Following McKinsey's nine "golden rules" is associated with a substantially higher success rate. These figures cover corporate reorgs broadly, not just engineering reorgs, but they set the base rate: reorganizing is a high-risk intervention that should not be the first thing you reach for.
When coordination is irreducible
Some coupling cannot be designed away: safety-critical systems, shared physical constraints, and regulated domains impose genuine cross-team dependencies. Team Topologies treats regulatory compliance as a legitimate fracture plane precisely because it forces coordination. The honest framing is that the diagnostic question is not "coordination or architecture" in every case; sometimes the coupling is essential, and the right move is to make coordination cheap and explicit (clear contracts, versioned interfaces, well-defined team APIs) rather than to pretend you can eliminate it. [15]
Recommendations
Staged diagnostic approach, from cheapest to most invasive:
- Run the DORA independence test first. Ask each team: can you make a large change, test it, and deploy it without permission from or coordination with another team? This is the single best-validated discriminator in the literature. If the answer is no, you have structural coupling, and reorganizing people or adding ceremonies will not fix it. Change the architecture (or the ownership boundaries) instead.
- Apply the "same room" heuristic. If you put the two teams in one room with unlimited bandwidth, would the problem disappear? If yes, it is a coordination or information-flow problem: fix communication, ownership clarity, or team boundaries. If no, the problem is baked into the system structure (shared database, tangled dependency graph) and needs an architectural fix.
- Map the dependencies before you reorganize. Use dependency mapping, value stream mapping, or a design structure matrix (DSM) to compute propagation cost (the share of the system a single change can reach). High propagation cost plus slow delivery points to architecture; low propagation cost plus slow delivery points to process or coordination.
- Match the intervention to the diagnosis. If it is coordination: clarify ownership, reduce handoffs, align or co-locate teams, and only then consider structural change. If it is architecture: change the boundaries (fracture planes), decouple the database, or introduce backward-compatible versioned APIs so teams can deploy independently, and change team structure to match (inverse Conway) rather than bolting a governance layer on top.
- Treat governance bodies and reorgs as last resorts with a kill switch. Given the base rates (only about 23% of reorgs succeed; ARBs reliably become bottlenecks), set explicit success metrics and a review date up front, and prefer request-driven advice and embedded architects over mandatory approval gates.
Thresholds that would change the recommendation: if teams already deploy independently but still coordinate heavily, the problem is process, not structure. If propagation cost is high and lead time is dominated by waiting on other teams, invest in decoupling before any reorg. If the domain is safety-critical or regulated, accept irreducible coordination and optimize for making it cheap and explicit.
The best candidate diagnostic questions
- "Can team A ship without team B's calendar?" (Can a team make a large change, test, and deploy without permission from or coordination with another team?) This is the DORA loosely-coupled construct, the best-validated predictor of delivery performance. If no, it is architecture.
- "If the two teams were in the same room with unlimited bandwidth, would the problem disappear?" If yes, coordination or communication; if no, structural coupling. This maps directly to the definition of coupling: shared state, behavioral or temporal dependency.
- "If we change one component, how much of the system can it affect?" (propagation cost via DSM). High propagation cost is the fingerprint of an architecture problem masquerading as a coordination problem.
- "Does breaking this apart give us more autonomous teams with lower cognitive load, or just more moving parts to coordinate?" The Team Topologies test. If it produces more coordination, you are heading for a distributed monolith.
- "Is the coordination essential (safety, regulation, shared physical constraint) or accidental (tooling, ownership, history)?" Only accidental coordination can be designed away; essential coordination should be made cheap and explicit.
Caveats
- Much of the cross-team dependency wait-time data (2 to 4 weeks per dependency) and some Team Topologies operational claims come from vendor and practitioner sources, not peer-reviewed research. They are plausible and widely repeated but should be flagged as anecdotal.
- The Amazon Prime Video case is frequently misrepresented; it was one internal service, and credible microservices proponents (notably Adrian Cockcroft) dispute the "monolith won" narrative.
- The mirroring hypothesis is supported in roughly 69% of studies, not universally; open-source projects are a notable exception (56% of descriptive studies of open collaborative projects did not support it).
- Dunbar's number is scientifically contested and should be treated as a heuristic, not a hard limit.
- Reorg failure statistics (McKinsey's 23% success) cover corporate reorgs broadly, not engineering-specific reorgs; the base rate is suggestive, not a precise prediction for a software org.
- Conway's Law is a correspondence claim; the inverse Conway maneuver's effectiveness is supported by correlation and practitioner experience, not controlled experiments.
- The DORA "factor" and effect-size figures are cross-sectional survey correlations, so they show strong association, not proven causation; DORA reports use language like "predictor" advisedly.
- Wikipedia — https://en.wikipedia.org/wiki/Conway%27s_law
- Lawsofsoftwareengineering — https://lawsofsoftwareengineering.com/laws/conways-law/
- Oxford Academic + 2 — https://academic.oup.com/icc/article-abstract/25/5/709/2198460
- Harvard University — https://dash.harvard.edu/server/api/core/bitstreams/7312037e-7d81-6bd4-e053-0100007fdf3b/content
- umd — https://www.cs.umd.edu/~basili/publications/proceedings/P125.pdf
- AlgoCademy + 2 — https://algocademy.com/blog/why-your-microservices-might-just-be-a-distributed-monolith/
- MuleSoft Blog — https://blogs.mulesoft.com/dev-guides/microservices/is-this-the-end-of-microservices/
- Rework + 3 — https://resources.rework.com/libraries/project-management/spotify-model
- Thoughtworks — https://origin.thoughtworks.com/radar/techniques/inverse-conway-maneuver
- Libsyn — https://thoughtworks.libsyn.com/reckoning-with-the-force-of-conways-law
- Umbrex + 2 — https://umbrex.com/resources/frameworks/organization-frameworks/team-topologies/
- Ruzora — https://www.ruzora.com/blog/brooks-law-why-adding-engineers-can-slow-you-down
- Streamaligned — https://www.streamaligned.com/articles/why-team-dependencies-are-killing-your-engineering-velocity
- goodreads — https://www.goodreads.com/notes/59012409-the-devops-handbook/1004303-keith/6bad518c-6dc6-4343-8ccf-771c5320286a
- Danlebrero — https://danlebrero.com/2021/01/20/team-topologies-summary/
Commissioned from our research desk. Subject to final editorial discretion.