Big Ball of Mud 也是一个组织问题
Code 只是最小的一部分问题
Section titled “Code 只是最小的一部分问题”让我真正开始从组织视角看 Big Ball of Mud 的,不是某篇 Architecture Paper,而是一位 Principal Architect 的一句话。
他看过系统之后,大意是:
Code 其实是最小的问题。
这句话当时听起来很奇怪。毕竟 Repo 里是真实存在的几千个文件、Cycle、Global State、Shared Service、被绕过的 Boundary 和很难预测的 Side Effect。
但他的意思不是 Code 不严重。
而是:Code 至少还能被分析。
可以画 Dependency Graph,可以写 Test,可以切 Module,可以提取 Slice,可以逐步消除 Global State。真正困难的是,什么样的 Organisation 能够在几年里持续支持这种改变,而且在短期利益与长期结构发生冲突时仍然保护新边界。
这就是本篇的核心:
Big Ball of Mud 最终在 Repository 中可见,但塑造它的力量远不止 Repository。
Big Ball of Mud 在 Repository 里——但原因不只在那里
Section titled “Big Ball of Mud 在 Repository 里——但原因不只在那里”Software Architecture 会记录很多组织决定,即使这些决定从未被写成 Architecture Decision。
谁和谁需要协调,哪个 Team 可以独立 Release,哪个 API 必须经过 Central Team,哪些 Requirement 长期不清楚,哪些 Deadline 一直压倒结构工作,哪些共享区域没有真正 Owner——这些都会留下技术痕迹。
如果一个 Feature 每次都需要三个 Team 同步,Code 往往也会积累对应的 Shared Dependency。
如果 Organisation 不断切换 Ownership,系统里容易留下“大家都能改、没人真正负责”的 Shared Area。
如果一项业务能力在组织中横跨多个 Department,它也可能在 Code 中长期横跨多个 Module。
这不是机械的 Conway’s Law。Organisation 并不会一比一复制成文件夹结构。
但长期看,沟通路径、决策权和技术 Dependency 会相互影响。
Architecture 不只表达技术结构,也会沉积 Organisation 如何工作。
局部理性,全局昂贵
Section titled “局部理性,全局昂贵”Big Ball of Mud 很少需要“糟糕决策”才能形成。
很多时候,真正危险的是大量局部合理的决策。
假设 Team A 在 Deadline 前需要一个数据。最干净的方案也许需要新 Contract、Domain Boundary 和 Team B 配合。但直接 Import Team B 的 Internal Service 两天就能完成。
对于 Team A 当前 Sprint,这可能是合理选择。
Immediate Benefit 很清楚:客户按时得到 Feature。
长期成本却被推迟并分散:未来 Coupling 增加,Team B 修改内部实现必须考虑 Team A,Regression Radius 变大,更多 Change 需要 Coordination。

如果做决定的人立即获得收益,而未来成本由其他 Team、后续项目或下一任 Owner 支付,那么 Local Optimization 会自然获胜。
这不是道德问题,而是 Incentive Structure。
一个局部最优选择,完全可能在系统层面不断累积成全局昂贵的结构。
三种组织逻辑,同一个技术终点
Section titled “三种组织逻辑,同一个技术终点”同样的技术症状可以在非常不同的 Organisation 中形成。
这很重要,因为它说明 Big Ball of Mud 不能简单归因于“大企业 bureaucracy”或“Startup 太快”。
企业集团里的 Big Ball of Mud
Section titled “企业集团里的 Big Ball of Mud”大型 Organisation 通常重视标准化、治理与复用。
于是会出现 Central Platform、Shared Service、Canonical Model、Architecture Committee、统一 Identity、统一 Release Process。
这些东西本身完全可以有价值。
问题出现在中央化能力逐渐变成中央依赖。每个 Team 都必须使用同一个 Shared Model,每个 Feature 都要经过同一 Integration Layer,每个 Release 都等待中央 Approval。
技术上看,Organisation 获得了“复用”。
系统层面却可能失去自治。
当 Central Service 需要满足十个 Domain,Model 越来越抽象、越来越包含 Sonderfall。最后没人敢改,因为每个 Consumer 都可能依赖某个历史行为。
Startup 里的 Big Ball of Mud
Section titled “Startup 里的 Big Ball of Mud”Startup 的动力几乎相反。
速度本身可能是生存条件。
Team 小,Role 重叠,Product-Market Fit 还没找到,今天最重要的是验证需求而不是设计十年 Boundary。
这完全合理。
一个五人团队把 Frontend、Backend、Data 和 Deployment 快速放在一起,可能正是正确选择。
风险出现在 Organisation 已经从 Experiment 进入长期产品,却继续使用 PoC 的节奏和结构。
Temporary Path 不再 Temporary,Feature 继续叠加,过去“以后再抽”的逻辑变成核心 Flow。
不是 Startup 模式本身错了。
而是成功之后,系统身份改变了,Architecture Investment 却没有跟着改变。
外部供应商场景里的 Big Ball of Mud
Section titled “外部供应商场景里的 Big Ball of Mud”外部 Auftragnehmer 又有另一套逻辑。
他们通常围绕明确 Scope、Budget、Milestone 和 Abnahme 工作。
如果合同只支付 Feature Delivery,而没有显式 Budget 用于结构 Reconstruction,那么 Supplier 很难长期免费承担所有历史 Architecture Debt。
项目结束后可能换 Supplier。新 Supplier 接手 Code,却不自动继承所有 Historical Context。
下一次 Vertrag 又有新的 Scope 和 Deadline。
于是 Organisation 把责任切成多个项目,但 Software System 本身继续保持连续。

三种 Organisation 的目标完全不同,却可能得到非常相似的 Code:Shared Responsibility、历史 Dependency、缺少长期 Owner、局部 Optimization 和不断增长的 Coordination。
Software 的记忆比 Organisation 长
Section titled “Software 的记忆比 Organisation 长”Team 会换人。
Leadership 会变化。
Department 会 Reorganisation。
Supplier 会更换。
Code 继续存在。

这使 Software 成为一种异常持久的组织记忆。
问题在于,它通常保存的是 Decision Result,而不是完整 Reasoning。
十年前某个 Team 为什么加这个 Flag,也许当时非常合理。五年后那个 Team 不存在了,Requirement 已变化,Flag 却仍在控制核心 Flow。
新 Team 看见的是技术事实:
这里必须判断这个 Flag。
却未必知道:
这个判断最初只是某个客户 Migration 的 Temporary Workaround。
Organisation 会忘记。
Software 会继续执行。
Ownership 不只是 Organigram 里的名字
Section titled “Ownership 不只是 Organigram 里的名字”很多 Organisation 很快就能回答:
Billing 属于 Team A。
但这只是 Formal Ownership。
真正的 Ownership 还需要能力与 Decision Right。
Team A 是否能独立理解 Billing?能修改?能测试?能 Release?能运行?能处理 Incident?能改变 Contract?
如果每一次 Change 都必须等 Team B 的 Shared Library、Team C 的 Deployment、Central Architecture Committee 的 Approval 和某个 Expert 的 Review,那么 Team A 的 Ownership 很有限。
Ownership 不是一个 Label,而是对 Change 的端到端责任与行动能力。
Formal Ownership 与实际 Dependency 不一致时,会产生一个危险现象:Organisation 认为责任已经分配,Engineering 却每天在隐式协商真正的责任边界。
当业务不清晰最终进入 Code
Section titled “当业务不清晰最终进入 Code”不是所有 Architecture Problem 都从技术误判开始。
很多问题的起点是 Domain 本身没有被澄清。
如果 Stakeholder 对“Customer”“Case”“Status”“Approval”有不同定义,Frontend、Backend、Database、Report 和 Test 很可能各自建立自己的解释。
Technical Team 无法用 Layer、Pattern 或 Microservice 自动解决尚未形成共识的业务语义。
于是 Uncertainty 会通过 API、DTO、Enum 和 Condition 进入 Code。
一个 if 可能不是 Developer 不懂 Architecture,而是 Organisation 从未决定两个业务场景是否真的相同。
这也是 Ubiquitous Language 和 Requirements Engineering 为什么属于 Architecture Work 的原因。
业务边界没有被组织澄清时,技术边界很难长期稳定。
Survival Mode 会缩短时间视野
Section titled “Survival Mode 会缩短时间视野”持续 Incident、Deadline、Escalation 和 Customer Pressure 会改变 Decision Horizon。
当 Team 每周都在救火,问题自然从:
这个 Domain 两年后应该怎样组织?
缩短成:
这周五之前怎样安全上线?
后一个问题完全合理。
但如果 Organisation 长期停留在 Survival Mode,短期方案就不再是 Exception。
Architecture Work、Consolidation、Knowledge Distribution、Boundary Improvement 总是输给眼前 Delivery。
Technical Debt Research 中关于 Time Pressure、Budget、Stakeholder Influence 的工作也反复显示,Debt Management 不只是 Developer Discipline,而与 Business Constraint 和 Decision Process 相关。
当紧急状态成为常态,短期优化就会成为默认架构方法。
Architecture 需要一个值得投资的未来
Section titled “Architecture 需要一个值得投资的未来”结构投资需要一个时间窗口。
如果 Product 预计半年后关闭,为它做两年期 Reconstruction 很可能不合理。
但另一种情况同样危险:Organisation 连续五年都告诉 Team,“这个系统可能明年重写”。
如果每个人都相信当前系统只是 Temporary,就没人愿意为它建立长期 Boundary。
结果是所谓 Temporary System 反而活十年。
Architecture Improvement 因此需要某种可信 Future:
- Product 还会存在多久?
- 哪些领域会继续变化?
- Organisation 是否愿意在当前系统上投资?
- Replacement 是实际计划还是多年口号?
没有对系统未来的可信承诺,就很难要求 Team 为未来结构投资。
当 Responsibility 比 System 变化得更快
Section titled “当 Responsibility 比 System 变化得更快”频繁 Reorganisation 会导致另一个问题:Responsibility 的半衰期比 Software 短。
一个新 Team 接手 Subsystem,先用几个月理解历史,再开始建立新规则。刚刚有一点 Momentum,下一次 Org Change 又重新切分。
新 Manager 有新的 Priority,新 Owner 重新解释 Domain,Architecture Initiative 重新开始。
技术系统却继续携带上一轮 Decision。
长期 Reconstruction 需要 Continuity。
如果 Direction 每六个月重置,Architecture 很难通过多年积累得到改善。
这不意味着 Reorganisation 本身错误。
Organisation 必须变化。
但如果 Software 的结构生命周期远长于 Org Chart,Organisation 就需要某种跨 Reorganisation 的 Architecture Memory 和 Ownership。
“这不应该由 Geschäftsführer 告诉我吗?”
Section titled ““这不应该由 Geschäftsführer 告诉我吗?””在严重 Legacy System 中,有时会出现一种期待:
Management 应该知道 Architecture 已经出问题。
但 CEO、Director、Product Executive 通常不会看到 Dependency Graph、State Ownership 或 Regression Radius。
他们看到的是 Budget、Delivery、Incident、Revenue、Headcount、Deadline。
如果 System 仍然运行,而 Engineering 只用“Code 很乱”“Coupling 太高”表达风险,Stakeholder 很难独立判断严重程度。
因此 Engineering Leadership 的职责之一,就是翻译。
把结构问题翻译成:Lead Time、Predictability、Release Risk、Opportunity Cost、Key-Person Dependency、Security Response、Onboarding Cost。
不能要求 Stakeholder 诊断一个他们没有观测工具的系统属性。
Management 不主动发现并不自动意味着“不在乎”。
有时技术组织从未把问题表达成可以做 Business Decision 的形式。
技术解耦也可能触碰组织控制
Section titled “技术解耦也可能触碰组织控制”Architecture Transformation 会改变 Decision Right。
如果一个 Team 从 Shared Monolith 中真正获得自治,它可能同时获得自己的 Release、Data Ownership、API Decision 和 Operational Responsibility。
这会改变 Organisation 的 Power Distribution。
Central Platform Team 可能失去某些控制,Architecture Board 的 Approval Scope 可能缩小,Product Team 获得更多 Budget Responsibility。
因此,“技术上拆开”并不永远是 Engineering-only Decision。
它可能影响:
- Release Approval;
- Budget Responsibility;
- Platform Control;
- Architecture Authority;
- Product Ownership。
如果 Transformation Plan 假装这些 Social Consequence 不存在,Resistance 就会看起来像“大家不理解好架构”。
实际上,有些 Boundary Change 正在重新分配真实组织资源。
当过去突然也被带上讨论桌
Section titled “当过去突然也被带上讨论桌”Architecture Critique 很容易被听成 Historical Blame。
当一个 Architect 说:
这个系统结构已经严重 Erode。
某些参与者可能听到:
你们以前做错了。
于是技术讨论变成 Identity、Status 和 Responsibility Discussion。
过去 Decision 的参与者开始防御,新的 Team 又容易过度简单地评价旧系统。
更有效的方式通常是把 Focus 放在当前 System Behavior 与未来 Risk:
- 哪些 Boundary 今天不再有效?
- 哪些 Change Cost 已经出现?
- 哪些 Future Requirement 会受影响?
而不是试图重建一张十年前的 Schuld Map。
No Blaming 不等于不能评价 Decision。
只是评价应该帮助未来,而不是制造历史审判。
Big Ball of Mud 是一个 Feedback System
Section titled “Big Ball of Mud 是一个 Feedback System”Big Ball of Mud 不只是静态结构。
它可以形成自我强化循环。
Economic Pressure 让 Local Optimization 更有吸引力。
Local Optimization 引入更多 Structural Shortcut。
Shortcut 增加 Coupling。
Coupling 提高 Future Change Cost 和 Uncertainty。
更高 Change Cost 又制造新的 Economic Pressure。
于是下一次 Deadline 下,Team 更有理由再次选择局部方案。

结构侵蚀可能自我强化:Change 越贵,短期压力越大;短期压力越大,局部 Shortcut 又越有吸引力。
这解释了为什么一句“Team 应该更有 Discipline”通常不够。
如果 Incentive、Planning Horizon、Ownership 和 Coordination Structure 不变,新的 Team 可能在同样条件下重复非常相似的局部决策。
Organisation 不需要主动“制造坏架构”,也可以通过重复相同条件不断复制它。
研究说明了什么——又没有说明什么
Section titled “研究说明了什么——又没有说明什么”研究并没有一个模型,可以根据某种 Organisation Structure 精确预测“这里一定会产生 Big Ball of Mud”。
Software Development 太依赖 Context。
Conway 描述 Communication Structure 与 System Design 的关系;Mirroring Hypothesis 研究 Product Architecture 与 Organisational Architecture 的对应;Socio-Technical Congruence 关注 Technical Dependency 产生的 Coordination Need;Technical Debt Research 讨论 Time Pressure、Budget、Stakeholder、Management 与 Decision Process;Turnover Research 研究 Knowledge Loss;Requirements 与 Product Owner Research 又说明 Domain Clarification 强烈依赖 Communication 和 Organisational Embedding。
这些研究都没有说:Management 会制造坏架构。
同样也不支持另一极端:Developer 只是 Organisation 的受害者,没有 Agency。
Technical Decision 仍然是 Technical Decision,Developer、Architect 和 Team 都拥有 Responsibility。
但经验研究提示我们必须认真对待 Decision Context。
Organisation 决定 Coordination 有多贵,分配 Decision Right,定义 Budget 和 Planning Horizon,决定 Stakeholder,保存或丢失 Knowledge。
它不能决定每一行 Code。
但它改变了那一行 Code 被写出来时的选择空间。
为什么 Code 仍然可能是“最小问题”
Section titled “为什么 Code 仍然可能是“最小问题””现在可以重新理解 Principal Architect 的那句话。
Code 绝不是“小问题”。Thousands of Files、Cycle、Global State、缺失 Boundary 和不可预测 Side Effect 都是真实技术问题。
但这些问题至少有技术操作可以开始处理:Analysis、Dependency Visualization、Test、Module Extraction、State Decoupling、Boundary Reconstruction。
即使技术状态很糟,也存在 Engineering Entry Point。
真正可持续的变化却还需要其他前提:
- 足够长期的 Direction Continuity;
- Domain Clarity,用来定义新的技术边界;
- 当短期利益与长期结构冲突时,有真实 Decision Capability;
- 当新 Boundary 触碰旧 Ownership 时,有 Organisational Legitimacy;
- 一个足够长的 Time Horizon,让结构投资可能回本。
因此,一份技术上完全正确的 Reconstruction Strategy 仍然可能失败,如果 Organisation 无法长期承载它。
这并不会让技术问题消失。
它只是把诊断扩展到更完整的 Socio-Technical System。
- Besker, Terese; Martini, Antonio; Bosch, Jan (2022): The use of incentives to promote technical debt management. Information and Software Technology, 142, 106740. DOI: 10.1016/j.infsof.2021.106740.
- Cataldo, Marcelo; Herbsleb, James D.; Carley, Kathleen M. (2008): Socio-technical congruence: A framework for assessing the impact of technical and work dependencies on software development productivity. ESEM 2008. DOI: 10.1145/1414004.1414008.
- Conway, Melvin E. (1968): How Do Committees Invent? Datamation.
- Eilam, Galit; Shamir, Boas (2005): Organizational Change and Self-Concept Threats: A Theoretical Perspective and a Case Study. The Journal of Applied Behavioral Science, 41(4), 399–421. DOI: 10.1177/0021886305280865.
- Klotins, Eriks; Unterkalmsteiner, Michael; Gorschek, Tony (2019): Software engineering in start-up companies: An analysis of 88 experience reports. Empirical Software Engineering, 24(1), 68–102.
- MacCormack, Alan; Baldwin, Carliss Y.; Rusnak, John (2012): Exploring the duality between product and organizational architectures: A test of the “mirroring” hypothesis. Research Policy, 41(8), 1309–1324. DOI: 10.1016/j.respol.2012.04.011.
- Nielsen, Mille Edith; Madsen, Christian Østergaard (2022): Stakeholder influence on technical debt management in the public sector: An embedded case study. Government Information Quarterly, 39(3), 101706. DOI: 10.1016/j.giq.2022.101706.
- Ramač, Robert et al. (2022): Prevalence, common causes and effects of technical debt: Results from a family of surveys with the IT industry. Journal of Systems and Software, 184, 111114. DOI: 10.1016/j.jss.2021.111114.
- Rigby, Peter C.; Zhu, Yue Cai; Donadelli, Samuel M.; Mockus, Audris (2016): Quantifying and mitigating turnover-induced knowledge loss: Case studies of Chrome and a project at Avaya. Proceedings of ICSE 2016, 1006–1016. DOI: 10.1145/2884781.2884851.
- Robillard, Martin P. (2021): Turnover-induced knowledge loss in practice. Proceedings of ESEC/FSE 2021, 1292–1302. DOI: 10.1145/3468264.3473923.
- Wiese, Marion; Borowa, Klara (2023): IT managers’ perspective on Technical Debt Management. Journal of Systems and Software, 202, 111700. DOI: 10.1016/j.jss.2023.111700.
当 Principal Architect 用“Code 是最小的问题”总结系统状态时,他并不是说 Architecture 无害。
它已经严重侵蚀。
只是 Code 拥有一个很多其他问题没有的特性:它可以被分析,也可以通过技术手段改变。
更困难的是,什么样的 Organisation 能在足够长时间里维持这种变化。
Dependency 和 Global State 当然由 Developer 写出来,Technical Responsibility 仍然属于做技术决策的人。
但形成这些 Decision 的条件远远超过 Code:Requirement、Economic Pressure、Contract、Team Structure、Ownership、Knowledge Loss、Leadership、Planning Horizon、Decision Right。
Big Ball of Mud 因此既不是纯技术问题,也不是纯组织问题。
这些层面相互作用。
Local Decision 在 Code 中变成长期事实,而 Organisation 继续变化。新 Team 继承前人的技术决定,又在新的 Constraint 下继续做新的局部决定。几年后,同一个 Codebase 中可能沉积了多代 Requirement、Budget、Project、Leadership、Team Boundary 和 Contract。
Big Ball of Mud 在 Repository 中可见,但它的原因并不只存在于 Repository。
如果技术、业务、组织和经济力量共同参与了这种状态的形成,那么可持续改变也需要超过纯技术视角。
Architecture 活在 Code 中,但塑造它的力量远远超出 Code。