跳转到内容

AI 与 Big Ball of Mud

Big Ball of Mud 很难理解。User Flow 分散在 Component、Service 和 Global State 中,Side Effect 出现在可见结构无法预测的位置,Inheritance 把业务上几乎无关的区域绑在一起。一个看似局部的 Change,可能沿 Dependency 穿过多个 Domain。

对人来说,这种分析最终会变得非常昂贵。问题不一定是每个局部都特别复杂,而是一次修改需要同时考虑的潜在相关关系越来越多。

然后 Coding Agent 出现了。

它可以搜索整个 Repository,追踪 Reference,打开 Service、Base Class 和 Test,重建 Call Chain,寻找 State Writer;如果一条路径又指向新的 Dependency,它就继续打开下一份文件。

人类 Developer 在几个小时之后会开始问:“我为什么会走到这个文件里?”

Agent 只是继续搜索。

这绝不是小改进。现代 Coding Agent 已经能够承担大量过去对人极其耗费 Attention 的 Repository Exploration。Repository-level Benchmark 和 2025/2026 年研究也早已不再只关心单函数 Generation,而是在研究 Navigation、Search、Multi-file Reasoning、Tool Use、Memory 与 Context Selection。

时间边界:2026 年 8 月。 本文讨论的是当前 Coding Agent、Context Mechanism 和经济模型。没有人能够严谨预测五年或十年后 Agent 能做到什么。也许未来真的可以只说一句“扔掉重做”。截至 2026 年 8 月,这还不是可靠 Architecture Strategy。

因此,有意思的问题不是 AI 能不能帮助 Big Ball of Mud。

它今天已经能。

真正有意思的是:AI 到底改善了什么?

Coding Agent 拥有一些非常适合复杂 Legacy Codebase 的能力。

它可以 Repository-wide Search,可以一次检查大量文件,可以追 Symbol 与 Reference,可以执行 Test、解释 Log,也可以在第一条推理路径失败之后换另一条路径继续探索。

尤其重要的是,它不需要像人一样长期把所有分析结果同时放在 Working Memory 里。Agent 可以重新加载 Context,重复 Search,压缩 Intermediate Result,再把之前走过的 Path 重建一次。

这并不意味着它“理解了整个系统”。这个说法太强。

但它确实能够以机器化方式处理大量显式 Code Context,并比人更便宜地遍历大 Search Space。

Repository Exploration 本身已经可以被测量。2026 年 FastContext 的 Preprint 分析了 SWE-bench Multilingual 上 300 条 Coding-Agent Trajectory:Reading 与 Searching 合计占主 Agent Tool-use Turn 的 56.2%,占主 Agent Token 的 46.5%。把 Exploration 外包并压缩后,实验中的主模型 Token Consumption 明显下降,在部分 Benchmark 上最高接近 60%。这些值不能被当作普遍定律,但它们清楚说明一件事:找到相关 Code,本身就是 Agent Workload 的重要部分。

ContextSniper 的 Pilot Experiment 同样显示 Repository-level Repair 可以显著降低 Token Use,代价是部分 Submitted-Resolution Rate 略有下降。SWE Context Bench 则说明,正确选择和压缩 Experience 可以降低 Token Cost 与 Runtime,而无差别加入 Context 并不会稳定帮助,甚至可能伤害效果。

这些研究仍然很年轻,很多在 2026 年 8 月还是 Preprint。

但基本问题已经不是 Hypothesis:

Agent 在改变 Code 之前,必须先找到哪些 Code 真正相关。

而 Big Ball of Mud 正是让这个问题变难的结构。

假设 Agent 为一个 Ticket 重建出这样一条关系:

A 调用 B

B 修改 C

CD 中触发 Side Effect。

共同 Base Class 又覆盖部分行为,因此额外执行 E

Agent 找到了这条 Chain,甚至可以解释为什么它与 Ticket 有关,补一个 Test,然后在正确位置实施 Change。

这很有价值。

但任务完成之后,系统仍然是:

A → B → C → D → E

业务 Boundary 并没有因此更清晰,Responsibility 没有重新切分,Dependency 不会突然变成单向,Global State 仍然 Global,跨 Domain 的 Base Class 也继续耦合多个区域。

AI 可以用分析补偿缺失的 Architecture,但分析能力本身不是 Architecture。

能够重建一条 Dependency Path,与让这条 Path 有合理边界,是两件不同的事。

这回到了本系列一直出现的机制:Big Ball of Mud 可以很长时间继续工作,因为 Organisation 不断补偿结构缺陷。Specialist 记住危险区域,Review 捕获 Risk,Regression Test 保护 Critical Path,Release Process 变得更谨慎,某些人知道哪些“看起来没问题”的 Change 实际上绝对不能做。

Coding Agent 可能成为又一层、而且非常强的 Compensation。

系统不一定因此更简单。

我们只是变得更擅长在复杂系统里继续工作。

老系统还有第二个限制:一次 Change 所需的所有 Context,不一定存在于当前 Code 中。

为什么有这个 Sonderfall?它来自 2019 年 Production Incident 吗?当时的技术限制今天已经消失了吗?这个奇怪的额外 Request 是不是为了一个早已下线的 Backend?某个 if 是真实业务规则,还是一次历史 Exception?这个 Dependency 是设计决定,还是 Deadline 下留下的 Accident?

Agent 可以分析当前 Code,也可以分析 Git History、Issue、ADR、Documentation、Log、旧 Pull Request——只要这些东西仍然存在并且可以访问。

这恰恰也是 Agent 的优势。它可以比人更经济地搜索大量历史 Artefact,把分散信息联系起来。

但限制非常明确:

信息必须还存在。

如果 Decision 从未记录,Ticket 已经删除,当年 Developer 已离职,只剩由 Decision 产生的 Code,那么 Agent 最多形成 Hypothesis。

它可以识别 Pattern,提出 Plausible Explanation。

却不能从零可靠恢复丢失的历史。

再大的 Context Window 也无法恢复根本没有被保存的历史。

这不是 AI 特有问题。新加入的人类 Developer 同样受到限制。

Agent 改变的是:我们能够以可接受成本搜索多少仍然存在的 Material。

这正是 Agentic Coding 中 Architecture 更有意思的地方——Architecture 不只是 Code Structure,也是 Knowledge Requirement 的结构。

传统 Architecture 会用 Maintainability、Separation of Concerns、Information Hiding、Clear Responsibility 来解释好结构。

对 Coding Agent,可以再增加一个视角:

Architecture 限制一次 Change 的相关 Search Space。

如果修改发生在 planning Capability 中,而系统有真正清晰的 Slice,那么 Agent 一开始就有强 Prior:这里的 Public Contract、State、Use Case、Test 是主要相关区域。对外 Dependency 经过窄 Interface,内部 Implementation 不需要在每个 Change 中重新理解。

这可以被视作 Context Compression。

系统当然仍然复杂,但 Architecture 把复杂度装进 Boundary 中,让每次任务不必重新加载全部 Complexity。

Parnas 1972 年关于 Module Decomposition 与 Information Hiding 的经典工作,核心正是让一个 Module 隐藏 Design Decision,使其他 Module 不需要知道内部细节。

对 Agent 来说,这个效果非常具体:需要读取的 File 更少,需要探索的 Dependency 更少,不相关的 Code 更少进入 Context。

同一个 Planning Change:模块化系统把相关 Context 压缩到小 Slice 与少量 Contract;Big Ball of Mud 则把分析扩散到大量连接区域。

好的 Architecture 并不会“让 AI 变聪明”。

它减少 AI 在一次任务里必须同时推理的世界大小。

Big Ball of Mud 缺少的正是这种 Boundary

Section titled “Big Ball of Mud 缺少的正是这种 Boundary”

在 Big Ball of Mud 中,Agent 不能简单说:

这是 Planning Change,所以只读 Planning。

它必须先证明 Planning Boundary 是否真的存在。

它会 Search Consumer、找 State Writer、追 Event、检查 Shared Service、读 Base Class、比对重复 Model、运行 Test、检查 Runtime Wiring。

Repository Exploration 于是从准备工作变成任务主体的一部分。

Agent 做得越好,Big Ball of Mud 越可维护。

但不要混淆原因和补偿:

Agent 的强 Exploration 能力,是对缺失 Boundary 的补偿,而不是 Boundary 已经存在的证明。

清晰 Architecture 让 Context Selection 更便宜。

糟糕 Architecture 迫使 Context Selection 先成为一个搜索问题。

Agentic Coding 在好 Architecture 中也不便宜

Section titled “Agentic Coding 在好 Architecture 中也不便宜”

这里也不要制造反向 Myth。

即使系统结构很好,复杂 Feature 仍然需要大量 Agent Work:Requirement Context、Domain Rule、Integration、Test、Build、Lint、Tool Call、Review、Failure Recovery。

一个复杂 Use Case 不会因为 Vertical Slice 就突然只消耗 2k Token。

区别不在“好 Architecture 免费、坏 Architecture 昂贵”。

而在于 Cost 是否主要用于真正的 Domain Work,还是大量花在重新发现系统隐藏结构。

良好 Architecture 仍然会消耗 Context。

只是这个 Context 更有可能是任务相关 Context,而不是“为了先找到任务相关 Context 而加载的 Repository 噪声”。

Coding Agent 非常擅长完成局部目标。

Test 红了,让它变绿。

Edge Case 没处理,补一个 Condition。

Build Broken,调整 Type。

如果 Architecture Constraint 不清晰,最便宜的成功 Path 很可能就是再加一个局部 Exception。

Ticket 完成,Test 通过,Agent Score 很好。

系统的 Dependency Graph 却没有改善,只多了一个新的 Sonderfall。

Agent 成功完成 Ticket、Test 变绿,但复杂 Dependency Network 只多了一个局部 Sonderfall,结构没有改善。

这是 Agentic Coding 的一个关键风险:

任务成功不等于结构成功。

Agent 的 Optimization Target 通常是“满足 Requirement、通过 Test、完成 Ticket”。

如果“新的 Code 必须尊重 Boundary”“不允许增加 Shared State”“Component 不允许 Orchestrate Use Case”没有成为 Explicit Constraint,那么 Agent 没有理由自动选择更昂贵的 Structural Path。

Coding Agent 会放大团队给它的目标。

目标局部,它就可能极其高效地放大局部优化。

Big Ball of Mud 常常已经有多层 Compensation:Expert Knowledge、Regression Test、Review Ritual、特殊 Release Process、Documentation、Manual Checklist。

Agentic Repository Analysis 可以成为新的一层。

它替 Team 搜索隐藏 Dependency,帮助解释老 Code,生成 Test,预测 Affected Area。

Big Ball of Mud 上叠加专家知识、测试、Review、Release Ritual,最上方再增加 Agentic Analysis,而底层结构仍然不变。

这并不是坏事。

Compensation Mechanism 本来就有价值。一个能够把高风险 Change 从三天缩短到三小时的 Agent,是巨大 Productivity Gain。

有趣的是 System Effect:

如果 Agent 足够便宜地补偿 Architecture Erosion,Organisation 对根本 Reconstruction 的即时经济压力可能下降。

过去“没有 Expert 就改不了”的区域,现在 Agent 也能导航。

过去需要两天 Repository Exploration,现在可能 30 分钟。

结构没有改善,但它重新变得经济可操作。

AI 可能不是先消灭 Big Ball of Mud,而是先提高 Organisation 对 Mud 的承受能力。

Agentic Workload 的 Cost 不再只有 Human Hour。

高度耦合系统会产生重复循环:Explore、Read、Search、Run、Fail、Retry、Re-plan、再读更多文件。

模块化系统也会发生这些步骤。

区别可能在于 Scope 与次数。

同一个 Task 在清晰 Boundary 中,也许只需要一个局部 Context;在 Big Ball of Mud 中,需要多轮扩大 Search Space 才能确定“哪些东西不能漏”。

当 Tool 使用 Flat-rate Subscription 时,这种差异最初很难被看见。

一个 Developer 可能在一天内烧掉原本计划一个月使用的高端 Agent Contingent,却只看到“今天这个 Ticket 很复杂”。

一旦进入 Usage Cap、Credit、Request-based Billing 或 Enterprise Cost Allocation,Repository Structure 会逐渐表现成可以观察的 Compute Demand。

这让 Architecture Economics 多了新的计量单位。

今天的研究已经说明什么——又没有说明什么

Section titled “今天的研究已经说明什么——又没有说明什么”

截至 2026 年 8 月,Repository-level Coding Research 已经越来越关注 Context Selection、Repository Exploration、Graph Representation、Long-context Reasoning 与 Iterative Memory。

但这篇文章必须非常明确地限制 Claim:

目前还没有直接受控研究,把“架构清晰系统”与“Big Ball of Mud”在真实 Agent Token / Compute Cost 上进行系统比较。

所以,“Big Ball of Mud 会让 Token Cost 指数增长”不能被写成 Research Result。

更可靠的说法是:Boundary 失效会扩大潜在相关 Context;而最新 Repository-level Agent Research 已经表明 Context Selection 与 Exploration 对性能、Token Use 和 Runtime 很重要。

这是两个有 Evidence 的事实。

把它们组合成具体 Architecture Cost Model,仍然是 Hypothesis 和未来 Research Question。

FastContext、ContextSniper 等工作说明 Search 与 Read 并非 Agent 外围动作,而可能占很大比例的 Tool Turn 和 Token。

Agent 在生成 Patch 之前,必须先建立局部 Repository Model。

因此,Architecture 如果能让相关区域更容易定位,就有可能影响 Agent Resource Consumption。

我们不能从单个 Benchmark 推出 Production Cost 的固定百分比,但“Exploration 不是免费的”已经非常清楚。

Context Selection 比最大 Context 更重要

Section titled “Context Selection 比最大 Context 更重要”

LLM 时代一个常见直觉是:Context Window 越大,直接把 Repository 塞进去就好。

现实研究越来越不支持这种简单策略。

无关 Context 会增加 Noise,错误 Experience 可能降低效果,选得准比放得多更重要。

这与 Software Architecture 有一个直接连接点:清晰 Boundary 为 Context Selection 提供结构先验。

如果 Domain、Public API、Ownership 和 Dependency Direction 明确,Agent 不需要完全从 Token-level Evidence 重新推断“哪里相关”。

CodexGraph、CodeMEM 以及其他 Code Graph、AST-guided Memory、Repository Representation 工作尝试把结构信息提供给 Agent。

这说明 File Tree 并不是唯一有价值的 Context。

Symbol Relation、Call Graph、Dependency、AST、Memory 都可以减少 Agent 每次从头搜索的负担。

好的 Architecture 并不等同于这些 Representation。

但 Architecture 越能让实际 Dependency 与声明结构一致,这些结构化表示就越有机会成为可靠导航工具。

目前不能严谨地说:

Vertical Slice 比 Big Ball of Mud 节省 43% Token。

也不能说:

Layered Architecture 自动让 Agent Benchmark 提高 X 个百分点。

这样的数字需要在控制 Task、Repository Size、Model、Harness 等变量后进行 Architecture-level Comparison。

截至本文时间,这种 Evidence 仍然缺失。

因此本文的 Claim 刻意保持较弱:Architecture 决定 Context Boundary,而 Context Boundary 对 Agentic Work 具有可测量重要性。

账单越来越以 Token、Context 和 Compute 出现

Section titled “账单越来越以 Token、Context 和 Compute 出现”

传统 Big Ball of Mud Cost 已经包括 Human Analysis Time、Expert Dependency、Regression、Coordination、Lead Time、Onboarding。

Agentic Coding 又增加新的 Resource:Input Token、Output Token、Cached Context、Retrieval、Tool Call、Agent Run、Compute、Runtime,以及之后仍然存在的 Human Review 与 Correction Loop。

这些 Resource 并不都由每个 Provider 分开收费。有些打包在 Subscription,有些通过 Cache 便宜,有些藏在 Credit / Usage Limit 后面。

但技术上仍然被消耗。

AI 不会让 Big Ball of Mud 免费变得可理解。我们只是越来越不只用 Human Attention 支付这笔费用。

这可能给 Technical Debt 带来一个新的观察维度。

过去,“理解这个 Change 需要多少 System Context”很难度量。

Agent Harness 则越来越能记录 Token、Context Size、Cache、Tool Call、Run Duration、Exploration Trajectory。

未来也许可以额外问:

为了安全定位一个局部 Change,这个系统需要多少 Machine Context?

它不会成为 Universal Architecture Metric。困难 Feature 在好架构里仍然困难;不同 Model 与 Harness 差异巨大;Token 也不是标准化 Cognitive Complexity Unit。

但作为 Supplementary Signal,它可能很有意思。

传统坏架构成本之外,Agentic Coding 又增加 Context、Token、Tool Call、Compute 与 Agent Run 等可计量成本。

AI 不一定消除 Complexity Cost,它可能把其中一部分转移到 Machine Processing。

这类关系长期很容易被 Subscription 隐藏。

Developer 每个月付固定价格,处理很多 Ticket,月底仍然看到同一笔费用。

于是主观 Cost Function 很简单:

AI 每个月 X 欧元。

但这与某个 Task 实际消耗的 Machine Work 不是同一个量。

Subscription 通过 Included Usage、Credit、Session / Weekly Limit、Fair Use、Model Class 抽象 Resource Consumption。超过边界之后,额外使用可能被 Limit 或按 Usage 计算。

2026 年几个大 Provider 的变化清楚显示:Flat Fee 与真实 Agentic Consumption 的关系正在重新调整。

GitHub 从 2026 年 6 月 1 日开始把 regular Copilot Plan 转向 GitHub AI Credits,计算基于 Input、Output、Cached Token 和具体 Model Price。GitHub 明确说明,Long Agentic Session 带来更高 Compute / Inference Demand,旧 Premium-request Model 无法长期表达这种差异。部分 Annual Legacy Plan 暂时保留旧 Model。

需要特别小心一个数字:旧年度 Plan 对 GPT-5.5 在 2026 年 8 月显示 Premium-request Multiplier 57。

不等于“GPT-5.5 Token 价格涨了 57 倍”。

它只是一个即将退出的 Request Billing Unit 的 Multiplier。Token Price、Request Multiplier、Included Credit 与真实 Inference Cost 是不同概念。

OpenAI 也在 2026 年 4 月把 Codex Credit Rate Card 更直接绑定 Input、Cached Input 和 Output Token。Codex 仍包含在部分 ChatGPT Plan 中,但超出 Included Limit 后可以使用额外 Credit。

Anthropic 的 Claude / Claude Code 付费 Plan 同样有共享 Usage Limit,超过后按 Plan 可启用 Usage Credit 或转向 Usage-based API。Anthropic 也说明 Usage 受到 Interaction Length、Complexity、Model 和 Feature 影响。

这些事实不证明 Subscription 对 Provider 亏损,也不允许我们推断内部 Cost。

它们只说明:

Subscription Price 与某个 Task 实际需要的 Machine Work 不是同一个量。

当 Architecture Cost 消失在 Subscription 中

Section titled “当 Architecture Cost 消失在 Subscription 中”

设想两个 Team。

Team A 在一个 Modular System 中工作,Domain Boundary 清楚,Dependency 定向。

Team B 在一个高度 Entangled 的系统中工作。

两边使用同一个 Coding-Agent Product。

如果 Accounting 只看到每月同样 Subscription Fee,两边“AI Cost”看起来完全一样。

这不意味着 Resource Consumption 相同。

Team A 的 Agent Budget 可能更多用于 Implementation、Test、Review。

Team B 可能先花更大比例探索:哪些 File、Service、State、Side Effect 真正相关。

这个差异在真实 Production Repository 中究竟有多大,我们目前没有足够 Empirical Evidence。

但只要 Billing 更透明、更 Usage-based,额外 Search、Context、Agent Run 和 Compute 就可能直接变成 Cost。

于是过去藏在 Human Frustration、Expert Time 与 Slow Lead Time 中的一部分 Architecture Cost,开始以另一种 Unit 出现。

坏 Architecture 的 Cost 没有消失。Flat-rate AI 只可能暂时隐藏我们正在用什么单位支付它。

未来讨论 Technical Debt,也许除了问:

Developer 需要多久?

还会问:

我们的 Development System 需要多少 Context 和 Machine Exploration,才能把这个 Change 安全地限制到一个范围?

Big Ball of Mud 一直都很贵。

Agentic Coding 可能让部分账单从小时变成 Token 和 Compute。

这里可能出现最危险的误判。

过去某类 Legacy Change 需要 Developer 两天,现在优秀 Agent 两小时就能做。

这是非常真实的 Productivity Gain。

Management 很容易进一步推导:

那 Architecture Problem 好像已经没那么重要了。

因为可见症状确实变小。Developer 不再手工搜索两天,Agent 找到 Global State、Base Class、Side Effect,补 Test,再实现一个 Sonderfall。

Delivery 又变快了。

但没有一个 Slice 因此恢复,没有一个 Layer 因此变清楚,没有 Ownership 自动出现,Global Dependency 没减少,Historical Workaround 也没消失。

Agent 找得到路,不等于系统重新拥有了可持续结构。

甚至从 Organisation 视角看,成功补偿可能降低 Modernization Pressure。

这仍然是 Hypothesis,但它与 Big Ball of Mud 过去的稳定机制非常一致:只要 Compensation 足够便宜,继续运营旧系统就可以保持经济理性。

Coding Agent 因此可能不只是改变 Productivity。

它还可能改变“坏 Architecture 什么时候变得经济不可承受”的时间点。

当然,而且很可能帮助很大。

让 Agent 成为强 Compensation Layer 的能力,同样适合 Architecture Work。

它可以分析 Dependency Structure,从 Git History 研究 Change Coupling,重建 User Flow,识别 Code Gravity Center,生成 Characterization Test,为 Refactoring 建 Safety Net,并在 Incremental Migration 中反复检查新 Boundary 仍有哪些 Violation。

Architecture Recovery 本身也是自然 Use Case。系统已经失去可见 Architecture 时,Machine Analysis 可以帮助重新发现真正有效的结构。

Code Graph、Repository Representation、Context-aware Navigation 研究都说明这类 Structural Information 对 Agent 有价值。

区别不在 Tool,而在 Goal。

“在当前系统里实现这个 Ticket。”

主要把 Agent 当 Compensation。

“帮助我理解真实结构,并逐步恢复可靠 Boundary。”

使用的是同样能力,但目标变成 Architecture Reconstruction。

两者都可以合理。

只是完全不同的 Objective Function。

AI 真正改变了 Big Ball of Mud 什么?

Section titled “AI 真正改变了 Big Ball of Mud 什么?”

AI 给 Big Ball of Mud 带来一种很有意思的 Ambivalence。

它可以让本来对人类分析已经极其昂贵的系统继续保持可变更。它能更快搜索大量 Code、追 Reference、汇总 Constraint、重建 Dependency、补 Test、完成过去难以想象速度的局部 Change。

这是真实而重要的能力。

但结构问题并不会因此消失。

Agent 仍然要补偿缺失 Boundary:搜索、选择 Context、处理更大相关范围,消耗 Tool Call、Token、Compute 和 Runtime。

而只要下一个 Ticket 再次成功,这种 Compensation 甚至可能减少 Organisation 立即处理 Root Architecture Problem 的压力。

AI 可以让 Big Ball of Mud 更容易工作,而不让它因此更少成为 Big Ball of Mud。

未来 Agent 是否能够高度自主地重建大型 Legacy 的 Hidden Domain Model,并安全执行结构 Transformation,仍然是开放问题。

截至 2026 年 8 月,我们还没有到那里。

今天更值得注意的结论更微妙:

AI 可能首先不是 Big Ball of Mud 的终结者,而是延长它寿命的新能力。

Agent 已经能够找到很多穿过 Mud 的路。

地面不会因此自动变干。

以下研究与 Provider 信息构成本文关于 2026 年现状的主要依据:

  • Parnas, D. L. (1972): On the Criteria To Be Used in Decomposing Systems into Modules. Communications of the ACM 15(12), 1053–1058. DOI: 10.1145/361598.361623.
  • Le Hai, N.; Nguyen, D. M.; Bui, N. D. Q. (2025): On the Impacts of Contexts on Repository-Level Code Generation. Findings of NAACL 2025, 1496–1524. DOI: 10.18653/v1/2025.findings-naacl.82.
  • Liu, X. et al. (2025): CodexGraph: Bridging Large Language Models and Code Repositories via Code Graph Databases. NAACL 2025, 142–160. DOI: 10.18653/v1/2025.naacl-long.7.
  • Wang, P.; Zhang, L.; Liu, F.; Tao, C.; Zhu, Y. (2026): CodeMEM: AST-Guided Adaptive Memory for Repository-Level Iterative Code Generation. Findings of ACL 2026, 16903–16917. DOI: 10.18653/v1/2026.findings-acl.834.
  • Wang, Y. et al. (2026): RepoReasoner: Evaluating Repository-Level Code Reasoning Ability of Long-Context Language Models. FSE 2026.
  • Qiu, J. et al. (2025): LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering. arXiv:2511.13998 (Preprint).
  • Zhu, J.; Hu, M.; Wu, J. (2026): SWE Context Bench: A Benchmark for Context Learning in Coding. arXiv:2602.08316 (Preprint).
  • Zhang, S. et al. (2026): FastContext: Training Efficient Repository Explorer for Coding Agents. arXiv:2606.14066 (Preprint).
  • Luk, C. et al. (2026): ContextSniper: AntTrail’s Token-Efficient Code Memory for Repository-Level Program Repair. arXiv:2607.01916 (Preprint).
  • Ma, D. et al. (2026): LLM Agents Can See Code Repositories. arXiv:2606.14061 (Preprint).
  • Salim, M.; Latendresse, J.; Khatoonabadi, S.; Shihab, E. (2026): Tokenomics: Quantifying Where Tokens Are Used in Agentic Software Engineering. arXiv:2601.14470 (Preprint).
  • GitHub (2026): GitHub Copilot is moving to usage-based billing 以及 Model multipliers for annual plans on request-based billing (legacy).
  • OpenAI(截至 2026 年 8 月):Codex rate card 以及额外 Codex Usage Credit 的文档。
  • Anthropic(截至 2026 年 8 月):Claude Code usage and limitsUsage Credits for paid Claude plans.