2026 Agentic Software Engineering: A Terminology and Literature Placeholder
Table of Contents
1. Overview
This is a placeholder. The terminology below is unsettled as of September 2026: different papers and vendors name overlapping pieces of the same problem differently, and no umbrella term has stabilized yet. This note tracks the landscape rather than argues a position, and it will need revisiting as usage converges.
The problem itself is concrete: many AI-agent-authored pull requests compete to land on one trunk. The umbrella term that appears most often for the broader practice is agentic software engineering, sometimes called "SE 3.0." The specific problem — many agent-authored changes competing for one trunk — is usually called integration of agent-authored pull requests, or the integration bottleneck.
2. Names by sub-problem
No single name covers all four parts. Each sub-problem already has an established name from pre-agent software engineering research, and a newer, narrower name that 2026 papers are using for the agent-specific case.
| Sub-problem | Established name | 2026 agent-specific name |
|---|---|---|
| Ordering and batching merges | Merge queues; "keeping master green"; speculative CI | Agent-scale or speculative merge queues |
| Concurrent changes that conflict | Merge conflict prediction; semi-structured merge | Cross-agent conflicts; merge contention |
| Predicting that a change will break production | Just-in-time (JIT) defect prediction; change risk assessment | Merge-outcome predictors for agent PRs |
| Deciding whether a change may deploy | Release engineering; change management (ITIL change enablement) | Agent governance; deploy gating |
| Coordinating agents before they write code | Task decomposition; concurrency control | Pre-write admission; multi-agent code co-synthesis |
3. 2026 research
3.1. Conflict measurement
AgenticFlict (Ogenrwot and Businge 2026) (Ogenrwot and Businge, arXiv:2604.03551, AIware 2026) simulates merging each agent PR against its base branch. Across more than 142,000 agent PRs in more than 59,000 repositories, 27.67% conflicted. Conflict rate increases with PR size (lines added plus deleted), and rates differ by agent.
AI Agent Pull Requests on GitHub: Frequency, Structure, and Merge Conflict Rates (Xu, Subramanian, and Karthik 2026) (arXiv:2607.04697) uses the AIDev-pop dataset. In 40.2% of repositories, some pairs of agent PRs were open at the same time, and those pairs account for 79.4% of all agent PRs. The authors list lost CI compute and token spend on conflict resolution as hypothesized costs; they did not measure them.
3.2. Merge-outcome prediction
When AI Teammates Meet Code Review (Nachuma and Zibran 2026) (Nachuma and Zibran, MSR 2026, arXiv:2602.19441) fits a logistic regression to 33,596 agent-authored PRs. Substantive reviewer engagement is the strongest positive predictor of merge. Repositories reviewed only by automated agents merge 45% of agent PRs; repositories with human review merge 68%.
3.3. Coordination before writing
ATM: CID-Brokered Pre-Write Admission for Multi-Agent Code Co-Synthesis (Huang 2026) (arXiv:2607.00041) admits agent writes to shared code before they happen. It applies concurrency-control ideas (read-set tracking, isolation levels) to agents writing code, rather than resolving conflicts at merge time.
3.4. Trust boundaries
Knowledge-Based Pull Requests (Zhang and Sun 2026) (arXiv:2606.26721) proposes that an external contributor's code and agent trace serve as input knowledge. A project-owned agent then regenerates the code inside the receiving project. This separates two decisions: whether the knowledge should enter the project, and whether a specific implementation should merge.
3.5. Merge queues at agent scale
Tian Pan's July 2026 essay applies queueing theory to a serial merge queue. With 30-minute CI, the queue lands at most two PRs per hour. Agents raise the arrival rate while the service rate stays fixed.
The same essay notes that batch failures force bisection, which costs O(log n) CI runs. Flaky tests set how often this happens, so agents that re-queue failed PRs increase the cost.
Mergify's speculative checks test the cumulative merges (PR1), (PR1+PR2), (PR1+PR2+PR3) in parallel lanes.
4. Pre-agent foundations
The agent-specific work above builds on an older, non-agent-specific literature:
- Ananthanarayanan et al., "Keeping Master Green at Scale" (Ananthanarayanan et al. 2019) (EuroSys 2019), describes Uber's SubmitQueue. SubmitQueue predicts build outcomes to schedule speculative builds; it is the main academic treatment of merge scheduling.
- Kamei et al., "A Large-Scale Empirical Study of Just-in-Time Quality Assurance" (Kamei et al. 2013) (TSE 2013), is the standard reference for JIT defect prediction: assigning a risk score to each commit.
- Google's TAP publications cover presubmit/postsubmit testing and culprit finding at monorepo scale.
- Ghiotto et al., "On the Nature of Merge Conflicts" (Ghiotto et al. 2020) (TSE 2020), and Apel et al.'s semistructured merge (Apel et al. 2011) (FSE 2011) cover conflict characterization and resolution.
- The DORA metrics (change failure rate, lead time) and ITIL 4 change enablement (standard, normal, and emergency changes; the change schedule) are the operations-side terms for risk-gated integration.
5. Search terms
A practical list for finding more of this literature: "agent-authored pull requests", "AIDev dataset", "agentic PR integration", "merge queue speculative", "keeping master green", "just-in-time defect prediction", "change risk assessment", "multi-agent code co-synthesis", and "SE 3.0". MSR 2026 and AIware 2026 are the venues where the agent-specific work is concentrated.
6. Closing synthesis
Two other projects were mentioned in passing as covering pieces of this problem: a "slipway" deploy-gating pipeline for the change-risk and ITIL-change-enablement part, and a "wake" post-merge survival analysis for the outcome-measurement part. Neither name matches anything findable in this repository's own site content or history (checked via search and a full-tree grep) — they read as external projects, not something documented here, so they are reported as context rather than linked. On the same framing: the merge-scheduling and pre-write coordination parts of this problem are the ones the ATM (Huang 2026) and SubmitQueue (Ananthanarayanan et al. 2019) work addresses above.