← → navigate  ·  G grid  ·  F fullscreen
The Iteration Manager in the Age of Agents
Hybrid Delivery 2026
26
Delivery Leadership in Human+AI Teams

The Iteration
Manager in the
Age of Agents

Own what agents can generate but can't decide.
Presenter
Adnan Ali
Version
v3.0.0
aha agile
The Anomaly in the Squad

Your developers have never shipped more. Your squad's delivery hasn't moved.

441%
PR review time — one triangulating signal, not the whole story; the real story is that output and delivered value came apart. Velocity didn't break today — agents just made gaming it free.
Faros AI, "AI Engineering Report 2026" (22,000 devs, 4,000+ teams) — telemetry, not the DORA survey
02
aha agile
Locate Yourself
Where the squad is
  • Using AI agents in daily work
  • Several ways of working with AI, all informal
  • Coordination quietly getting harder
  • Extra review work nobody is tracking
Where you are
  • Still mostly using chatbots
  • No agreed way to govern agent work
  • No deliberate choice of how humans stay involved
  • Same ceremonies as before
That gap is the opportunity. You close it from inside the team — not by waiting for a mandate.
03
The Loop

An agent doesn't stop when you stop typing. It loops until it hits a gate you designed.

A year ago "agent" meant a chatbot you prompt; now it means something that acts on its own — and that gate, the human-in-the-loop decision point, is where the IM governs.

Agentic loop: Receive → Think → Act → Observe → Decide → (complete or loop back to Think)
04
aha agile
The Shape of the Work

Humans frame it. Agents build the middle. Humans judge what's good.

On familiar work the construction step is now near-free, so the binding constraint moved to the ends — the framing at the front and the judgment at the back.

Frame
Humans
The what and the why — planning, intent, the problem worth solving.
Collapse
Agents
Construction — the middle of the work, compressed to near-zero cost.
Judge
Humans
What's good enough to ship — the call the binding constraint moved to.
METR, "Measuring the Impact of Early-2025 AI on Experienced OS Developer Productivity" (July 2025 RCT)  ·  shape: synthesis of SDLC-phase / human-sandwich / 80%-problem framings
05
aha agile
From the Trenches
The first boundary I owned
was one nobody was watching.
The fix wasn't more proofreading — it was naming who signs off before an agent ships.
06
aha agile
The Pivot
What happens when a delivery team runs with AI — not beside it?
07
aha agile
The Broken Metric

Output and value came apart. Velocity now measures the wrong one.

Individual output — up
  • PRs/dev +98% (2025) → +16.2% (2026)
  • Tasks/dev +21% (2025) → +33.7% (2026)
Faros AI, "AI Engineering Report 2026" — telemetry
Squad delivery — flat or worse
  • Incidents/PR +242.7%
  • Bugs/dev +54%
Faros AI, "AI Engineering Report 2026"  ·  Scrum.org, "From Velocity to 'Agent Efficiency'" (2026)
Different levels, same story: METR's 2025 RCT put experienced devs 19% slower on familiar code, and velocity was already soft (Jeffries, #NoEstimates, 2012) — agents just made gaming it free.
08
The Evidence

For complex, generative work the upside is real — and conditional on task type.

The advantage is real and conditional — the same evidence proves both; task type decides which.

+40%
quality uplift, human+AI vs solo — for complex, generative work
BCG RCT · Organization Science 2026 · n=758
90%
of Malone's 106 studies tested decision/classification tasks — where human+AI underperforms the best solo performer
Malone et al. · Nature Human Behaviour 2024
both
task type is the whole reconciliation — full breakdown in the appendix
Malone in full — appendix 28
BCG RCT · Organization Science 2026 (n=758)  ·  Malone et al. · Nature Human Behaviour 2024 (106 studies)
09
aha agile
What's Happening Now
Agents are generating this now
  • Requirements drafts
  • User stories
  • Sprint tracking
  • Status reports
  • Dependency mapping
  • Stand-up summaries
Nobody's governing this
  • Design choices
  • Risk thresholds
  • Approval gates
  • Production calls
  • Agentic workflow architecture
The automation is here. The governance gap is where the IM gets bypassed.
10
The Boundary
Agents generate
  • Drafts
  • Tracking
  • Analysis
  • Dependency detection
  • Test coverage
  • Communication summaries
Agents cannot decide
  • Risk tolerance calls
  • Approval gates
  • Production decisions
  • Competing stakeholder interests
  • Ethical edge cases
The boundary is not permanent. The job is to own it right now.
11
Sensors Aren't Decisions

Agents didn't retire your job. They moved it from the metric to the decision the metric forces.

Engineering reads the signals
  • Incidents per PR
  • Code churn
  • Durability of shipped code
  • Review latency
Signals live on the tech lead's dashboards · Faros AI · GitClear
You own the decisions the signals force
  • What counts as done
  • Planned review capacity
  • Where WIP is capped
  • What the sprint commits to
Flow control — Reinertsen / Theory of Constraints (decades-old)
Whoever owns the delivery outcome owns these calls — down to what stakeholders are told — and capping flow to a moving constraint has always been delivery craft, not tooling.
12
Centaur Teams

Weak human + machine + better process beat the grandmaster.

Human
  • Judgment
  • Accountability
  • Context
  • Final sign-off
AI
  • Speed
  • Throughput
  • Pattern matching
  • Draft generation
Process design is the variable. That is the IM's job.  ·  Kasparov · NYRB 2010 · 2005 Freestyle Chess Championship
13
The Mechanism

You decide which pattern. The agent can't.

This is workflow design — set before the agent runs. That's architecture, not the in-task judgment calls where adding AI degrades the decision.

AI drafts → Human reviews
Code, docs, plans
Your move: define the review gates in the Definition of Done.
Human steers → AI executes
Complex analysis, research
Your move: maintain the backlog the agents pull from.
AI monitors → Human intervenes
Ops, alerts, QA
Your move: set the escalation thresholds.
HITL taxonomy — LangGraph · DEV Community · Anthropic agent guidelines
14
aha agile
The Fork
Fragmentation path
  • Tech lead absorbs the workflow
  • A new AI-governance specialist takes the risk
  • No single owner of human-agent coordination
  • The role dissolves into its parts
Evolution path
  • You own the workflow architecture
  • You govern the human-in-the-loop patterns
  • You own risk at the human-AI boundary
  • One accountable owner — the role expands
One accountable owner beats handoffs between separate specialists.
15
The Evolved Role

Three capabilities. One survival argument.

01
Workflow architecture
Compose agent lifecycles inside the delivery system — select HITL patterns per work type, maintain the backlog agents pull from.
02
Agentic governance
Prompt evaluation, tool-usage governance, model evaluation on new releases, permission scoping, oversight gate design.
03
Risk at the boundary
Own human-agent decision splits — design approval gates, take central AI capability into the team's operating rhythm.
The organised version of capability development already happening informally — credentials at appendix 42
16
aha agile
A Proposal Primary Framework

Budget the scarce thing. It isn't hours anymore — it's judgment.

The need is real
~11 hrs/week saved; ~6.4 go straight back to supervising the agent; 69% admit shipping unverified work. Review capacity is the scarce resource.
~58% of the saving erodes (derived: 6.4÷11) · Glean Work AI Institute, "Work AI Index" (n=6,000)
The mechanism is proven
Cap work-in-progress to the binding constraint — Reinertsen / Theory-of-Constraints flow control, 15+ years old. Only the constraint moved.
Developer hours → human review capacity
The proposal
Name human-review capacity as the squad's explicitly budgeted, planned WIP limit — the attention budget. Open territory; my proposal.
No published framework budgets attention yet
Need: Glean Work AI Institute, "Work AI Index" (n=6,000, US/UK/AU, Dec 2025–Jan 2026)  ·  Mechanism: Reinertsen / Theory of Constraints WIP  ·  Synthesis: proposed here
17
aha agile
A Proposal Companion Framework

Plan agent spend like capacity — read against what shipped, not how much.

Subordinate to the attention budget: attention is the binding constraint, tokens are the purchasable one. Grounded need · a 15-year-old mechanism · my synthesis. Per-dev agent spend commonly $150–250/month, heavy shops far more; Gartner expects >40% of agentic projects canceled by end-2027.

Rate card
$/sprint per role
Tuned by the squad; squad budget = Σ (headcount × role rate).
Breach = a decision
Not an auto-halt
Top-up, descope, or downgrade tier — logged, never an automatic stop.
Reports at sprint review
Next to judged outcomes
Burn vs budget read against shipped, verified work — never against raw output.
Need: enterprise-spend + Gartner (2026)  ·  Mechanism: Reinertsen breach semantics + $X/dev industry practice  ·  Synthesis: proposed here
18
aha agile
Practice Changes VERDICT · Srivastav & Saxena 2026

The ceremony survives as a judgment ritual. The metric inside it does not.

Not an AI sticker — velocity out; judged outcomes, attention budget, agent cost in.

V · Validation
Definition of Done
Decide which agent actions need a gate before they run.
E · Evidence
Daily Scrum
Read agent action logs for deviations.
R · Runtime control
Stop-the-line
Define who can halt an agent mid-run, and when.
D · Decisions
Retrospective
Inspect agent reasoning; debug the failures.
I · Identity
Working agreements
Give every agent a named human owner.
C · Cost & compliance
Sprint review
Read agent spend against judged outcomes — shipped and verified — never against velocity.
T · Transparency
Stakeholder reporting
Make agent activity a dashboard line, not a black box.
VERDICT governance model · Srivastav & Saxena (2026, practitioner framework)
19
aha agile
The Urgency

No owner at the boundary. Here is what that looks like.

Replit · July 2025
Agent deleted a production database during a code freeze. Missing infrastructure-level approval gate.
Fortune · The Register · AI Incident Database #1152
Meta · December 2024
Agent posted sensitive user data to a public forum without authorization. 2-hour exposure. Sev 1.
SecurityBrief Asia · AI Magazine
Context-leak pattern
Data in the agent's context window becomes data the agent acts on. No authorization check.
Design implication: scope context windows explicitly at workflow-design time
20
aha agile
Action Path This Week

Start with a question, not a plan.

  • 01Who is using AI tools — chatbots only, or agents?
  • 02For what work? (code, docs, analysis, stand-up summaries, testing?)
  • 03What HITL pattern is in play — even informally? (Who reviews? Who steers? Who monitors?)
  • 04Where does the team sit on the maturity ladder? (Unseen / Observed / Controlled / Autonomous)
  • 05Where do gaps between individual adoption and team coordination show up?
  • 06What does your squad count as done — and does that number still mean anything now an agent can produce it in seconds?

The audit takes one sprint. It tells you where to put governance before someone else decides for you.

21
aha agile
Action Path

The 12-month path

This quarter
Map one workflow
Interpret the audit. Pick the workflow with the highest coordination risk, map it to a HITL pattern, design one approval gate, present the governance map to the team. Not a committee proposal — an IM making a delivery decision.
This year
Three or four workflows governed
HITL patterns visible in the Definition of Done. The squad's AI maturity is a known quantity, not an informal assumption. The IM is the conduit between the org's AI capability and the team's operating rhythm.
What success looks like
The role has not fragmented
The IM owns the boundary. Engineers move fast; the governance layer moves with them, not behind them. The 12 months did not start with a consultant — they started with a question in a sprint.
Bottom-up is not a philosophy — it is a delivery practice. The governance decisions happen in the work, not in a workshop.
22
aha agile
The Move
Own the decisions agents can't make.
The capability gap closes on its own. The governance gap is yours to close.
Homework: run the six-question squad audit; track one honest metric beside velocity for a sprint — judged-outcome cycle-time or 30-day code survival.
Or pilot the attention budget with me — speak to Adnan after the session.
Start the squad audit — this week  →
23
aha agile
Credits

The team behind this deck

Adnan AliAA
Adnan Ali
Presenter
PaxPX
Pax
Research
RexRX
Rex
Argument
CodaCD
Coda
Production
IrisIR
Iris
Design
AriaAR
Aria
Narrative & Persuasion
VeraVR
Vera
Adversarial
LarryLR
Larry
Orchestration
NolanNL
Nolan
Team
24
aha agile
Appendix
A
REFERENCE MATERIAL

Appendix

25
aha agile
Appendix Decoupling Deep-Dive

Output up, activity flat — measured twice, two ways.

+59.1%
completed story points (281 → 447) — but planned rose +≈150% (447 → 1,155, significant): the planned-vs-completed gap is the Goodhart signal, not delivery improving
Tomaz et al., FORGE '26 (arXiv:2602.13766)
p=0.928
developer activity — committed lines of code — no significant change
Same study — the P-A-E divergence
~82%
of developers reported working faster — perceived, not delivered
Same study — satisfaction high but task-dependent
+98%
PRs/dev, 2025 (+16.2% in 2026) — individual output up
Faros AI telemetry — cite year
+242.7%
incidents per PR — squad quality degrading
Faros AI, "AI Engineering Report 2026"
+54%
bugs per developer — the value side, flat or worse
Faros AI, "AI Engineering Report 2026"
Tomaz et al., FORGE '26 — 3 teams, 13 months; the gain can't be fully separated from team maturation  ·  Faros AI, "AI Engineering Report 2026"
26
aha agile
Appendix The Pre-Concession

Velocity was soft before agents. Agents only made gaming it free.

"I may have invented story points, and if I did, I'm sorry now." — Ron Jeffries, co-creator of story points
#NoEstimates (2012) and Goodhart's law said it years ago: when a measure becomes a target, it stops being a good measure. The flaw predates agents.
What changed is the cost of inflating the number fell to near zero — an agent produces points at near-zero cost, so the gaming goes vertical.
Ron Jeffries (co-originator, story points) · #NoEstimates (2012)
27
aha agile
Appendix Malone in Full

Malone et al. in full — why BCG and P&G results reconcile

106
studies, 370 effect sizes
Nature Human Behaviour 2024
90%
decision/classification tasks — where human+AI underperforms the best solo performer
Task-type breakdown
10%
generative/creative tasks — where BCG and P&G results live
The conditional scope
Malone et al., Nature Human Behaviour 2024 (106 studies, 370 effect sizes)
28
aha agile
Appendix Strong-Middle Detail

The strong-middle shape — four framings, one RCT.

The shape, four ways: the SDLC-phase compression table (what AI changes vs what stays human), the human-sandwich (Frame → Collapse → Judge), and the 80%-problem (Osmani) all describe the same division of labour.
METR's 2025 RCT: 16 experienced developers, 246 real issues, randomized per task — 19% slower with AI on familiar code.
METR's own generalizability caveats: self-selected, small-N, specific repos — the study does not claim AI fails to speed up most developers. The 2026 follow-up is selection-biased per METR itself and is not a stable re-measurement. The shape triangulates; the 19% is one context-bound number.
METR, "Measuring the Impact of Early-2025 AI on Experienced OS Developer Productivity" (July 2025 RCT) · SDLC-phase / human-sandwich / 80%-problem framings
29
aha agile
Appendix The Adoption Gap
Adoption
83%
of agile practitioners already use AI tools
AI4Agile Practitioners Report 2026 · n=289
Integration
55%
spend 10% or less of work time with AI
Same survey
Wide. Not deep.
30
Appendix The Governance Gap
Adoption
84%
of agile teams use AI tools
Digital.ai · 18th State of Agile 2025 · n=350
Governance
49%
of those teams have governance guardrails
Same survey
The 35-point gap is where the risk — and the role — lives.
31
Appendix The Structural Problem

Most businesses don't redesign. They bolt AI on.

<40%
of businesses report measurable profit gains from AI. They layer AI on top of legacy workflow — the workflow never changes.
McKinsey MGI · AI adoption research
32
aha agile
Appendix Age of Agents

Agent capability is a ladder.

As agents climb it, your job moves from doing the work, to steering it, to governing it.

  • AssistedYou operate — the agent suggests; you do the work.
  • AugmentedYou review — the agent drafts and recommends; you execute and edit.
  • CollaborativeYou supervise — the agent plans and executes alongside you; you steer.
  • OrchestratedYou architect — agents hand off to each other; you set the gates and escalation.
  • AutonomousYou govern — agents run end-to-end; you audit and hold the kill-switch.

Synthesis · autonomy framing after Feng, McDonald & Zhang (Univ. of Washington, 2025); cf. SAE J3016, Cloud Security Alliance 2026

33
aha agile
Appendix What Is an Agent

Chat ends when you stop typing. Agents don't.

LLM Chat vs Agent: Chat is stateless single-turn; Agent wraps the LLM with Context (memory · knowledge), Connections (tools · APIs), Capabilities (skills · actions), and Cadence (schedules · triggers)
34
aha agile
Appendix The Collision

Agents don't slot in. They collide — unless you compose them.

  • 01Pull vs. push lifecycle
  • 02Learning vs. execution orientation
  • 03State transition mismatch
  • 04Handoff ambiguity
  • 05Shared context gaps

Five failure modes · Webframp practitioner analysis

35
aha agile
Appendix Governance Maturity

AI Governance Maturity. Most teams are at Level 1–2.

Level 1 · Unseen
No inventory, no oversight
Agents run with no organisational awareness. Most teams are here.
Level 2 · Observed
Visibility without control
You know what's running; you can't yet govern or stop it.
Level 3 · Controlled — your target
Policies enforced, owners assigned
Actions logged, gates enforced, every agent owned. This is the move.
Level 4 · Autonomous
Agents monitoring agents
Continuous compliance, board-ready. Few are here.
VERDICT governance model · Srivastav & Saxena, 2026 (practitioner framework)
36
aha agile
Appendix Governance Incidents

Governance failure — primary sources for the named incidents

Replit · July 2025 · AI Incident Database #1152
Agent deleted a production database during a declared code freeze. Root cause: missing infrastructure-level approval gate. Sources: Fortune, The Register, AI Incident Database #1152. Post-incident fix: mandatory human approval for destructive operations.
Meta · December 2024
Agent posted sensitive user data to a public forum without authorization. 2-hour exposure, Sev 1. Sources: SecurityBrief Asia, AI Magazine, PointGuard AI. Root cause: missing authorization check — context-window data the agent was permitted to act on without a review gate.
Gartner · 2026
More than 40% of agentic deployments predicted to be decommissioned by end-2027 on uncontrolled cost, unclear value, or governance failures. Primary press release available on request.
37
aha agile
Appendix Historical Precedent

Wave 1 dissolution — a delivery role has thinned before

49%
Scrum Master training enrollment at peak (2020)
Wolpers enrollment data
<5%
Scrum Master enrollment by 2024 — a four-year collapse
Wolpers · same longitudinal data
Wolpers enrollment data · Scrum Alliance 2024 (18% of orgs now report dedicated Agile Coach demand). Dormant Q&A backup — the main line does not use the Wave narrative.
38
aha agile
Appendix Attention Budget in Full

How to actually run an attention budget.

The math: ~11 hrs/week saved, ~6.4 back to supervising the agent — ~58% of the saving erodes (derived: 6.4÷11). 69% admit shipping unverified work.
The mechanism: cap work-in-progress to the binding constraint (Reinertsen / Theory of Constraints). The constraint moved from developer hours to human review capacity.
The signal: review-queue health — pickup time, % merged unreviewed. Stop pulling agent work when review capacity is spent.
Setting and breaching the budget: name the planned review capacity as the squad's WIP limit; a breach means stop pulling new agent work, not ship faster. The budget number itself is a pilot output — no invented figure.
Glean Work AI Institute, "Work AI Index" (n=6,000, US/UK/AU, Dec 2025–Jan 2026) · Reinertsen / Theory of Constraints
39
aha agile
Appendix Token Budget in Full

The token budget — the full instrument for a pilot.

Rate card
Role → tier → $/sprint
Squad budget = Σ (headcount × role rate). Tuned by the squad, not handed down.
Breach semantics
Top-up / descope / downgrade
Logged and visible at sprint review. Never an automatic halt — a decision trigger.
Review contract
Burn vs budget
Read next to judged outcomes. Constraint, not a score — dollars as the ledger; Goodhart guard.
Growth path
Not day one
Cost-per-judged-outcome, routing rules, eval-gated top-ups — once the baseline is running.
Need: enterprise-spend + Gartner (2026) · Mechanism: Reinertsen breach semantics + $X/dev industry practice · Synthesis: proposed here
40
aha agile
Appendix Sources

Every load-bearing number — one place to check.

  1. 1Faros AI, "AI Engineering Report 2026" (22,000 devs, 4,000+ teams) — telemetry, not the DORA survey. PR review time +441%; incidents/PR +242.7%; PRs/dev +98% (2025) / +16.2% (2026); tasks/dev +21% (2025) / +33.7% (2026); bugs/dev +54%.
  2. 2METR, "Measuring the Impact of Early-2025 AI on Experienced OS Developer Productivity" (RCT, July 2025). Experienced devs 19% slower on familiar code.
  3. 3GitClear, "The AI Code Quality Maintainability Gap" (623M changes, 2023–2026). Refactoring 21%→3.8%; duplication +81%.
  4. 4Glean Work AI Institute, "Work AI Index" (n=6,000; US/UK/AU, Dec 2025–Jan 2026). 11 hrs saved / 6.4 botsitting / 69% shipping unverified.
  5. 5Tomaz et al., "Impacts of Generative AI on Agile Teams' Productivity" (FORGE '26, arXiv:2602.13766). Story points +59.1% with flat activity.
  6. 6Scrum.org, "From Velocity to 'Agent Efficiency'" (2026). Velocity as a vanity metric.
  7. 7Malone et al., Nature Human Behaviour 2024 (106 studies, 370 effect sizes; 90% decision/classification, 10% generative).
  8. 8BCG RCT · Organization Science 2026 (n=758); P&G RCT · NBER 2025 (n=776). +40% quality on complex, generative work.
  9. 9Kasparov, "The Chess Master and the Computer," NYRB 2010; 2005 Freestyle Chess Championship.
  10. 10Digital.ai, 18th State of Agile 2025 (n=350). 84% adoption / 49% governance guardrails.
  11. 11McKinsey MGI · AI adoption research. <40% report measurable profit gains.
  12. 12AI4Agile Practitioners Report 2026 (n=289). 83% use AI tools / 55% spend ≤10% of time with AI.
41
aha agile
Appendix Credentials

Start with three. Seven paths in all.

Start here · Scrum.org
PSM-AI — 1-day, launched Feb 2026
Maps to: HITL pattern governance and ceremony redesign
Start here · PMI
AI in Agile Delivery — 5-module, PDU credit
Maps to: workflow architecture and agentic governance
Start here · IAPP
AIGP — AI governance & risk professional
Maps to: risk at the human-AI boundary, context leak, governance-failure prevention
Scrum Alliance
AI for Scrum Masters
PMI
CPMAI — advanced AI project management
CI Agile
AI-Driven Agile
Agile Seekers
AI for Agile Leaders
Full credential landscape — start with the first three, each mapped to a capability from slide 16
42
aha agile
01 / 42