Skip to main content

There's plenty of advice on how individuals should use AI coding tools. Less clarity on how teams should work together differently, or how leaders can tell whether the new workflow produces better software.

The leverage comes when the entire team changes how it works. Sprint planning, documentation, testing, reviews, security, ownership, and measurement need to support AI-native engineering together. Faster first drafts are useful only when the team can verify them, ship them reliably, and maintain what arrives in production.

The operating principle of this playbook remains the same: agents produce implementations; engineers build and operate the system that verifies them. The measurement layer tests whether that shift delivers useful, durable work rather than additional activity.

Change Your Defaults: From Coding to Verification

The traditional sequence is to write code, review it, test it, and ship. In a verification-led AI-native workflow, engineers invest first in the specification, architecture, constraints, and evidence required to accept a change. Agents generate an implementation within those boundaries. Humans remain accountable for design, security, review, and the accepted result.

This is an operating-model choice, not a requirement to maximize AI code share or eliminate manual coding. Some tasks warrant direct implementation, some suit assistance, and others can be delegated to an agent. Choose the mode that fits the task and its risk.

The New Workflow

1. Make the Specification the First Engineering Artifact

Spend more effort on design and implementation planning before generation. Pair engineers for design and specification review, not only for coding. Break work into small phases with independently testable outcomes.

A useful specification states the problem, interfaces, relevant context, constraints, edge cases, non-goals, security requirements, acceptance criteria, and definition of done. Make architectural tradeoffs explicit: an intentional compatibility shim or temporary duplication should be explained before an agent tries to simplify it away.

The Superpowers framework and Harper's design-before-implementation guidance offer practical approaches to planning. Use a process your team can apply and verify consistently rather than assuming one framework fits every repository.

2. Use Independent Review Without Treating Model Agreement as Proof

The original playbook describes using Opus 4.5 and GPT 5.2 to review implementation plans. Keep the underlying practice: use another model or review approach to challenge assumptions, missing cases, and proposed architecture.

Agreement between models can help surface issues, but both may share a blind spot. Engineers must resolve disagreements and validate the plan against the system's actual requirements. Independent tests, security checks, and human judgment remain necessary.

3. Put Test-Driven Verification in the Workflow

For behavior changes, write tests that fail for the right reason before delegating implementation. Use tests as an acceptance boundary, not as a metric to inflate. End-to-end tests can verify user-visible behavior; integration tests check interactions; focused unit tests isolate logic where they add useful evidence.

Keep tests human-readable and grounded in the specification. A test suite generated from the same mistaken assumptions as the implementation can pass while the result is wrong. Review the cases, failures, and coverage appropriate to the change.

4. Increase Useful Constraints

Use types, linting, interface contracts, architecture rules, and explicit dependency policies to narrow the space of acceptable implementations. Strong compile-time checks can catch classes of error; they do not prove correctness or security.

The original approach increases linting and typing in TypeScript and Python rather than requiring a language migration. Add constraints that address observed failure patterns in your codebase.

5. Select Models by Verified Task Performance and Total Cost

More capable models may justify their price on complex work, but the highest-priced model is not automatically the best choice for every task. Test model fit against comparable work, acceptance criteria, reliability, latency, review burden, and cost.

Read benchmark scores versus real-world coding-agent performance as two complementary evaluation layers. A held-out benchmark informs capability selection. Production sessions show how an agent investigates your repository, completes the task, verifies changes, and performs consistently under your operating conditions.

Do not convert a benchmark score into a production success estimate or assume success on one attempt means repeatable success. Use representative tasks and inspect failures as well as the happy path.

6. Automate Pre-Commit Checks and Preserve Human Review

Run relevant tests, linting, simplification checks, and security analysis before changes advance. Claude Code tooling can support review and execution, but a passing automated check is not a substitute for architecture and risk review.

Attach the specification, verification results, changed assumptions, and known limitations to the pull request. Reviewers should be able to see what was checked and what remains uncertain.

7. Make Documentation Useful Context

Maintain modular code and current documentation covering plans, interfaces, designs, decisions, and sprint context. Give agents the reason for past choices as well as the current implementation. A stale design document can be worse than missing context if an agent treats it as authoritative.

Record where context came from and what changed during the task. Keep documentation maintenance part of done when behavior or architecture changes.

8. Treat Security as a Separate Verification Responsibility

  • Static analysis: Run appropriate tools such as Semgrep, CodeQL, or Bandit to identify known patterns.
  • Security constraints: Specify safe practices for database queries, credential handling, authorization, and data access.
  • Human security review: Ask how the change could fail dangerously, not only whether it works in the expected case.
  • Dependency auditing: Check packages and versions with tools such as npm audit, pip-audit, or Snyk, and apply an approved-dependency policy where appropriate.
  • Execution boundaries: Scope agent permissions, network access, and secrets to the task. Use explicit approvals for consequential actions.

Measuring AI-generated code quality at scale extends these checks into downstream defects, incidents, rework, and durability. Clean syntax, test coverage, or a successful merge alone does not establish production quality.

Context Engineering and Agent Infrastructure

Lance Martin's Effective Agent Design provides a useful foundation for designing the information and action space around an agent.

  • Provide a controlled environment with the filesystem and execution tools required for the task.
  • Use a small, understandable set of actions and introduce additional context when it becomes relevant.
  • Offload large intermediate results and retain enough state to continue the work.
  • Use caching where it preserves correctness and improves economics.
  • Isolate parallel tasks when they do not compete over shared implementation state.

The Ralph Wiggum pattern uses repeated execution with persistent context and verification hooks. Add bounded retries, budget limits where enforceable, escalation, and stop conditions. “Continue until done” is not a safe substitute for a definition of done or an intervention plan.

Containerized development environments and clean CI/CD support repeatable agent work. Sandboxing reduces exposure but does not eliminate risk. Validate permissions, secrets, production access, rollback, and auditability before allowing unattended execution.

Separate Adoption, Integration, and Impact

The adoption–integration gap separates three questions: who has access, who actively uses AI, and how AI participates in engineering work. A weekly login can describe occasional documentation help or sustained implementation work; it cannot distinguish them by itself.

Track access, active usage, workflow depth, and AI-assisted contribution separately. Then examine delivery, quality, and cost. Low activity may reflect onboarding, IDE fit, security delays, or task mismatch rather than resistance. High activity with weak outcomes may require workflow redesign or better verification, not additional licenses.

Larridin's elite adoption benchmarks combine usage, AI code share, delivery speed, quality, and economics. Treat them as internal benchmark context with methodology and population limits, not universal quotas. The right AI code share depends on codebase, task, and risk; forcing it upward can reward volume that later becomes rework.

Measure Shipped Work Rather Than Activity

Measuring AI's impact on engineering output distinguishes work that reaches production from the activity that precedes it. Features, resolved incidents, migrations, and useful debt reduction are different from commits, suggestions accepted, or pull requests opened.

Define the outcome for the workflow. A merged PR is an accepted intermediate artifact, not necessarily a customer-facing production outcome. Preserve that distinction in reporting and cost calculations.

The developer productivity metrics guide recommends connected views of adoption and attribution, complexity-adjusted work, quality and rework, workflow performance, and cost. A single count cannot establish whether the team solved harder problems or only generated more boilerplate.

Use Complexity and Quality Decomposition

Engineering Output is Larridin's composite measure of merged engineering work. It assesses complexity across logic, scope, architecture, risk, and novelty, maps that assessment to a superlinear weighting, and applies penalties for warranted missing tests and defined low-quality patterns.

Use the decomposition to understand a trend: did work mix change, did test gaps rise, or did low-value volume expand? The score is a leading indicator, not direct proof of customer value or lasting quality. Pair it with durability, delivery, security, and business outcomes. Ask for definitions and calibration before using any composite score in leadership reporting.

Keep measurement at team and workflow level. Assignment mix influences complexity and output, so the metric should diagnose the operating system rather than rank individual engineers.

Find Where the Delivery Constraint Actually Lives

The AI code review bottleneck analysis and change lead time guidance explain why faster generation may not shorten delivery. Review is a possible constraint, not the default explanation in every team.

Measure each stage using explicit boundaries:

  • Implementation and readiness for review.
  • Reviewer pickup or waiting before active review.
  • Active review and requested corrections.
  • Tests, security checks, CI validation, and failure recovery.
  • Approval-to-merge waiting.
  • Merge-to-production deployment.

PR-open-to-merge cycle time and commit-to-production change lead time are not interchangeable. Neither automatically measures all time spent specifying and generating the change. Keep elapsed waiting separate from active effort where possible.

If pickup time grows, investigate ownership, routing, work in progress, scarce expertise, and PR size. If validation dominates, examine CI capacity, test failures, and flaky checks. If merge or deployment dominates, examine release coordination and policy. Adding reviewers or generation capacity without diagnosing the stage can optimize the wrong constraint.

Keep AI-assisted changes small and independently testable. Include the context needed for review and constrain concurrent agent work when downstream capacity cannot absorb it. Shorter review is not an improvement if necessary scrutiny was removed.

Keep DORA, Add AI Attribution and Verification Context

AI-native DORA guidance complements delivery metrics with AI contribution, verification, and rework. The coding-agent deployment example reinforces the central lesson: rising deployment frequency can coexist with longer lead time and worsening failures.

The source articles use both the classic four-metric terminology and the updated five-metric delivery model. State the version and definitions used. A current scorecard can include deployment frequency, change lead time, change fail rate, failed deployment recovery time, and deployment rework rate.

DORA describes delivery performance; it does not independently identify AI's causal effect or the business value of each deployment. Segment where attribution supports it and read speed alongside failures, recovery, rework, and code durability. Do not turn these measures into individual rankings or compare unlike services as if their workloads and risk were identical.

Verify Quality After Merge and Preserve Code Comprehension

Verification continues after the pull request closes. Track defects, security findings, incidents, rework, turnover, and reversions over consistent observation windows. Recent changes need the same opportunity to exhibit problems as older changes before comparing cohorts.

The revert-rate comparison reports substantially different AI-assisted and human-only results in one customer environment. That is evidence from that environment, not proof that one source always produces better code. Revert rate measures direct reversal; code can remain while creating maintenance debt or defects fixed through other changes.

Keep revert, turnover, and rework definitions distinct. Inspect refactoring intent, task difficulty, reviewer practices, and incident context before attributing a quality difference to the coding tool.

Reviewers Must Understand What They Accept

The code-comprehension article discusses a July 2026 preprint involving 54 students building small websites. It found comprehension concerns associated with accepting agent edits without reading or adapting them. The study is preliminary, not peer reviewed, and does not establish the same effect in professional teams or production code.

The operational implication is worth testing: engineers should be able to explain the relevant behavior, dependencies, failure modes, and maintenance path of a change they accept. Use walkthroughs, debugging exercises, and review discussions where appropriate. Session signals such as verification and intervention can reveal practices; they do not directly measure understanding.

Make junior development part of the operating model. Delegation should create space for learning architecture, testing, and debugging, not leave engineers dependent on output they cannot evaluate.

Measure the Full Coding Stack Without Double Counting

Multi-tool productivity measurement requires a common model for identities, teams, repositories, work types, time periods, costs, and outcomes. A developer using Copilot, Cursor, and a coding agent is still one developer. One accepted change may involve all three tools.

Use native dashboards as inputs, not interchangeable totals. Normalize definitions and preserve the evidence for contribution. Where joint work cannot be attributed reliably, label it mixed or uncertain rather than assigning all of the value to every tool.

AI-assisted versus human-only comparisons should match task type, repository conditions, complexity, experience, and measurement windows. Controlled research has produced both speedups and slowdowns under different conditions; neither is a universal forecast for your team.

Use a comparable control or phased rollout where practical. Track changes in staffing, task mix, process, and release pressure. Developer feedback helps explain friction and confidence, but does not replace system-observed results or establish financial savings.

Follow Agent Runs Through Accepted Outcomes

Agentic coding productivity measurement starts with the run or assigned task. Capture its owner, team, tool, model, repository, purpose, start and end, resulting changes, usage, and costs where available.

Distinguish autonomous agent work, human AI-assisted work, human-only work, and mixed or unattributed work. Track task acceptance, planned approval, unplanned intervention, retries, incomplete cleanup, and corrections before merge. A technically completed run may still hand the difficult work back to an engineer.

Inspect investigation depth, completion, verification, and consistency across real sessions, then follow the output into review and production. Unknown trace or tool coverage should be visible. Neither raw run count nor a single accuracy score can explain all of these failure modes.

Make the Budget Follow Verified Work

Cost per shipped outcome gives finance and engineering a common unit: attributed AI spending for a period divided by consistently defined accepted outcomes associated with that spending.

Include usage from retries, abandoned sessions, and unsuccessful work in the stated cost scope. A low cost per successful session can hide money spent on failed sessions if those charges disappear from the calculation.

Define whether the denominator is merged PRs, production features, resolved tasks, or verified durable changes. A tool-spend-per-merged-PR measure is not the total cost of production delivery. Show broader economics separately: licenses, tokens, infrastructure, implementation, human review, debugging, remediation, and ongoing support.

Compare matched work and read cost alongside complexity, quality, durability, and delivery. A more expensive tool may be justified if it reduces total effort or supports harder work. A cheaper one may create more recovery. Separate cash savings, redeployable capacity, and additional output when expressing the financial benefit.

Give the CTO a Decision Scorecard

The CTO scorecard brings five metric sets together:

Metric SetRead TogetherDecision It Supports
DeliveryEnd-to-end lead time, stage waits, deployment frequency, failures, recovery, and reworkWhich pipeline constraint needs attention?
Adoption and proficiencyAccess, sustained usage, contribution, workflow fit, and verification practicesWhere is enablement or redesign needed?
Quality and riskTests, security, incidents, reverts, turnover, and comprehension checksAre generation gains creating downstream debt?
Cost and valueAttributed spending, failed-run costs, accepted outcomes, and broader workflow effortWhat should expand, change, or stop?
Developer experienceReview burden, cognitive load, context switching, feedback, and workflow frictionIs the operating model making useful work easier?

Use trends with drill-down by team, repository, work type, and risk. Every material finding needs an owner and next action. Weekly operational checks can catch cost spikes and quality regressions; a leadership review can address workflow, tool, and investment decisions. Choose a cadence that fits the risk rather than waiting for board reporting.

Team Structure and Anti-Patterns

Create explicit coordination responsibility for the human-agent workflow. A producer or similar role can keep specifications, reviews, environment readiness, and handoffs moving. Protect focused work while preserving fast feedback and shared learning, whether teams are co-located or distributed.

  • Do not skip required verification because the change looks plausible.
  • Do not advance unresolved review findings without an explicit risk decision.
  • Do not let parallel agents modify shared implementation state without isolation and coordination.
  • Do not optimize PR count, AI code share, or benchmark labels at the expense of accepted outcomes.
  • Do not assume low revert rates prove maintainability or passing tests prove security.
  • Do not reward individual engineers for activity counts that reflect task assignment and tool behavior.
  • Do not leave agent permissions, budgets, and stop conditions implicit.

Framework-specific rules need context. For example, how an agent receives a plan or whether implementation tasks can run in parallel depends on the harness, task dependencies, and isolation. Preserve the intent, which is clear context and conflict prevention, rather than treating every tool convention as a universal engineering rule.

Roll Out the Operating Model With Evidence

  1. Choose a bounded workflow. Define representative tasks, repository, risk, acceptance criteria, and human accountability.
  2. Establish a baseline. Capture delivery, quality, effort, and cost over a period representative of the work. There is no universal history requirement; disclose missing baseline evidence.
  3. Build the verification harness. Provide tests, context, security checks, sandboxing, permissions, and recovery.
  4. Instrument the chain. Connect tool activity and agent runs to commits, reviews, delivery, and downstream quality, preserving mixed and unknown attribution.
  5. Test the changed workflow. Compare like work and examine failed runs and human repair, not only successful demos.
  6. Fix the constraint. Improve specifications, context, routing, CI, review capacity, or model fit based on the evidence.
  7. Scale what holds. Transfer validated practices and continue tracking quality and economics as tools and workloads change.

Larridin's Developer Productivity, Agent Effectiveness, Workflow Intelligence, and Token Spend & Insights connect parts of this measurement chain. Confirm integration coverage and methodology for your environment, and retain specialist security and code-review controls where needed.

Frequently Asked Questions

Does AI-native engineering mean developers stop coding?

No. It shifts more effort toward specification, orchestration, architecture, verification, and maintenance. Direct coding remains appropriate where it best serves the task. Engineers still need to understand the systems and changes they accept.

What should we measure first?

Start with a bounded workflow and a baseline for accepted work, quality, effort, and cost. Track actual AI involvement separately from access. Add the stage-level and agent-session signals needed to explain the current constraint.

Are DORA metrics obsolete when agents write code?

No. They remain delivery measures. Add AI attribution, verification, quality, durability, and cost context, and state the metric definitions used. Faster delivery counts as progress only when the surrounding outcomes and guardrails hold.

Does low revert rate prove AI code is better?

No. It indicates fewer direct reversals within the stated window. Pair it with turnover, rework, defects, security findings, incidents, and maintenance evidence, using comparable work and equal observation windows.

Should every team target elite AI usage or code share?

No. External and internal benchmarks provide context, not universal quotas. Determine healthy integration from task fit, risk, durable output, verification, and economics. Raising usage without stronger outcomes is not the objective.

How do we compare multiple coding tools fairly?

Normalize identities, repositories, tasks, periods, costs, and outcome definitions. Compare like workflows, disclose overlapping contribution, and include human review and recovery. A universal tool ranking can conceal differences in work and context.

What does a public coding-agent benchmark tell us?

It tests capability on a defined task set. It does not establish production performance in your repository. Combine benchmark context with real-session investigation, completion, verification, consistency, and downstream outcomes.

Sources and Further Reading

This playbook synthesizes practitioner guidance and the source articles below. The linked research and customer examples use different populations, periods, and definitions; they should not be combined into a universal productivity multiplier. Internal benchmarks are context, not causal proof or individual performance targets.

Output, Attribution, and Adoption

Delivery, Review, and Quality

Agent Evaluation, Cost, and Leadership

Frameworks and Practitioner References

Thanks to Justin McCarthy for insights on running AI-native teams, Sam Schillace for experience with Amplifier, Jesse Vincent for building Superpowers, and James Cham for connecting practitioners.

The objective is a team that can specify useful work, delegate appropriately, verify rigorously, and ship software whose quality and economics hold after the first draft. Build the verification system, then measure the work it enables.