// Open Source — Test Strategy

AI-Native
Quality
Assurance
Strategy

AI agents compress Build from weeks to hours. Delivery only speeds up if quality moves to the edges: sharper intent before Build, automated evidence after it, and near-zero waiting in between.

Start exploring ↓ Download PDF
← Back to Publications

Faster code is not faster delivery

Faster code generation does not equal faster delivery. When Build shrinks, the constraint moves to requirements clarity, review capacity, test confidence and release approvals. Teams that simply let agents write more code will ship more defects faster.

Quality is an input to agent speed, not a tax on it.

Every hour we take out of Build must be reinvested in three places:

  1. 01

    Specify

    Executable acceptance criteria, design decisions and risk classification written before an agent starts.

  2. 02

    Verify

    Layered, largely automated validation that treats all AI-generated change as untrusted until proven.

  3. 03

    Flow

    Removing hand-offs and queues so reclaimed time reaches the customer instead of sitting in a review backlog.

Interactive

The speed paradox

Make agents build faster, then reinvest the time you save. Watch where the lead time actually goes.

Show the numbers
StageBefore agents (days)Your scenario (days)
How this model works

An illustrative model of one typical change, not benchmark data. It exists to make the dynamics visible — replace the numbers with your own value-stream data.

  • Before agents: Plan 3 days, Design 2, Build 10, Test 5, Deploy 1, waiting in queues 12 and rework from unclear intent 3 — 36 days, with a 15% change failure rate.
  • Agent speed-up divides Build time. More change volume grows waiting (× 1 + 0.1·log₂ speed-up), rework (× 1 + 0.15·log₂) and test effort (× 1 + 0.05·log₂), and unverified AI change raises the change failure rate (× 1 + 0.25·log₂).
  • Specify brings Plan to 2.5 days, Design to 1.5 and rework to 0.75, and cuts the change failure rate by 15%.
  • Verify brings Test to 1.5 days with layered, automated evidence, stops the failure rate growing with speed and cuts it by 40%.
  • Flow brings waiting to 3 days and Deploy to half a day through review SLAs, WIP limits, instant environments and progressive delivery.

This strategy sets out what we compress at each SDLC stage, the capabilities we must add to protect quality, the guardrails and metrics that govern it, and a questionnaire each team completes to chart its own journey.

Seven principles. Break one and it is not ready.

Seven principles govern every team's adoption; a practice that breaks one is not ready for production.

  1. 01 Intent before implementationNo agent writes production code without an approved spec, acceptance criteria and a risk tier. Specify

    In practice

    Executable acceptance criteria (Given/When/Then), a risk tier per change and explicit non-functional requirements, approved by the product owner before the build agent starts.

    Proven by

    Spec rework rate (stories reopened for unclear requirements) trending down, and 100% spec-to-test traceability for Tier 2 and 3.

  2. 02 AI output is untrusted inputGenerated code, tests and configuration are verified with the same rigour as a third-party dependency. Verify

    In practice

    At every tier, tests are generated from the spec by a separate agent or a human — never by the agent that wrote the code. From Tier 2, mutation testing proves they can fail; secrets, licence and provenance checks run on every change.

    Proven by

    Mutation score on changed code (above 70% for Tier 3 as a starting target) and defect density in AI-generated versus human-written code.

  3. 03 Humans own the decision, agents own the toilAccountability for merge and release stays with a named engineer; agents do the drafting, checking and evidence gathering. Accountability

    In practice

    Every stage keeps a human decision point: the product owner approves intent, the architect signs off Tier 3 designs, the engineer owns the diff, the quality owner confirms the evidence and the named service owner releases Tier 3 change. When automated gates release a Tier 1 change, the engineer who owns it stays accountable.

    Proven by

    Developer satisfaction up and review load per engineer down — agents absorb the toil, not the accountability.

  4. 04 Risk-proportionate assuranceControls scale with blast radius. A copy change and an authentication change do not take the same path. Risk tiers

    In practice

    Three risk tiers set the minimum assurance and who approves a release: automated for Tier 1, an engineer on the team for Tier 2, a named service owner for Tier 3.

    Proven by

    A full evidence trail for 100% of Tier 3 AI-generated change, with the heaviest controls reserved for the largest blast radius.

  5. 05 Evidence over attestationEvery gate produces machine-readable evidence (tests, scans, traces, approvals) rather than a checkbox. Evidence

    In practice

    Policy-as-code gates, an evidence bundle attached to every pull request, and an audit trail of AI involvement.

    Proven by

    The share of AI-generated changes with a full evidence trail trending up — 100% for Tier 3.

  6. 06 Shift left and shift rightPrevent defects with better specs; detect escapes fast with SLOs, observability and automated rollback. Prevent + detect

    In practice

    Spec review and threat modelling on the left; canary releases, SLOs, error budgets and automated rollback on the right — with production feedback flowing back into specs and tests.

    Proven by

    Escaped defects per release and mean time to restore trending down, with more defects contained in the lower layers.

  7. 07 Measure flow end to endSuccess is lead time and change failure rate for the whole value stream, never lines of code or PR count. Flow

    In practice

    A scorecard that pairs speed with stability: lead time, queue time, deployment frequency, change failure rate and time to restore.

    Proven by

    Lead time falling while change failure rate stays flat or better. Anything else has moved cost downstream.

Interactive

Ship it or stop?

Ten real-world practices. Decide whether each one is ready for production, then see which principles it puts at stake.

Scenario 1 of 10Score 0 / 0

Loading scenarios…

What we compress. What we add.

Build compresses most, but every stage has toil agents can remove; Plan, Design and Test grow in rigour even as their elapsed time falls.

Interactive

Walk the lifecycle

Select a stage to see the toil agents take on, the quality we add in return, and the decision that stays with a human.

Plan

− Compress with AI

  • Drafting epics and stories from discovery notes
  • Backlog de-duplication
  • Impact and dependency analysis

+ Add to protect quality

  • Executable acceptance criteria (Given/When/Then)
  • Risk tier per change
  • Explicit non-functional requirements (latency, SLO, security, compliance)

◆ Human decision point

Product owner approves intent and risk tier

Design

− Compress with AI

  • Generating option papers and ADRs
  • Sequence and data-flow diagrams
  • Threat-model first drafts

+ Add to protect quality

  • Design review against architecture standards
  • Threat model and data classification
  • Contract-first APIs
  • Testability designed in

◆ Human decision point

Architect or tech lead signs off design and threat model — mandatory for Tier 3

Build

− Compress with AI

  • Code, unit tests, refactors and migrations
  • Docs and boilerplate generated by agents

+ Add to protect quality

  • Spec-grounded prompts and context files
  • Approved model and tool allow-list
  • Secrets and licence guardrails
  • Small, reviewable diffs

◆ Human decision point

Engineer owns the diff and its explanation

Test

− Compress with AI

  • Test generation from acceptance criteria
  • Test data synthesis
  • Exploratory and regression runs by agents
  • Flaky-test triage

+ Add to protect quality

  • Independent verification (tests not written by the same agent as the code)
  • Mutation testing
  • Property-based and contract tests
  • AI-assisted review of every change, plus human code review from Tier 2

◆ Human decision point

Quality owner confirms evidence meets the tier; for Tier 1, policy-as-code gates confirm it

Deploy

− Compress with AI

  • Release notes and change records
  • Pipeline config
  • Rollout plans

+ Add to protect quality

  • Policy-as-code gates
  • Progressive delivery (canary, feature flags)
  • Automated rollback on SLO breach
  • Audit trail of AI involvement

◆ Human decision point

Automated for Tier 1; an engineer on the team for Tier 2; the named service owner for Tier 3

Maintain

− Compress with AI

  • Log and trace triage
  • Incident summaries
  • Dependency upgrades
  • Self-healing runbooks

+ Add to protect quality

  • SLOs and error budgets per service
  • Production feedback loop into specs and tests
  • Drift and regression detection for AI-generated code

◆ Human decision point

On-call engineer approves any remediation beyond pre-approved runbooks

Waits between stages

− Compress with AI

  • Auto-triaged reviews
  • Agent pre-review before human review
  • Parallel pipelines
  • Instant environments

+ Add to protect quality

  • Clear ownership and SLAs for review and approval
  • WIP limits
  • Flow metrics on queue time

◆ Human decision point

Team lead owns queue health

The waits between stages are usually the largest single source of lost time. Measure them before investing further in Build acceleration.

Three pillars carry the quality load

Twelve capabilities across three pillars carry the quality load that agent speed removes from Build; none is optional for high-risk services.

Interactive

Capability check

Tick what your team already has. The gaps show where agent speed will leak out as rework, escaped defects or waiting.

Goal: agent-speed delivery with no loss of quality

Specify

Measure: spec rework rate0/4

Verify

Measure: escaped defects0/4

Flow

Measure: end-to-end lead time0/4

Foundation — shared and built once by the platform and quality engineering function

Tick the capabilities your team already has.

Environments at agent speed

Agent-speed Build fails if environments stay slow, shared and hand-provisioned; every environment must be defined as code, have one clear purpose, and be available on demand.

When agents raise ten times the pull requests, a shared integration environment becomes the new queue. We move from a few long-lived, contended environments to ephemeral, per-change environments, and keep only the long-lived ones that need production-like scale, data or integrations.

Interactive

Promotion path

Pick a risk tier to see which environments a change visits, then select an environment to see what it is for and what it takes to leave it.

    Environments and what each one is for
    EnvironmentPurposeLifecycleDataTests run hereAgent roleExit criteria to promote
    Local and agent sandboxFast feedback while code is written; isolated space for agents to build and run codePer developer or per agent task; disposableSynthetic onlyUnit, component, static analysis, secrets scanAgents build and self-test here with least-privilege access and no production credentialsUnit tests pass; scans clean
    Ephemeral PR environmentValidate a single change in a realistic, isolated stackCreated on PR open, destroyed on merge or close (target under 10 minutes to provision)Synthetic plus contract stubs; service virtualisation for depend­enciesContract, API, integration of the changed service, UI smoke, accessibilityTest agents generate and run acceptance tests from the spec; agent pre-reviewAcceptance criteria verified; evidence attached to PR
    Integration (SIT)Prove the change works with real neigh­bouring services, not stubsLong-lived, contin­uously deployed from mainSynthetic and maskedEnd-to-end journeys, cross-service regression, data-flow checksAgents run regression, triage failures and flaky testsE2E suite green; no new defects above agreed severity
    Performance and resilienceVerify non-functional require­ments (latency, throughput, failure behaviour)On demand, production-like scale, torn down after runsProduction-shaped synthetic volumesLoad, soak, chaos, failover, capacityAgents generate load models and analyse results against SLO targetsNFR and SLO targets met (mandatory for Tier 3)
    Security testingProbe the running system for vulner­abilitiesOn demand or scheduled; can share the perf stackSyntheticDAST, penetration testing, authZ and data-leak checks, prompt-injection tests for AI featuresAgents run DAST and triage findings; humans lead pen testsNo open critical or high findings
    Staging (pre-production)Final rehearsal of the release in a production mirrorLong-lived, config­uration-identical to productionMasked production copy under data governanceRelease rehearsal, migration dry-runs, UAT, deployment and rollback verificationAgents verify config drift against production and run smoke suitesRollback proven; business sign-off for Tier 3
    Production (progressive)Validate with real traffic at controlled blast radiusPermanentRealCanary analysis, synthetic monitoring, feature-flag experimentsAgents watch SLOs and trigger auto-rollback on breachSLOs hold across canary stages

    Not every change visits every environment: Tier 1 changes may go from the PR environment straight to a production canary, Tier 2 changes add the integration environment and a staging rehearsal, and Tier 3 changes pass through all seven.

    Capabilities we must add

    • Environments as code — every environment defined in version-controlled infrastructure-as-code templates, reproducible by agents and humans alike.
    • On-demand ephemeral environments — self-service provisioning per PR, with automatic teardown and a time-to-live.
    • Service virtualisation and contract stubs — so a change can be tested without waiting for dependent teams or third-party systems.
    • Test data management — a synthetic data generator per domain, masking pipelines for any production-derived data, and no unmasked customer data outside production.
    • Configuration drift detection — automated comparison of staging against production, with drift raised as a defect.
    • Isolation for agents — sandboxed runtimes, scoped credentials, and network egress controls so an agent cannot reach production or exfiltrate data.
    • Environment observability and cost control — the same tracing and logging as production, utilisation dashboards, and spend limits per team.
    • Clear ownership — the platform team owns templates and tooling; each product team owns its environment configuration and data; quality engineering owns promotion criteria.

    Environment metrics

    MetricTarget direction
    Time to provision an ephemeral environmentDown, target under 10 minutes
    Time changes spend waiting for an environmentDown, target zero
    Environment-caused test failures (not code defects)Down
    Staging-to-production configuration drift itemsZero
    Environment cost per deployed changeDown

    Stack slices so the holes rarely line up

    No single test layer is trustworthy on its own, least of all when agents write both the code and its tests; we stack independent layers so their gaps rarely line up.

    The Swiss cheese model comes from safety engineering (James Reason, 1990). Each defence is a slice of cheese with holes, which are its blind spots. A hazard only causes harm when the holes in every slice happen to align. Safety-critical engineering and SRE teams rely on it, and it maps directly onto software verification.

    Interactive

    Swiss cheese simulator

    Fire defects through seven layers of defence. Switch slices off, or let the agent that wrote the code write its own tests, and watch the holes line up. Select a slice to inspect its blind spots.

    Where 100 defects are caught — expected for this set-up

    What each layer catches and misses
    LayerCatchesTypical holesHow AI helpsHow we shrink the holes
    Spec reviewWrong intent, missing edge cases, unclear acceptance criteriaUnstated non-functional needs; assumptions nobody questionsAgents critique specs and generate edge cases and negative scenariosGiven/When/Then template; risk tier; NFR checklist
    Static analysis and AI code reviewStyle, known bug patterns, secrets, vulnerable dependencies, licence issuesBusiness-logic errors; context outside the diffAI reviewer flags logic smells and explains the diff to the human reviewerTuned rulesets; human review mandatory for Tier 2 and 3
    Unit testsLogic errors in functions and classesIntegration issues; tests that mirror the code's own bugsFast generation of tests and edge cases from the specMutation testing to prove tests can fail; independent generation
    Contract and component testsBroken APIs, schema drift, service behaviour in isolationReal dependency behaviour; timing and data volumeContract generation from OpenAPI or event schemasConsumer-driven contracts checked in the pipeline
    Integration and E2EBroken user journeys and cross-service data flowsSlow, flaky, and thin on edge casesSelf-healing locators; journey generation from specs; failure triageKeep few and critical; quarantine flaky tests
    NFR and securityLatency, capacity, resilience and security weaknessesUnrealistic load or data; untested failure modesLoad models, chaos scenarios and DAST triageProduction-shaped data; SLO targets as pass criteria
    Canary and SLO rollbackAnything that escaped every earlier layer, at limited blast radiusSlow-burn defects and low-traffic pathsAnomaly detection and automated rollback decisionsSynthetic monitoring for low-traffic paths; error budgets
    How this model works

    An illustrative model, not measured data. Ten kinds of defect, weighted by how often they occur, each have a chance of being caught by each layer, based on what that layer catches and misses above. Switching off independent test generation lowers the chance that unit tests catch logic errors and missing edge cases, because tests written by the agent that wrote the code inherit its misunderstandings.

    Three rules follow from the model:

    1. Rule 01

      Independence shrinks the holes

      Two layers built from the same assumptions share the same holes. Tests generated by the agent that wrote the code inherit its misunderstandings, so tests are generated from the spec by a separate agent or a human.

    2. Rule 02

      Catch early, but never rely on one slice

      Most defects should die in the cheap, fast layers on the left. The expensive layers on the right exist for what slips through, not as the primary net.

    3. Rule 03

      Every escape reveals a hole

      When a defect reaches a later layer or production, the post-incident review asks which earlier slice should have caught it and adds a test there.

    Thicken the base. Keep the top thin.

    AI makes tests cheap to write at every level, which tempts teams to generate large, slow E2E suites; we use AI to thicken the base instead and keep the top thin and deliberate.

    Think of it like inspecting a building. Every brick and beam is tested at the factory: cheap, fast, thousands of checks (unit). Each joint is checked as it is fitted (contract and component). The whole building gets a handful of walkthroughs before handover (E2E). You would never find a weak brick by walking through the finished building, and you should not find a logic bug through a browser test.

    Interactive

    Shape your suite

    Start from a preset or drag the sliders. The outline is the AI-era target — shares are indicative targets by test count, not a quota. Select a level to see what to automate there first.

    Model assumptions

    Illustrative per-test run times: unit 5 ms, component and contract 0.5 s, integration and API 5 s, E2E UI 45 s, spread across 8 parallel runners. Share of tests that turn out flaky: unit 0.1%, component and contract 0.5%, integration and API 2%, E2E UI 8%. Replace them with your own pipeline data.

    The inverted pyramid, or “ice-cream cone” (mostly manual and UI tests), is the most common pattern in teams starting this journey. AI does not fix it by itself; used naively, it makes the cone bigger.

    Increasing automation efficiently, layer by layer

    LayerAutomate firstAI acceleratorEfficiency leverProof the tests are good
    Foundation (static)Linting, type checks, secrets and dependency scans on every commitAI reviewer comments on each PR with a plain-English diff summaryPre-commit hooks; fail fast before any test runsFalse-positive rate tracked and tuned
    UnitBusiness rules, calculations, validation and error paths on Tier 2 and 3 codeSeparate agent generates tests from acceptance criteria and edge cases; backfills legacy gapsRun on every save and commit; target under 5 minutes for the suiteMutation score, not line coverage
    Component and contractEvery API and event schema between teamsGenerate consumer-driven contracts from OpenAPI or AsyncAPI specsReplace cross-team E2E tests with contracts; run in the PR environmentContract broken in pipeline before merge, never after
    Integration and APICritical service-to-service and data-store pathsGenerate API test suites and synthetic test dataTest impact analysis runs only affected tests per changeDefects found here trace back to a missing lower-level test
    E2E UITop 10 to 20 revenue- or risk-critical user journeys onlySelf-healing locators; journeys generated from specs; automatic failure triageParallel runs; quarantine any test that flakes twiceFlaky rate under 2%; every test maps to a critical journey
    ExploratoryRisk-based charters on new features and Tier 3 changesAgents run exploratory sessions and propose new automated tests from findingsHumans focus on judgement, usability and ambiguityEach session produces either new tests or no findings, with evidence

    Efficiency levers that apply across the pyramid

    1. Start from risk, not coverage. Automate Tier 3 services and critical journeys first; a 90% covered admin page is worth less than a 60% covered core customer journey.
    2. Push every test down. When a test can move a level lower and still catch the defect, move it; it becomes faster, cheaper and more stable.
    3. Trace tests to specs. Every acceptance criterion links to at least one automated test, so gaps are visible before code is written.
    4. Run only what matters. Test impact analysis and parallelism keep PR feedback under 15 minutes as suites grow.
    5. Treat flakiness as a defect. Quarantine, fix or delete within a sprint; a flaky suite trains people to ignore red builds.
    6. Prune as well as add. AI-generated tests accumulate fast; delete duplicates and tests that never fail under mutation.
    7. Convert manual effort into charters. Manual regression scripts become automated tests; human time moves to exploratory work.

    Same eight steps. Different depth.

    Every AI-assisted change follows the same eight steps; the risk tier decides how deep each step goes, not whether it happens.

    Tiers apply to changes and to services: every change is tiered by its own blast radius, and every service carries the tier of the riskiest change it can receive.

    Interactive

    Route a change

    Pick a change. Its risk tier sets the minimum assurance, who approves the release, how deep each verification step goes and which environments it visits.

    The verification process for every change

    1. 1

      Spec locked

      Acceptance criteria, risk tier and NFRs are approved before the build agent starts.

    2. 2

      Tests derived from the spec

      At every tier, a test agent independent of the build agent generates unit, contract and acceptance tests from the criteria. They fail first (red) before code exists.

    3. 3

      Build against the tests

      The build agent writes code until the independent tests pass. It may add its own tests, but they do not count toward the gate.

    4. 4

      Foundation checks

      Static analysis, types, secrets, dependency and licence scans, and an AI review summary on the PR.

    5. 5

      Test-quality check

      Every new test must be seen to fail before it passes; from Tier 2, mutation testing on changed code proves the tests would catch a real fault.

    6. 6

      Environment verification

      Contract, API and smoke tests run in the PR environment for every change; integration and E2E tests run in the integration environment from Tier 2; NFR and security tests run for Tier 3.

    7. 7

      Review

      An agent pre-reviews every change; from Tier 2 an engineer reviews the diff, the tests and the evidence bundle, and Tier 3 adds a second human reviewer.

    8. 8

      Production verification

      Every change is released through a canary with SLO guards, synthetic monitoring and automatic rollback; any escape triggers a “which slice missed it?” review.

    Anti-patterns to stop

    • Same agent writes code and tests

      Tests confirm the code's own mistakes; both slices have the same holes.

      Independent test generation from the spec

    • Snapshot tests of current behaviour

      Locks in bugs as “expected”.

      Assert against acceptance criteria, not output

    • Coverage as the target

      Agents easily reach high coverage with assertion-free tests.

      Mutation score and spec traceability

    • AI-generated E2E sprawl

      Slow, flaky suites that block the pipeline.

      Cap E2E at critical journeys; push tests down

    • Skipping human review because “tests pass”

      Tests only check what someone thought to specify.

      Risk-tiered human review, always for Tier 2 and 3

    • Retrying flaky tests until green

      Hides real intermittent defects.

      Quarantine and fix within a sprint

    Pyramid and verification metrics

    MetricWhat it tells usTarget direction
    Defect containment by layerWhich slice caught each defect; escapes show where holes alignMore defects caught in lower layers; production escapes down
    Mutation score on changed codeWhether tests can actually detect faultsUp; above 70% for Tier 3 as a starting target
    Test distribution across the pyramidWhether we are building a pyramid or an ice-cream coneBase widening quarter on quarter
    Spec-to-test traceabilityShare of acceptance criteria with an automated test100% for Tier 2 and 3
    PR feedback timeTime from push to full test resultUnder 15 minutes
    Flaky test rateTrust in the pipelineUnder 2%
    Manual regression effortHours of scripted manual testing per releaseDown, towards zero

    Targets are indicative starting points; each team sets its own baseline in the questionnaire and improves from there.

    Judge flow and stability together

    Assurance scales with risk tier, and progress is judged on flow and stability together — never on AI usage alone.

    Scorecard

    DimensionMetricDirection
    SpeedLead time for change (idea to production)Down
    SpeedQueue time between stages as % of lead timeDown
    SpeedDeployment frequencyUp
    StabilityChange failure rateFlat or down
    StabilityMean time to restoreDown
    QualityEscaped defects per release; defect density in AI-generated code vs human-writtenDown
    QualityMutation score on Tier 2 and 3 servicesUp
    IntentSpec rework rate (stories reopened for unclear requirements)Down
    Trust% of AI-generated changes with full evidence trailUp, target 100% for Tier 3
    ExperienceDeveloper satisfaction and review load per engineerUp / down

    A team whose lead time falls while change failure rate rises has not improved; it has moved cost downstream.

    Operating rhythm

    • MonthlyTeam review of the scorecard with the quality engineering lead.
    • QuarterlyMaturity reassessment using the questionnaire below.
    • Always onA central register of AI tools, models and agent permissions, reviewed by security and risk.

    Level 2 today. Level 4 in 15 months.

    Most teams will self-assess at Level 2 today; the target is Level 4 for all product teams within 15 months, with Tier 3 services never above Level 4.

    Interactive

    Where are you today?

    Choose the description that matches your team in each dimension. The next level tells you what to build next.

    The full maturity model
    LevelNamePlan and designBuildTestDeploy and operate
    1Ad hocInformal requirements; no AI useManual coding; individual AI experimentsMostly manual, late testingManual releases and approvals
    2AssistedAI drafts stories; criteria varyInline AI code assistance; no shared standardsSome AI-generated unit testsCI/CD in place; manual gates
    3IntegratedSpec templates and risk tiers used consistentlyAgents build from approved specs with repo contextIndependent test generation; mutation testing on key servicesPolicy-as-code gates; canary releases
    4AgenticAgents draft specs, designs and threat models for human sign-offMulti-agent build with guardrails and full evidence trailAgentic regression and exploratory testing; evidence per tierSLO-driven auto-rollback; agent-assisted incident triage
    5Autonomous with guardrailsProduction signals draft spec updates for human approvalAgents deliver Tier 1 changes end to end; a named engineer stays accountableContinuous verification in productionSelf-healing runbooks, pre-approved by on-call, act within error budgets

    Four phases, three gates

    Four phases move every team to AI-native delivery within 15 months. Timings count from the start of adoption.

    1. Phase 1

      Foundations

      Months 1 to 3

      • Platform, guardrails
      • Risk tiers agreed
      • Baseline scorecard
      • Questionnaires in
    2. Gate ABaseline captured
    3. Phase 2

      Pilot

      Months 4 to 6

      • 2 to 3 pilot teams
      • Spec and test agents
      • Evidence gates live
      • Playbook v1
    4. Gate BPilot CFR holds
    5. Phase 3

      Scale

      Months 7 to 12

      • All teams onboard
      • Shared skills library
      • Tier 3 controls live
      • Review SLAs enforced
    6. Gate CLead time down
    7. Phase 4

      Optimise

      Month 13 onward

      • Agentic ops loops
      • Self-healing runbooks
      • Tier 1 fully automated
      • Continuous tuning
    Interactive

    Gate check

    Each gate is a go/no-go: a phase does not start until the previous phase's scorecard shows lead time falling with change failure rate flat or better. Set the trends and try it.

    Lead time

    Change failure rate

    Write your own AI-native journey

    Each team completes this in a 90-minute working session with its tech lead, product owner and quality lead; the answers become its baseline for Gate A.

    Answer honestly: this is a baseline, not an audit. Where you do not know, write “unknown” — that is a finding in itself.

    Interactive

    Team questionnaire

    Answers are saved in this browser only — nothing is sent anywhere. Export them as Markdown for your team's working session.

    0 of 61 answered

    Part A — Team profile
    Part B — Where are you today?

    Rate each stage against the maturity model (1 to 5) and give one piece of evidence.

    Plan
    Design
    Build
    Test
    Deploy
    Maintain
    Waits between stages
    Part C — Stage deep dive
    Part D — Flow and waiting
    Part D2 — Test environments
    Part D3 — Test pyramid and verification
    Part E — Risk and guardrails
    Part F — Your journey
    Part G — 90-day commitments

    Fork it. Adapt it. Make it yours.

    This strategy is published openly so any team can adapt it to its own context. The page — text, simulators and all — lives in this site's public source repository under the MIT License.