12.4

View in English

12.4 Maturity self-assessment

Every chapter in this guidebook ends with a “Maturity model” that describes how a practice typically evolves. This appendix consolidates every chapter’s maturity model into one reference so you can assess a team, a domain, or a whole organisation at a glance.

The shared five-level scale

All chapters describe the same progression. The exact wording varies slightly between chapters, but the intent maps cleanly onto these five levels:

  • Level 1, Initiate. Ad hoc, reactive, and personality-driven. Practices exist only where an individual chooses them, so outcomes depend on heroics and luck.
  • Level 2, Develop. Basic practices exist, but they are inconsistent across teams, partly manual, and often bypassed under pressure.
  • Level 3, Standardise. Practices are documented, standardised, and enforced org-wide. This is the audit and compliance floor: the level most enterprise and government work must reach to be dependable and auditable.
  • Level 4, Manage. Practices are measured and controlled with data and metrics against baselines. You know quantitatively how each practice performs, and you act on the numbers.
  • Level 5, Orchestrate. Practices are continuously improved, integrated across the organisation, and adaptive. The safe or correct path is the default, and the organisation learns and evolves deliberately.

How to use it for self-assessment

  1. For each chapter relevant to your context, read the five cells below and pick the level that honestly describes your typical behaviour, not your best team on its best day, and not your written policy, but what actually happens.
  2. Score each chapter 1 to 5. Round down when in doubt; a practice that is inconsistent is Level 2, not Level 3.
  3. Average scores within a part to see where a whole domain stands, then look at the spread: a part at “3 on average” that hides a Level 1 chapter still has a Level 1 risk.
  4. Re-assess periodically and track the trend. Movement matters more than any single snapshot.

Maturity is a means, not an end

Higher maturity is not automatically better. The goal is fit: enough rigour to manage the risk and scale you actually face, and no more. A small, low-stakes tool does not need Level 5 chaos engineering. Reaching for a high level as a trophy, rather than to solve a real problem, produces ceremony without value. Read every “Level 5” below as “appropriate when the stakes justify it,” and let risk, scale, and regulatory exposure decide how far to climb.


Part 1. People

TopicLevel 1 InitiateLevel 2 DevelopLevel 3 StandardiseLevel 4 ManageLevel 5 Orchestrate
Engineering culture and valuesCulture is accidental and personality-driven; incidents mean blame; knowledge lives in a few heads.Some teams run postmortems and write docs, but practice is inconsistent and unreinforced by leadership.Blameless learning, ownership models, and a writing culture are org-wide norms with clear expectations and tooling.Culture health is measured (psychological-safety surveys, incident-learning rates, retention) and tracked against baselines and acted on.Culture is continuously improved and practices spread between teams; leadership adapts norms as the organisation grows and learns.
Team topologiesTeams form by accident or headcount; structure mirrors legacy hierarchy; dependencies everywhere.Some stream-aligned teams exist, but shared bottlenecks and functional silos persist.The four team types and explicit interaction modes are used deliberately; platforms and InnerSource cut dependencies.Cognitive load, flow, and dependency counts are measured per team against targets; boundaries are adjusted when the numbers slip.The organisation continuously reshapes teams and interaction modes to sustain flow as products and platforms evolve.
Roles, career ladders, growthNo written ladder; promotions and pay are ad hoc and personality-driven.A basic ladder exists but is applied inconsistently; no calibration; hiring is unstructured.Dual tracks, a clear competency matrix, calibration, and structured hiring are standard.Progression rates, pay equity, and time-in-level are measured against baselines; calibration outcomes are analysed for bias.The framework evolves continuously with the work; sponsorship and apprenticeship are deliberate and org-wide as roles change.
Ways of workingProcess is ad hoc or cargo-cult; communication is meeting-driven and undocumented; estimates are treated as promises.A methodology is followed consistently, but ceremonies are rote and cross-team coordination is heavy.Practices are chosen to fit context; async, docs-first communication is the norm; estimation informs, not controls.Flow metrics (lead time, work in progress, throughput) are tracked against baselines and reviewed each cycle.Teams continuously tune their way of working from those metrics; coordination need is minimised at the source and good practice spreads org-wide.
Decision-making and governanceDecisions are ad hoc and unrecorded; governance is absent or a blanket bottleneck; debt is invisible.Some decisions are documented and some review exists, but process is inconsistent and mismatched to decision weight.ADRs, a paved road, reversibility-based delegation, and a debt inventory are standard and transparent.Decision cycle time, reversal rates, and debt levels are measured; scrutiny is calibrated by decision weight against those numbers.Governance is continuously tuned across the organisation; scrutiny targets irreversible decisions; debt and sourcing are managed as evolving portfolios.

Part 2. Software Programming

TopicLevel 1 InitiateLevel 2 DevelopLevel 3 StandardiseLevel 4 ManageLevel 5 Orchestrate
Coding standards and styleStyle is per-author; no shared configs; formatting is argued in review.Each team has a formatter and linter, but configs and rules vary across teams.Central shared configs per language; CI enforcement; new repos inherit standards via templates.Standard adoption, violation rates, and review-time impact are measured against baselines; configs are versioned and governed.Standards are continuously refined from that data and shared org-wide; enforcement is near-frictionless and adapts to new languages.
Software design principlesDesign is ad hoc; coupling accumulates; principles are unknown or invoked as slogans.Teams know the principles and apply them, but inconsistently and often dogmatically.Shared design vocabulary, deliberate coupling/cohesion analysis, and bounded contexts aligned to teams.Coupling, cohesion, and change-failure metrics inform design reviews against baselines; decisions are recorded.Design decisions are revisited as evidence accrues; principles are applied with nuance and paradigm choices adapt org-wide as the domain evolves.
APIs and interface designAPIs emerge from implementation; no shared conventions; breaking changes are common and unannounced.Teams follow basic REST conventions and version informally, but consistency and docs vary.Contract-first design, machine-readable specs, a deprecation policy, and consistent error/pagination conventions.Adoption, latency, error rates, and breaking-change frequency are measured per API against targets.APIs are governed products in a catalogue with strong DevEx; the practice adapts continuously and breaks are rare and well-managed org-wide.
Testing strategyTesting is manual and ad hoc; automated coverage is minimal; regressions are frequent.Automated unit and some integration tests exist, but the suite is slow or flaky and trust is low.A balanced, fast, reliable suite gates every change; flakiness is managed; non-functional testing is integrated.Coverage, flakiness, escaped-defect, and suite-duration metrics are tracked against baselines to target effort.Advanced techniques (property, mutation, fuzz) target high-value code; the strategy improves continuously and spreads across teams.
Code review and collaborationReview is inconsistent or skipped; mechanical issues dominate; feedback norms are unset.Review is required but slow and variable; automation is partial; PR size and quality vary widely.Small PRs, automated mechanical checks, clear standards and feedback norms, and monitored latency.Review latency, PR size, and defect-escape rates are tracked against targets; depth is matched to measured risk.The organisation continuously improves review from that data; pairing and AI assistance are adopted deliberately and practices spread between teams.
Version control and source managementAd hoc branching; long-lived branches; poor messages; no secret scanning; frequent merge pain.A consistent branching model and message conventions exist, but branches live too long and enforcement is partial.Trunk-based development, protected mainline, enforced commit conventions, secret scanning, deliberate repo structure.Branch lifetime, merge frequency, and revert rates are measured against delivery metrics and baselines.Automation enforces hygiene end to end; repo structure and workflow evolve continuously across the organisation as delivery needs change.
DocumentationDocumentation is sparse, scattered, and stale; knowledge lives in people’s heads.Key docs exist (READMEs, some runbooks) but are inconsistently maintained and hard to find.Docs-as-code with clear structure, generated API docs and changelogs, decision records, and update expectations.Doc coverage, freshness, and accuracy are measured against baselines; staleness is flagged automatically.Docs are living, largely generated or tested against the system, owned and discoverable; the practice improves continuously org-wide.

Part 3. Systems

TopicLevel 1 InitiateLevel 2 DevelopLevel 3 StandardiseLevel 4 ManageLevel 5 Orchestrate
Architecture fundamentalsArchitecture is implicit and lives in heads; no quality attributes or ADRs; decisions surface during incidents.Key diagrams exist and major decisions are sometimes recorded; quality attributes named but rarely quantified; docs drift.Quality-attribute scenarios and ASRs are specified; ADRs routine; C4/arc42 docs maintained near code; trade-off reviews happen.Fitness functions enforce quality attributes in CI and record measured results against baselines; trade-offs are quantified.Architecture evolves continuously with that data across the organisation; docs stay trustworthy enough for auditors as the system adapts.
Architectural styles and patternsOne tangled monolith or accidental distributed mess; boundaries follow layers or history; style chosen by fashion.Deliberate modular boundaries or a few coarse services; some cross-cutting concerns consistent; splits still ad hoc.Services aligned to bounded contexts owning their data; gateway/BFF where apt; clean/hexagonal layering standard.Style decisions are evidence-based, using measured coupling, latency, and change-cost data against baselines.A mature platform makes distribution cheap; the organisation re-consolidates when a split stops paying and adapts style as evidence changes.
Distributed systemsRemote calls treated as local; no/naive retries; failures cascade; debugging is per-machine log spelunking.Timeouts and basic retries exist but inconsistent; some idempotency; logs centralised but uncorrelated.Idempotency, backoff, circuit breakers, bulkheads via shared libs; sagas; distributed tracing; documented consistency per flow.Resilience is measured against SLOs; fault-injection results and failure rates are tracked against baselines.Resilience is the platform default, continuously tested with fault injection; graceful degradation is designed in and evolves org-wide.
Data architecture and storageOne database for every purpose; no migration discipline; incidental caching; scale by bigger machine.Storage choices mostly deliberate; a cache and maybe a warehouse; versioned migrations sometimes need downtime.Polyglot persistence matched to workloads, each store owned; automated zero-downtime migrations; explicit caching and replicas.Storage choices are measured against access patterns, latency, and cost baselines; sharding and caching decisions are data-driven.Data architecture is continuously reviewed and evolved across the organisation; migrations are automated and audited as workloads change.
Scalability, performance, resilienceSingle-instance or vertically scaled; server-side state; no load testing or budgets; failures cause full outages.Horizontally scaled stateless tiers; basic autoscaling; some pre-launch load testing; DR documented but rarely tested.Capacity planned with headroom; performance budgets in CI; resilience patterns standard; RTO/RPO defined and DR tested.Capacity is forecast from measured load; performance budgets and RTO/RPO are tracked against baselines.Multi-region automated failover, continuous chaos, and game days prove and improve recovery objectives as the system evolves org-wide.
Legacy modernisationLegacy is feared and frozen; no inventory; modernisation is all-or-nothing rewrite; knowledge in retiring heads.An inventory exists and some risk understood; legacy wrapped with APIs; still big-bang thinking; migration underestimated.Systems prioritised by risk and value; strangler-fig and branch-by-abstraction standard; migration reconciled with dual-running.Modernisation is portfolio-managed with measured risk, value, and progress against baselines.Modernisation is continuous across the organisation; incremental replacement is routine, reversible, and adapts as priorities shift.

Part 4. Security

TopicLevel 1 InitiateLevel 2 DevelopLevel 3 StandardiseLevel 4 ManageLevel 5 Orchestrate
Security foundations and cultureSecurity is reactive and centralised; reviews late if at all; no threat modelling; security is “someone else’s problem.”A security team defines standards; some threat modelling on major projects; basic training; security seen as a gate.Security champions embedded; threat modelling routine; secure SDLC documented; risk-based prioritisation; blameless reviews.Security metrics (threat-modelling coverage, finding-to-fix time, control adoption) are tracked against baselines.Security is genuinely everyone’s job; threat modelling is habitual; zero-trust is largely realised and the practice improves continuously org-wide.
Application securitySecurity depends on individual knowledge; no standard controls; secrets in code; stale deps; ad hoc auth.OWASP Top 10 awareness; some framework protections; secrets manager unevenly used; occasional dependency scanning.ASVS-based requirements per tier; parameterised queries; central identity with MFA; managed secrets; SBOMs and pipeline scanning.Vulnerability density, mean time to remediate, and control coverage are measured against baselines across services.Secure defaults ship in paved-road frameworks; short-lived credentials and full supply-chain assurance (SLSA) are verified continuously org-wide.
Infrastructure and cloud securityManual provisioning; broad permissions and static keys; flat networks; inconsistent encryption; no posture management.Some IAM roles and MFA; basic network tiers; encryption at rest for major stores; periodic manual reviews; partial IaC.Least-privilege RBAC/ABAC with short-lived creds; default-deny segmentation; encryption by default with KMS; CSPM with policy.Posture, drift, and policy-violation metrics are tracked against baselines; guardrail effectiveness is measured.Secure defaults ship in landing zones and IaC; micro-segmentation and preventive guardrails evolve continuously and drift is auto-remediated org-wide.
Security operationsSecurity testing manual and rare; no central logging or SIEM; no incident plan; ad hoc patching; never adversarially tested.Some scanners in the pipeline; central logging; a basic incident plan; loose patching timelines; annual pentest.Full DevSecOps scanning with risk-based gates; SIEM with some SOAR; rehearsed IR with tabletops; remediation SLAs; red teaming.MTTD and MTTR are measured against baselines; detection coverage is mapped to adversary techniques and tracked.Testing and response are highly automated; purple teaming and detection engineering improve continuously and adapt to new threats org-wide.
Privacy and data protectionPersonal data collected freely; no inventory, minimisation, or retention; consent an afterthought; no rights process.A privacy policy and basic consent exist; some retention awareness; rights requests handled manually and slowly.Privacy by design with DPIAs; data mapped and classified; retention enforced; lawful basis documented; rights met on deadline.Privacy posture is measured: data-inventory coverage, retention compliance, and rights-request turnaround against baselines.Privacy is a default engineering constraint; minimisation and automated retention are standard; rights requests are self-service and the practice adapts org-wide.
Compliance and governanceCompliance is reactive; no control framework; evidence assembled manually under deadline; frequent findings.Key frameworks identified; some documented controls; audits pass with heavy manual effort; accessibility considered late.A unified control framework cross-maps standards; evidence partly automated; accessibility tested; records and authorisations established.Control effectiveness and evidence coverage are measured continuously against baselines; findings are trended.Compliance is continuous with always-on evidence and compliance-as-code; new certifications are low-cost and the framework adapts org-wide, audit-ready at any moment.

Part 5. UI/UX Design

TopicLevel 1 InitiateLevel 2 DevelopLevel 3 StandardiseLevel 4 ManageLevel 5 Orchestrate
UX foundationsNo dedicated UX practice; decisions by opinion; research ad hoc; inconsistent flows and terminology.Some designers and occasional usability testing; personas unmaintained; UX is a phase, often bypassed.Continuous mixed-method research feeds prioritisation; shared personas, journey maps, and IA; UX quality gates in the DoD.UX metrics (task success, satisfaction, usability scores) are tracked against baselines alongside business metrics.Research is continuous and outcome-linked; controlled experiments close the loop and insights spread across teams as products evolve.
UI design and design systemsEach team builds its own UI; no shared components; inconsistent look; hard-coded colours and spacing.A partial style guide or component library exists but is optional and often out of sync between design and code.A tokenised design system with a maintained coded library, docs, and governance is used across teams; a11y built in.Design-code parity, component adoption, and drift are measured against baselines; versioning is tracked.The system is a governed product with a roadmap; it improves continuously org-wide and rebrands become token changes.
AccessibilityNo accessibility practice; issues found via complaint or lawsuit; non-semantic, untested markup.Awareness exists; some automated scanning and a pre-launch audit; a11y is a late checklist, often deprioritised.WCAG 2.2 AA is the standard; a11y built into the design system, tested, and in the DoD; teams trained with an owner.Accessibility conformance is measured in CI against WCAG baselines; defect rates and audit results are tracked.Accessibility is continuous; disabled people are involved in research; it is embedded in procurement, tokens, and CI and improves org-wide.
Content and communication designNo content practice; words written ad hoc; inconsistent terminology and tone; unhelpful errors and empty states.A style guide may exist; some plain-language awareness; content is still late-stage and per-team with little reuse.A content strategy, voice-and-tone guide, and glossary used across teams; plain language standard; shared patterns.Content is measured against outcomes (comprehension, task completion, error rates) versus baselines.Content is continuously improved from that evidence; dark patterns are prohibited and audited; patterns are localised and accessible by default org-wide.
Internationalisation and localisationSingle language; hard-coded strings; non-Unicode assumptions; new locales require code changes.Strings externalised and Unicode used, but localisation is a manual pre-launch batch; formatting and plurals inconsistent.Shared i18n architecture and locale-aware formatting; a TMS and continuous pipeline; pseudo-localisation and multi-locale CI.Localisation coverage, string freshness, and locale-defect rates are measured against baselines.i18n is enforced by tooling and lint across teams; localisation is continuous, cultural adaptation is systematic, and new locales launch fast.
Frontend engineeringAd hoc per-team frontend; heavy client code; no budgets; tested only on team devices; framework by hype.Some shared tooling and a component library; performance measured occasionally, not budgeted; limited cross-device testing.Framework and rendering chosen deliberately per surface; budgets enforced in CI with RUM; progressive enhancement standard.Performance, resilience, and reach are measured against real-user baselines and budgets; regressions fail the build.Those signals are tied to outcomes and continuously improved across surfaces as the frontend and its users evolve.

Part 6. Artificial Intelligence

TopicLevel 1 InitiateLevel 2 DevelopLevel 3 StandardiseLevel 4 ManageLevel 5 Orchestrate
AI strategy and readinessAd hoc experiments; no shared strategy; decisions driven by hype and individual enthusiasm.Problem framing on some projects; a first platform baseline; build-versus-buy discussed but inconsistent.A portfolio of use cases with clear metrics, a decision tree, readiness assessments, and lock-in/TCO analysis.Use-case value, adoption, and readiness are measured against baselines; portfolio ROI is tracked.AI strategy is integrated with business and risk planning; readiness is continuously maintained and systems are re-scoped on evidence org-wide.
MLOpsModels built ad hoc in notebooks; manual deployment; no data/model versioning; no monitoring.Some experiment tracking and a model registry; semi-automated deployment; basic monitoring for a few models.Shared platform with feature store, registry, reproducible pipelines, lineage; drift/quality monitoring; governed promotion.Model quality, drift, and business impact are measured against baselines; retraining is triggered on thresholds with gates.The lifecycle is fully automated and auditable; self-service paved roads and continuous evaluation improve models org-wide as data shifts.
Generative AI and LLM applicationsAd hoc prompting in isolated projects; no grounding, guardrails, or eval; hallucinations found in production.Some RAG and prompt versioning; basic output validation; a small manual evaluation set.Shared patterns for RAG, guardrails, and tool use; automated offline eval on every change; online metrics and human review.Offline and online eval scores, hallucination and injection rates are measured against baselines.Evaluation is tied to outcomes and improves continuously; injection defences, governed observable agents, and mitigation adapt org-wide.
AI-assisted software developmentIndividuals use assistants ad hoc; no policy; no measurement; secrets and IP at risk.Basic usage guidance and data rules; some security scanning; anecdotal productivity claims.Clear norms by risk level; mandatory review and scanning; honest outcome metrics; secure deployment and disclosure.Delivery and quality impact of assistance are measured against baselines; verification coverage is tracked.Verification is strong in the pipeline; skill development is deliberate and policy adapts continuously as tools and evidence change org-wide.
Responsible and trustworthy AINo fairness testing, explanations, or governance; responsibility undefined; issues found only after harm.Some bias testing and documentation; ad hoc oversight; framework awareness but partial adoption.Governance mapped to recognised frameworks; systematic fairness/safety/privacy testing; documented oversight and appeals; red-teaming.Fairness, safety, and privacy metrics are monitored in production against baselines and thresholds.Governance is integrated into delivery; responsibility is everyone’s job and the approach improves continuously across the organisation.
AI infrastructure and operationsAd hoc GPU allocation; no batching or caching; no cost visibility; unversioned prompts; minimal monitoring.Some shared scheduling and caching; basic cost tracking; prompts in version control; ad hoc evaluation.Shared platform with scheduling, quotas, batching, caching, right-sizing; vector infra; automated eval; cost attribution.Utilisation, cost per outcome, and latency are measured against baselines; budgets and quotas are enforced.Routing and scaling are automated, LLMOps observability is full, and utilisation and cost are continuously optimised with portability maintained org-wide.

Part 7. Data, Analytics, and Insight

TopicLevel 1 InitiateLevel 2 DevelopLevel 3 StandardiseLevel 4 ManageLevel 5 Orchestrate
Data strategy and governanceData undocumented and unowned; conflicting definitions; quality found when reports break; no catalogue or lineage.Some datasets have owners and docs; a partial catalogue; manual, reactive quality checks; policy written but weakly enforced.Critical data products have owners, contracts, SLAs; catalogue with automated lineage; continuous quality; federated governance.Data quality, contract compliance, and freshness are measured against SLAs and baselines.Data-as-product is the norm; contracts are enforced automatically, self-service guardrails adapt, and definitions are trusted enterprise-wide.
Data engineeringAd hoc scripts, manual runs, no tests or monitoring; failures found by consumers; costs unmanaged.Some orchestration and scheduling; basic transformations in version control; occasional tests; reactive firefighting.ELT with layered, tested, versioned models; orchestrated dependencies with retries/backfills; observability; costs tracked.Pipeline reliability, freshness, and cost are measured against SLAs; anomalies are detected against baselines.Pipelines are software with CI/CD, contracts, and testing; the platform improves continuously and new data products ship fast org-wide.
Analytics and business intelligenceReports built ad hoc in spreadsheets; inconsistent metrics; misleading charts; no governance.A BI tool with some shared dashboards; metric definitions still diverge; uncontrolled self-service and sprawl begins.A semantic layer defines core metrics once; certified vs experimental content; self-service within guardrails; managed lifecycle.Metric usage, freshness, and definition changes are tracked against baselines; certified content is monitored.Metrics are governed like APIs with owners and changelogs; analytics spans descriptive to prescriptive and embeds at decision points across the organisation.
Product analytics and experimentationLittle/inconsistent instrumentation; decisions by opinion; no experiments; vanity metrics; careless consent.Some events tracked but taxonomy inconsistent; occasional A/B tests without power analysis; north star proposed, not embedded.A governed, validated tracking plan; funnels/cohorts/retention routine; experiments on a shared platform; consent handled properly.Experiment volume, power, and win rates are measured against baselines; instrumentation coverage is tracked.Experimentation is the default; a shared results repository and owned instrumentation let the organisation learn cumulatively and adapt.
Decision science and data cultureDecisions by hierarchy and intuition; correlation treated as causation; uncertainty ignored; metrics surveil and are gamed.Data consulted selectively to justify decisions; some awareness of causal traps; uncertainty rarely communicated.Analyses tied to decisions with predefined criteria; correlation vs causation distinguished; uncertainty communicated; outcome focus.Decision quality and forecast calibration are tracked against outcomes and baselines.“What would change our mind?” is routine; causal rigour and honest uncertainty are norms and leaders update visibly on evidence org-wide.

Part 8. Automation

TopicLevel 1 InitiateLevel 2 DevelopLevel 3 StandardiseLevel 4 ManageLevel 5 Orchestrate
CI/CD and deliveryBuilds and deploys largely manual and inconsistent; late integration; infrequent, stressful releases; manual rollback.Automated builds and unit tests per commit; scripted but manually supervised deploys; artefacts may be rebuilt per stage.A standardised pipeline promotes one immutable artefact through envs with automated gates; canary/blue-green; auto change records.DORA metrics (lead time, deploy frequency, change-fail rate, MTTR) are tracked against baselines and gate rollbacks.Progressive delivery decouples release via flags; the pipeline self-improves and compliance evidence is automatic across the organisation.
Infrastructure as code and configurationInfrastructure provisioned manually; inconsistent, undocumented environments; slow, uncertain recovery.Some infrastructure scripted, but practices vary; inconsistent state; common drift; policy enforced by manual review.Declarative IaC standard from shared versioned modules with remote state; policy-as-code guardrails; regular drift detection.Drift, provisioning time, and policy-violation rates are measured against baselines; compliance evidence is automatic.Infrastructure is immutable, GitOps-driven, and self-healing; the module and policy library improves continuously and adapts org-wide.
Containers, orchestration, cloud-nativeContainers used ad hoc; hand-built unscanned images; manual deploy; no shared platform or isolation model.Teams containerise and use an orchestrator, but practices vary; inconsistent scanning and limits; ungoverned cost and tenancy.A standardised platform with hardened images, signing/scanning gates, namespace tenancy with quotas and network policy, cost allocation.Utilisation, density, and cost per workload are measured against baselines; FinOps optimisation is data-driven.A self-service, self-healing platform with strong multi-tenancy stays portable and hybrid/sovereign-ready and improves continuously org-wide.
Platform engineering and DevExNo platform; each team assembles its own tooling inconsistently; ticket-driven handoffs; high cognitive load.Some shared tools and templates, but fragmented and partly manual; limited self-service; DevEx unmeasured.A platform team runs golden paths, self-service provisioning, a developer portal, and scorecards; guardrails in paved roads; DevEx measured.Adoption, DevEx scores, and cognitive-load signals are measured against baselines and reviewed.A mature platform product improves continuously from that feedback; voluntary adoption is high and governance stays invisible in the workflow org-wide.
Test and process automationTesting and ops largely manual; inconsistent coverage; procedures in heads or stale docs; compliance evidence by hand.Automated tests exist but slow/flaky and run inconsistently; some operational scripts; manual remediation; periodic-review governance.Fast, parallel, reliable test infra; codified runbooks; ChatOps; auto-generated compliance evidence; governance as automated checks.Automation coverage, false-positive rates, and remediation times are measured against baselines.Routine incidents are auto-remediated with safeguards; compliance is continuous and audit-ready and humans focus on judgement org-wide.

Part 9. Operations, Reliability, and Observability

TopicLevel 1 InitiateLevel 2 DevelopLevel 3 StandardiseLevel 4 ManageLevel 5 Orchestrate
Site reliability engineeringOperations manual and reactive; no SLOs; reliability is opinion; the same incidents recur; firefighting dominates.Key services have basic SLIs/SLOs; some monitoring and alerting; toil acknowledged but unmeasured; inconsistent postmortems.Error budgets influence prioritisation; toil measured and capped; routine capacity planning; funded automation; PRR and engagement model.Error budgets, toil, and SLO attainment are measured against baselines and drive prioritisation.Error-budget policy is automated and respected; self-service ops and proactive capacity let the organisation trade velocity and stability on data and adapt.
Observability and monitoringBasic uptime checks and unstructured logs per machine; debugging means SSH; noisy, ignored alerts.Centralised metrics and log aggregation; some dashboards and threshold alerts; partial/absent traces; manual correlation.OpenTelemetry instrumentation with propagated trace IDs; structured logs, tracing, curated dashboards, SLO symptom alerting; sustainable on-call.Alert quality, MTTD, and telemetry cost are measured against baselines; burn-rate alerting is tuned to SLOs.High-cardinality, event-rich observability supports ad hoc investigation; retention is cost-optimised and telemetry informs decisions across the organisation.
Incident managementIncidents handled ad hoc by whoever notices; no roles, severities, or postmortems; informal on-call; failures recur.Basic on-call rotations and severities; some postmortems, but unclear roles and inconsistently tracked corrective actions.A formal incident command system with clear roles and criteria; blameless postmortems standard; actions tracked; on-call compensated.Incident frequency, MTTR, and on-call load are measured against baselines; recurring causes are trended.Response is rehearsed via game days; on-call stays sustainable and quiet and aggregate analysis drives structural investment as the organisation learns.
Cost, sustainability, green softwareCloud costs are a monthly surprise; no tagging, allocation, or carbon awareness; generous, unrevisited provisioning.Basic cost visibility and tagging; some reactive rightsizing and idle cleanup; sustainability acknowledged but unmeasured.A FinOps practice with attribution, budgets, forecasts, anomaly alerts, commitments, rightsizing; carbon measured for major services.Cost and carbon are measured per team against budgets and baselines; anomalies are flagged.Cost and carbon are continuous team-owned signals; efficient defaults, automated optimisation, and carbon-aware scheduling improve continuously org-wide.

Part 10. Project/Product/Programme Management

TopicLevel 1 InitiateLevel 2 DevelopLevel 3 StandardiseLevel 4 ManageLevel 5 Orchestrate
Portfolio and programme managementPriorities set ad hoc by whoever asks loudest; no portfolio view; dependencies surface as crises; annual funding scrambles.A periodically reviewed portfolio inventory; published objectives weakly linked to work; a dependency register; project-based budgeting.Strategy cascades via OKRs; a consistent prioritisation framework; cross-team planning manages dependencies; persistent team funding.Portfolio outcomes, delivery predictability, and dependency counts are measured against baselines.The portfolio is continuously rebalanced on outcome evidence; dependencies are designed away and funding cadence matches learning cadence org-wide.
Risk, audit, and assuranceRisk handled reactively after incidents; no framework or register; undocumented controls; painful manual audits.Risk registers for major systems; a control framework adopted, audits pass but manual and point-in-time; suppliers assessed at onboarding.Three-lines model and common framework org-wide; many controls automated; continuous monitoring; supplier/SBOM inventories; scheduled DR.Control effectiveness, open-risk counts, and audit findings are measured against risk appetite and baselines.Assurance is continuous and largely automated; auditors sample live evidence and supply-chain integrity is verified as risks evolve org-wide.
Procurement, open source, licensingOpen source added freely; no policy or inventory; licences unexamined; end-of-life found by accident; no owner.A basic policy and approved-licence list; some manual/late scanning; an inventory for major systems; ad hoc contribution.An OSPO owns strategy and tooling; automated licence/vuln scanning and attribution; SBOMs; clear contribution; EOL tracked.Licence compliance, dependency currency, and vulnerability exposure are measured against baselines.Open source is a managed strategic asset with fully automated compliance; upstream investment is deliberate and currency and EOL are managed continuously org-wide.
Sustaining large and long-lived systemsSystems depend on heroes; ownership by memory; undocumented knowledge; systems frozen until they break; retirements never finish.Ownership assigned and recorded for major systems; some docs and runbooks; obvious critical functions have a backup person; reactive maintenance.Team-level ownership in a catalogue surviving reorgs; bus factor measured and mitigated; decision records and runbooks; incremental modernisation.Bus factor, ownership coverage, and knowledge-transfer progress are measured against baselines.Stewardship is a funded discipline; no critical system is a single point of human failure and knowledge transfer and planned endings continue org-wide.
Ethics, accountability, public interestEthics unaddressed or reactive after scandal; accessibility ignored; opaque automated decisions with no recourse; bias untested.A code of conduct and some (late) accessibility; high-profile automated decisions get some oversight; occasional bias checks.Ethical review is part of the process; accessibility designed in and user-tested; consequential decisions carry explanation and recourse.Equity, accessibility, and algorithmic-accountability outcomes are monitored against baselines.Responsibility is embedded in how the organisation builds; equity is a non-negotiable default and algorithmic accountability is standard and continuously improved org-wide.

Overall maturity self-assessment

Use the matrices above to produce a lightweight, honest score.

Scoring rubric

  1. Score each chapter 1 to 5 using the level whose description best matches your typical reality. When behaviour is inconsistent, score the lower level.
  2. Average by part. Sum the chapter scores in a part and divide by the number of chapters. This gives a per-part maturity (e.g., “Part IV averages 2.5”).
  3. Record the minimum, not just the mean. A part that averages 3.0 but contains a Level 1 chapter carries that chapter’s risk regardless of the mean.
  4. Plot the trend. Re-score every quarter or two and watch direction of travel. A domain moving 2 → 3 is healthier than one stuck at a static 3.

A simple worksheet per part:

PartChapters scoredAverage (mean)Lowest chapterNotes / priority
I-Xcountmeanmin level…

Prioritising what to improve

Do not try to raise everything at once, and do not chase the highest average. Prioritise by risk-weighted maturity gap: attack the domains where a low level meets high consequence.

  • First: the lowest-maturity chapters in your highest-risk domains. For most organisations that means security, privacy, reliability, compliance, and any system whose failure harms people or violates law. A Level 1 here is urgent.
  • Next: foundational enablers (culture, ways of working, CI/CD, IaC, observability) that raise the ceiling for every other domain. Improving these makes later gains cheaper.
  • Later: domains that are already at Level 3 and could climb to Level 4 or 5. Only push beyond the floor where the stakes and scale justify the added investment.

The enterprise and government baseline

Enterprise and government contexts usually cannot stop at “it works.” To pass audits, sustain authorisations, and meet regulatory and public-accountability obligations, most domains must reach at least Level 3 (Standardise), the level where practices are standardised, documented, enforced across teams, and produce evidence. Level 2 typically fails audit because it is inconsistent and manually assembled under deadline; Level 1 fails outright.

Read Level 3 as the floor for anything auditable or safety-relevant, and the higher levels (4 and 5) as targets only where continuous assurance, scale, or public trust make the extra rigour worth it. Maturity remains a means: the aim is a defensible, proportionate level of control for the risk you actually carry, not a perfect score.