3.7

View in English

3.7 Software maintenance

Overview and motivation

Most software spends the overwhelming majority of its life not being built, but being maintained. The moment a system goes into production, it enters a phase (often lasting years or decades) of fixing defects, adapting to a changing environment, improving what already works, and preventing future trouble. In large enterprises, and especially in government, this phase dominates. Tax engines, benefits systems, defence platforms, and core financial ledgers are routinely maintained far longer than anyone who commissioned them expected. Software maintenance is the discipline of keeping delivered software correct, current, and valuable over its whole operational life.

Maintenance is chronically underestimated and undervalued, and that mistake is expensive. Study after study, across decades, places maintenance at well over half of total lifetime software cost, commonly cited in the range of 60 to 90 percent for long-lived systems. Yet organisations plan, budget, staff, and celebrate the initial build as if it were the whole endeavour. Then they treat everything afterward as an afterthought, funded from a shrinking pot and assigned to whoever is available. The result is predictable: brittle systems, demoralised maintainers, escalating cost of change, and eventually a crisis that gets framed as a “legacy problem” (chapter 3.6) when it was really an unmanaged-maintenance problem all along.

This chapter follows the SWEBOK (Software Engineering Body of Knowledge) Software Maintenance knowledge area and ISO/IEC 14764. It covers maintenance fundamentals and the four recognised categories; the key issues that make maintenance hard, including cost, staffing, and morale; the maintenance process; core techniques of program comprehension, reengineering, and refactoring; how to estimate maintenance cost; and (the highest-leverage idea in the chapter) how to design for maintainability from the start. The central conviction is that maintenance is not a lesser activity that follows engineering. It is the largest part of software engineering, and it must be planned for, resourced, and respected as such.

Key principles

  • Maintenance is the majority of the lifecycle, not an epilogue. Plan and budget for it from day one; it will cost more than the build.
  • The four categories are different work. Corrective, adaptive, perfective, and preventive maintenance have different drivers and cadences; most effort is not bug-fixing.
  • You cannot change what you do not understand. Program comprehension is the largest single activity in maintenance; make code and its history legible.
  • Maintainability is a design property. The cost of future change is set largely by decisions made during construction; design for it deliberately.
  • Small, safe, continuous change beats deferred big change. Refactor and modernise incrementally under a test safety net rather than accumulating a change debt.
  • Software ages even when it sits still. The environment moves (dependencies, platforms, regulations), so a static system silently rots; preventive maintenance is real work.
  • Maintainers deserve first-class status. Morale, knowledge retention, and staffing of maintenance teams directly determine long-term cost and risk.

Recommendations

Distinguish the four categories of maintenance and staff for all of them

ISO/IEC 14764 and SWEBOK recognise four categories, and confusing them is a common planning error. Corrective maintenance fixes defects found in operation. Adaptive maintenance keeps the software working as its environment changes: new operating systems, browsers, dependencies, hardware, regulations, or interfacing systems. Perfective maintenance improves the software for users and maintainers, through new features, better performance, improved usability, and enhanced maintainability. Preventive maintenance corrects latent faults and reduces future risk before it manifests, through hardening, cleanup, and modernisation of fragile areas. A useful further split groups corrective and preventive as correction (dealing with faults) and adaptive and perfective as enhancement (dealing with new requirements). Crucially, empirical studies consistently find that the majority of maintenance is not corrective; enhancement and adaptation dominate. Budget and staff accordingly, and track which category your effort actually falls into so you can manage it.

Invest in program comprehension

The single largest activity in maintenance is understanding the existing system well enough to change it safely. Maintainers routinely spend more time reading and reasoning about code than modifying it. Make this cheaper on purpose. Keep documentation close to the code and current (chapter 2.7). Preserve decision history through architecture decision records (chapter 1.6) and clean commit history (chapter 2.6). Use static analysis, dependency graphs, and code navigation tooling to map unfamiliar territory. Characterisation tests (tests that pin down current behaviour, including quirks) turn tacit understanding into executable, durable knowledge. When comprehension is expensive, every change is slow and risky. When it is cheap, maintenance becomes routine.

Refactor continuously under a test safety net

Refactoring is disciplined restructuring of code that improves its internal quality without changing its external behaviour. Done continuously and in small steps, it counteracts the natural drift toward complexity and keeps the cost of change flat rather than rising. The non-negotiable precondition is a reliable automated test suite (chapter 2.4). Without it, “refactoring” is just risky rewriting. Fold refactoring into everyday work: leave each module a little cleaner than you found it, rather than saving it for rare, large, dangerous cleanups. This is preventive maintenance in practice, and it is the cheapest maintenance there is.

Reengineer when incremental change is no longer enough

When a component has degraded to where routine change is too costly or risky, reengineering (examining and altering a system to reconstitute it in a new form) is the heavier tool. Reengineering typically combines reverse engineering (recovering design and intent from the implementation) with forward reengineering (rebuilding to a better structure while preserving behaviour). Prefer to reengineer in bounded, incremental slices using patterns such as strangler fig and branch-by-abstraction (chapter 3.6), rather than as a wholesale rewrite. Reengineering sits on the maintenance-to-modernisation continuum: refactoring for the small and local, reengineering for the structural, and modernisation for the platform-level.

Run a defined maintenance process

Maintenance benefits from an explicit, repeatable process, as described in ISO/IEC 14764: process implementation (establishing plans and procedures), problem and modification analysis (triage, reproduce, assess impact and cost), modification implementation, maintenance review and acceptance, migration, and retirement. Wrap it in disciplined change management: every maintenance request (whether a defect report or an enhancement) should be logged, classified by category, assessed for impact, prioritised, implemented under version control with tests, reviewed, and released through the normal pipeline (chapter 11.2). Impact analysis, understanding everything a proposed change might touch, is central and deserves real effort. Retirement is part of the process too: decommissioning a system safely, migrating its data and users, and preserving records is maintenance work that must be planned, not improvised.

Estimate maintenance cost explicitly and fund it

Do not treat maintenance as free, or as noise in the build budget. Estimate it. Common approaches include maintenance-effort ratios (the widely used rule of thumb that annual maintenance runs roughly 15 to 25 percent of the original development cost, though long-lived critical systems accumulate far more over their life), parametric models such as COCOMO II (the Constructive Cost Model) with its maintenance and reuse extensions, and metrics-driven forecasting from your own historical data on defect rates, change volume, and change cost. Feed these estimates into total-cost-of-ownership analysis and the economics discussed in chapter 10.10. A system’s purchase price or build cost is a down payment. The mortgage is maintenance, and it should appear in every business case.

Design for maintainability from the start

The most leverage over maintenance cost is exercised before maintenance begins. Maintainability (analysability, modifiability, testability, and modularity, in the vocabulary of ISO/IEC 25010) is a design quality that must be an explicit requirement, not a happy accident. Favour modular, loosely coupled, high-cohesion designs (chapter 2.2); clear interfaces and separation of concerns; strong automated tests; readable code and current documentation; and rich observability so that operators and maintainers can see what the system is doing (Part 9). Every one of these decisions trades a little more effort now for large, compounding savings across the decades a system will actually live. Building for maintainability is the highest-return investment in the entire lifecycle.

Trade-offs: pros and cons

ApproachProsCons
Continuous refactoring / preventive maintenanceKeeps cost of change flat, reduces risk, high ROIOngoing effort with no visible new features; needs strong tests
Deferring maintenance (“keep the lights on”)Cheapest this quarter; frees capacity for featuresChange debt compounds; eventual crisis and forced expensive action
Reengineering a degraded componentRestores maintainability and extends useful lifeSignificant effort and risk; behaviour must be preserved carefully
Designing for maintainability up frontCompounding lifetime savings; easier every future changeHigher initial cost and discipline; benefits are deferred and less visible

The recurring trade-off in maintenance is present cost versus future cost, and the temptation always runs toward deferral. Skipping refactoring, letting dependencies age, and starving the maintenance team all look free this quarter, because the bill arrives later: as a slower, riskier, more expensive system, and eventually as a “legacy crisis.” The discipline of good maintenance is to pay small, continuous, visible costs now to avoid large, sudden, career-defining costs later. Because the savings are deferred and invisible, this trade requires leadership that understands lifecycle economics, not just launch dates.

Questions to discuss with your team

  1. Who owns the maintenance number in your budget, and is it a first-class line or a residual scraped from whatever the build did not spend? Maintenance is the majority of lifetime cost, commonly 60 to 90 percent for long-lived systems, yet it is routinely funded as an afterthought and staffed with whoever is free. When the budget is residual, preventive work is the first thing cut, change debt compounds, and a predictable slide into a “legacy crisis” follows. Bring an actual estimate (a maintenance-effort ratio, a parametric model, or your own historical change-cost data) and name the person accountable for funding it across the system’s life. The fix is to budget maintenance explicitly in every business case, the way a mortgage sits beside a purchase price. Leadership that only celebrates launches will keep underfunding the phase where most of the money and risk actually live.

  2. Do you ring-fence capacity for preventive maintenance, or does it always lose to the next feature? Preventive work (refactoring under a test net, keeping dependencies current, hardening fragile areas) is the cheapest maintenance there is, because it keeps the cost-of-change curve flat instead of letting it climb. It is also the easiest to defer, since skipping it looks free this quarter and the bill arrives later as a slower, riskier system. A concrete mechanism helps: a standing allocation, and many strong teams protect roughly a fifth of capacity, that is guarded rather than negotiated away every sprint. Bring your cost-of-change trend as evidence; if it is rising, you are already under-investing. The discipline is paying small, visible costs now to avoid large, sudden, career-defining ones later, and that requires leadership that reads lifecycle economics rather than launch dates.

  3. What is your plan to retire a system, and when did you last actually decommission one? Retirement is an explicit part of the maintenance process (data migration, user cutover, records preservation, safe shutdown), yet organisations carry dead and redundant systems for years because decommissioning is unglamorous and unbudgeted. Every zombie system still consumes licences, security patching, integration surface, and the attention of people who could be elsewhere. Bring an inventory and flag systems with no active users or a full replacement already live, then plan their shutdown like any other work: migrate data, preserve what statute requires, and confirm nothing still depends on them. In government especially, records-retention law shapes how you retire, so involve compliance early. A maturity signal worth tracking: when did your organisation last deliberately turn something off?

  4. How much of every change is spent understanding the system before touching it, and what is your bus factor on the systems that matter most? Program comprehension is the single largest activity in maintenance, and its cost is set by how legible you have kept the code, its history, and its behaviour. When understanding lives only in a few long-tenured heads, each departure or retirement raises the price of every future change, and a single absence can stall a critical fix. Bring evidence: the ratio of reading-and-reasoning time to editing time on recent changes, the number of people who can safely modify each core module, and whether business rules and decisions are documented alongside the code or reconstructed from memory each time. The competing consideration is that documentation and characterisation tests cost effort now for savings that only show up later, so they are easy to skip. In enterprise and government, where systems outlive their original authors by decades and statutory rules are buried in a calculation engine nobody fully remembers, treat captured comprehension (architecture decision records, characterisation tests, current documentation) as an asset you fund deliberately, not a courtesy that happens when someone has spare time.

  5. Who actually staffs your maintenance work, and does its status and morale match its importance? Maintenance is the majority of lifetime cost and the hardest engineering there is, safely changing systems you did not build and may not fully understand, yet it is routinely handed to the least experienced people and framed as low-status “keeping the lights on.” That signal is corrosive: your best engineers avoid the work, knowledge concentrates and then walks out the door, and the cost of change climbs while nobody is watching. Bring the seniority profile of who maintains your longest-lived systems, your attrition and knowledge-retention data, and an honest read on whether maintenance is a career dead-end or a respected specialism in your organisation. The tension is real, since ambitious engineers want to build new things and leaders want to celebrate launches, so respecting maintenance takes deliberate structure. For a large enterprise or public body running systems that carry regulatory and financial risk for decades, staffing maintenance with respected senior engineers is a risk-management decision, and letting maintenance become a punishment posting is how you manufacture the next legacy crisis.

  6. Do you track which of the four categories your effort actually falls into, and are you measuring the cost of change as a leading indicator? Teams routinely plan for maintenance as if it were mostly bug-fixing, when empirical studies show enhancement and adaptation dominate, so a portfolio funded only for corrective work is mis-scoped from the start. Without category tracking you cannot see that a system is being reshaped by a steady stream of regulatory adaptations, and without a cost-of-change metric (change lead time, change failure rate, complexity trends) you cannot tell whether your curve is flat or quietly climbing toward a crisis. Bring your actual category breakdown for the last year, your cost-of-change trend if you have one, and an honest note on whether impact analysis is a real step or a formality skipped under deadline pressure. The competing pull is that measurement itself takes effort and can feel like overhead when the system still works. In enterprise and government portfolios, where many teams maintain many systems and a rising cost curve on any one of them is an early warning worth acting on, shared category tracking and cost-of-change indicators are what let leadership reengineer a module before it degrades rather than after it fails in public.

Sector lens

Startup. With a handful of engineers and little runway, you cannot afford a heavyweight maintenance process, but you also cannot afford a codebase nobody will touch. Carve out a small standing slice of each cycle (roughly a day in five) for preventive work: patch dependencies, clear small defects before they compound, and refactor the corners you already dread under whatever tests you have. The goal is to keep the code cheap to change while you pivot, so you never wake up at twenty engineers mistaking deferred maintenance for a “legacy problem.”

Small business. With no dedicated maintenance specialist and a tight budget, lean toward buying and hosting over building, so that adaptive maintenance (security patches, platform and dependency updates) is largely someone else’s job. Where you do own code, keep it small, boring, and well-documented, and make sure at least two people understand anything the business depends on. Track the handful of systems you cannot afford to lose, and budget a modest, explicit line for keeping them current rather than pretending maintenance is free.

Enterprise. At scale you are maintaining many long-lived systems across many teams, so the priority is a defined, repeatable process: a logged and triaged request pipeline, classification into the four categories, routine impact analysis, and a standing preventive allocation guarded rather than negotiated away. Fund maintenance as a first-class programme, measure cost-of-change indicators across the portfolio, and use a rising curve as the trigger to reengineer a module before it becomes a liability. Governance and audit expectations mean category tracking and change records are not overhead, they are the evidence that the estate is under control.

Government. Procurement rules, transparency, and public accountability shape maintenance as much as engineering does. Adaptive maintenance driven by statute arrives on hard annual deadlines that cannot slip, so budget maintenance as an indefinite operating cost and staff a stable expert team to retain knowledge of rules whose authors have long retired. Retirement is constrained by records-retention law, so plan decommissioning with compliance from the start, and favour contracts and architectures that keep the system maintainable and portable rather than trapping you with a single vendor for decades.

Examples

Startup. A startup that just shipped its MVP is tempted to pour every hour into new features, but its founding engineer carves out a standing slice of each sprint (roughly one day in five) for maintenance from the very first month. That budget keeps dependencies patched, clears small defects before they compound, and refactors the corners the team already dreads, so the codebase stays cheap to change while the product pivots. The startups that skip this reach twenty engineers with a codebase nobody wants to touch and mistake it for a “legacy problem” when it was deferred maintenance all along.

Enterprise. A global bank runs a payments platform that has been in production for fifteen years. It funds maintenance as a permanent, first-class programme rather than a residual budget line. Work is triaged into the four categories: a steady stream of adaptive changes tracks new regulations and partner-bank interface updates, perfective work adds features and improves throughput, corrective work clears defects against strict service-level agreements (SLAs), and a standing preventive allocation (roughly a fifth of team capacity) pays down complexity through continuous refactoring under a comprehensive test suite. The team measures change lead time and change failure rate, and treats a rising cost-of-change as an early warning to reengineer a module before it becomes a liability. Maintainers are senior, well-regarded engineers, not junior staff parked on “keeping the lights on.”

Government. A national tax authority maintains a system that has run for over thirty years and is amended every year as tax law changes. The dominant category here is adaptive maintenance driven by statute, with hard annual deadlines that cannot slip. The authority invests heavily in program comprehension: business rules are documented alongside the code, characterisation tests pin the behaviour of rules whose original authors have long retired, and impact analysis is a formal step before any change to the calculation engine. Because the environment (the law) changes continuously, the system can never be “finished,” so the authority budgets maintenance as an indefinite operating cost, staffs a stable expert team to retain knowledge, and modernises the surrounding delivery practices such as source control, continuous integration (CI), and automated testing, even while the core endures.

Business case: motivations, ROI, and TCO

The core business fact of software is that maintenance, not construction, is where the money goes. Across the industry and across decades of study, maintenance accounts for the clear majority of lifetime cost, frequently cited at 60 to 90 percent for systems that live a long time, which in enterprise and government is most of them. Any total-cost-of-ownership analysis that stops at go-live is off by a factor of several. The primary business case for taking maintenance seriously is simply accuracy: budget for the whole life of the system, or be repeatedly surprised by the bill.

The return on investment comes from bending the cost curve. In a neglected system, the cost of each change rises over time as complexity accumulates and comprehension decays, until change becomes prohibitively slow and risky. In a well-maintained system, continuous preventive work (refactoring, dependency currency, test coverage, documentation) keeps that curve flat, so the thousandth change costs about what the tenth did. Investing in maintainability and preventive maintenance is therefore not an expense to minimise. It is the lever that determines whether a system stays affordable to change or drifts into the escalating cost and risk of a legacy estate (chapter 3.6) and the sustaining challenges of chapter 10.4. Fund maintenance deliberately, measure the cost of change as a leading indicator, and treat a rising curve as a signal to act, not as a fact of nature. The economics are covered further in chapter 10.10.

Anti-patterns and pitfalls

  • Treating maintenance as an afterthought. Budgeting and celebrating only the build, then starving the far larger and longer maintenance phase.
  • Staffing maintenance with the least experienced people. Assigning the hardest work (safely changing systems you do not fully understand) to those least equipped, and signalling that maintenance is low-status.
  • Confusing maintenance with bug-fixing. Planning only for corrective work when adaptation and enhancement actually dominate the effort.
  • Deferring preventive maintenance indefinitely. Never refactoring, never updating dependencies, until change debt forces an expensive crisis.
  • Changing code without impact analysis. Making a “small fix” that ripples into unforeseen failures elsewhere.
  • Refactoring without a test safety net. Restructuring code with no way to prove behaviour was preserved: that is just risky rewriting.
  • Letting knowledge walk out the door. Failing to document business rules and decisions, so each retirement or departure raises the cost of every future change.
  • Never retiring anything. Carrying dead and redundant systems forever because decommissioning is unglamorous and unplanned.

Maturity model

  • Level 1: Initiate. Maintenance is unplanned and unfunded, handled reactively by whoever is free. It is seen as bug-fixing and as low-status work. There is no category tracking, no cost estimate, and knowledge lives in a few heads. Cost of change rises unnoticed until a fix stalls or a crisis forces attention.
  • Level 2: Develop. Some teams have begun logging and triaging maintenance requests and carrying a budget line, but practice is inconsistent across the organisation and the budget is usually a residual. Corrective work is tracked while adaptive and perfective effort is not clearly distinguished. Some tests and documentation exist in pockets, so change is partly controlled but comprehension remains expensive and uneven from team to team.
  • Level 3: Standardise. A defined maintenance process (per ISO/IEC 14764) is documented and enforced org-wide: work is classified into the four categories, impact analysis and change management are routine, and maintenance is estimated and funded explicitly in every business case. Preventive maintenance and refactoring are standard practice under a solid test suite, and maintainability (analysability, modifiability, testability, modularity) is an explicit design requirement rather than a local habit.
  • Level 4: Manage. Maintenance is measured and controlled with data against baselines. Cost-of-change indicators (change lead time, change failure rate, complexity and defect trends) are tracked per system, the four-category effort mix is quantified against expectations, and maintenance-effort ratios and parametric estimates are checked against actual historical change cost. A rising cost curve is detected as a leading indicator and triggers action, and preventive allocation is sized from evidence rather than guessed. Decisions to refactor, reengineer, or retire are made on measured thresholds, not intuition.
  • Level 5: Orchestrate. Maintenance is continuously improved and integrated across the organisation and its lifecycle economics. Lifecycle TCO drives portfolio investment, reengineering is applied deliberately before components degrade, knowledge is actively retained, and retirement is planned and routinely executed. The organisation rebalances maintenance effort as the environment shifts (regulations, platforms, dependencies), maintainers are respected senior engineers, and the whole estate adapts so the cost of change stays flat across systems that live for decades.

Ideas for discussion

  1. What fraction of your engineering effort actually goes to maintenance, and does your budget and staffing reflect that reality?
  2. Can you break your maintenance work down into the four categories, and does the mix match your assumptions?
  3. How much of a typical change is spent understanding the system versus modifying it, and what would make comprehension cheaper?
  4. Is your cost of change rising, flat, or falling over time, and are you measuring it at all?
  5. Do your teams have a reliable test safety net that makes continuous refactoring safe, or is restructuring too risky to attempt?
  6. Who maintains your longest-lived systems, how is their knowledge captured, and what is the status and morale of that work?

Key takeaways

  • Maintenance is the majority of software’s lifetime cost (commonly 60 to 90 percent for long-lived systems) and must be planned, budgeted, and staffed as a first-class activity.
  • The four categories (corrective, adaptive, perfective, preventive) are distinct work, and enhancement and adaptation, not bug-fixing, usually dominate.
  • Program comprehension is the largest single maintenance activity; make code, history, and behaviour legible to keep every change cheap.
  • Refactor continuously under a test safety net and reengineer degraded components incrementally to keep the cost of change flat.
  • Estimate maintenance cost explicitly and feed it into total-cost-of-ownership and economic decisions.
  • Design for maintainability from the start (it is the highest-return investment in the whole lifecycle) and treat maintainers as the senior professionals they need to be.

References and further reading

  • IEEE Computer Society, SWEBOK Guide (Software Engineering Body of Knowledge), Software Maintenance knowledge area
  • ISO/IEC 14764 / IEEE 14764, Software Engineering: Software Life Cycle Processes, Maintenance
  • ISO/IEC 25010, Systems and software Quality Requirements and Evaluation (SQuaRE): maintainability quality characteristics
  • Martin Fowler, Refactoring: Improving the Design of Existing Code
  • Michael Feathers, Working Effectively with Legacy Code
  • Thomas M. Pigoski, Practical Software Maintenance
  • Penny Grubb and Armstrong A. Takang, Software Maintenance: Concepts and Practice
  • Barry Boehm et al., Software Cost Estimation with COCOMO II (maintenance and reuse models)
  • Meir M. Lehman, “Laws of Software Evolution” (on why software must continually change or become less useful)
  • Robert C. Seacord, Daniel Plakosh, and Grace A. Lewis, Modernising Legacy Systems