In this guide
Ask what sits behind the percentage
  1. 01Which changes?
  2. 02Whose definition?
  3. 03Which sample?
  4. 04Which time frame?
Without a stated denominator and outcome definition, a failure percentage tells you little about your own change.

The statement that “70% of change initiatives fail” should not be used as a universal planning fact. It compresses many kinds of change, definitions of success, and observation periods into one number whose general applicability has not been established.

Mark Hughes’s 2011 article examined five published instances of the 70% claim and found no valid, reliable empirical support for that narrative. This is a critique of a frequently repeated statistic, not a new estimate of the true failure rate. Hughes, 2011

What the critique does—and does not—show

A claim can be unsupported even when the problem it describes is real. Organizations can struggle with implementation, lose intended benefits, or create harm. Rejecting an unreliable percentage does not imply that change is easy or that most initiatives succeed.

Hughes’s review is bounded: it examines prominent published claims and questions the idea of an inherent failure rate. It is not a census of every change initiative. Its conclusion therefore should not be inflated into “research proves that change rarely fails.”

The practical distinction is between an estimate for a defined population and a sweeping statement about all organizational change. A well-designed study could estimate outcomes for a specified class of projects over a specified period. It would still need a justification before its result could be applied to different settings.

Ask what is being counted

Before using any failure rate, define the unit. Is it a software installation, a merger, a new working practice, a restructuring, or a long program containing many smaller initiatives? If one program contains ten projects, counting the program and counting each project produce different denominators.

Then define the threshold. Missing an original deadline is different from producing no benefit. A project can deliver its planned output while failing to change how people work. Another can miss its first target, adapt, and eventually create value. A binary label may conceal these differences.

Finally, ask who is included. A voluntary survey of managers who completed a transformation is not automatically representative of all organizations that attempted change. Initiatives stopped early, organizations that disappeared, and people who declined the survey may be missing. These are questions to investigate in a source, not accusations that every survey has the same flaw.

Success has more than one observer

A sponsor may judge a change by cost savings. Employees may judge workload, autonomy, fairness, or the ability to serve customers. Users may care about access and reliability. Those perspectives can disagree without any one of them being a measurement error.

Oreg, Vakola, and Armenakis reviewed quantitative research on recipients’ reactions to change. Their review distinguishes multiple dimensions of reactions and considers the change process, context, content, and perceived benefit or harm. It provides a richer basis for asking how people experience change; it does not estimate a universal initiative failure rate. Review of change recipients’ reactions

A practical implication is to identify whose outcomes matter before implementation. Use stakeholder mapping to find relevant perspectives, then agree which measures capture them. Do not average away a serious adverse effect merely because another group benefits.

A CLOSER LOOK

Before repeating a percentage, inspect its denominator.

  1. What was counted?Which initiatives, organizations and sample selection?
  2. Who judged it?Whose criteria and whose experience?
  3. What was “failure”?Which outcome, threshold and point in time?
Figure 1. Original evidence-reading prompts. Missing definitions weaken a universal headline; they do not demonstrate that organizational change is easy or risk-free.Original illustration · Innovation & Change

Time changes the answer

A launch review, a three-month adoption review, and a two-year benefits review answer different questions. Early disruption may precede improvement, but delayed benefits should not become an excuse to postpone evaluation indefinitely.

Specify when each outcome should be visible and what would justify changing that expectation. Keep the original expectation in the record when targets change. This allows reviewers to distinguish learning and adaptation from quietly redefining success.

Also distinguish a stopped experiment from a failed implementation. If a small test is designed to reject an unpromising idea, stopping may be the intended decision. If a mandatory service rollout is abandoned after causing disruption, calling it “learning” alone would omit important consequences.

These distinctions are proposed evaluation practices. They do not require knowing an aggregate failure percentage.

A bounded study can still reveal serious risk

Research on a sample of IT projects illustrates why scope matters. Flyvbjerg and Budzier analyze project cost and schedule outcomes and draw attention to extreme overruns that an average can conceal. Their subject is a defined project domain and specific performance measures. It should not be converted into a percentage for all organizational change. IT project risk study

The transferable question is whether your own reporting hides the distribution of outcomes. An average across initiatives can look acceptable while a few produce severe losses or disruptions. Conversely, a headline failure count may hide partial benefits or successful adaptations.

Read a study’s sample and outcome definitions before borrowing its headline. Relevance to your decision matters more than how memorable the number is.

Hypothetical example: one change, three judgments

A company introduces a new case-management workflow. The sponsor’s initial objective is to shorten resolution time without reducing service quality. The team also commits to manageable staff workload and accessible customer support.

At launch, the system is available on schedule. At the first review, staff use it, but duplicate entry adds work and resolution time has not improved. At a later review, a revised handoff reduces delays, although some customers still struggle with the new contact route.

Calling the entire initiative either “success” or “failure” loses useful information. A clearer account says which objectives were achieved, which remain unmet, what adverse effects occurred, and what was changed. The record can still support an overall governance decision; it simply makes the basis visible.

This example is invented to demonstrate evaluation choices. Its sequence and outcomes are not a reported case study.

A CLOSER LOOK

One change can produce different judgments.

  1. Delivery perspectiveWas the planned system or process introduced?
  2. Working perspectiveCan people use it without unreasonable burden?
  3. Outcome perspectiveDid the intended benefit appear, and for whom?

Agree measures and review dates without hiding adverse effects.

Figure 2. Original evaluation lenses for the hypothetical example above. These are questions to reconcile, not measured results or a universal definition of success.Original illustration · Innovation & Change

Replace the headline with an evaluation agreement

Before a significant initiative, write a short agreement that answers the following questions:

  • What concrete behavior, service result, or capability should change?
  • What is the baseline, and how reliable is it?
  • Which outcomes matter to affected groups as well as the sponsor?
  • When will the team review delivery, adoption, benefits, and adverse effects?
  • What evidence would justify continuing, adapting, pausing, or stopping?
  • Who makes that decision, and how will the rationale be recorded?

Separate delivery measures from outcome measures. Training attendance records delivery of training; it does not establish competent use. A system login indicates access or activity; it does not establish that the intended workflow is happening correctly.

Use several sources where the stakes warrant it: operational data, observation, user feedback, and the experience of affected staff. State uncertainties and avoid claiming causality solely because an outcome changed after launch. Other events may contribute.

What to say in a presentation

A defensible replacement is: “Change outcomes vary with the initiative, context, and definition of success. We will evaluate this change against explicit delivery, adoption, benefit, and adverse-effect measures.”

If someone supplies a percentage, ask for the original source, the population, the definition of failure, and the observation period. If those cannot be recovered, remove the number. Do not replace it with an equally unsupported success rate.

The next useful step is a concrete change management plan and a review of its assumptions. A premortem can help surface vulnerabilities before implementation, provided its imagined explanations are then investigated rather than treated as evidence.

Evidence limits

This article combines a critical review of the 70% narrative, a research review of change recipients’ reactions, and a bounded study of IT project risk. Those sources answer different questions. Together they support careful interpretation and better evaluation questions; they do not supply a substitute universal failure rate.

The evaluation agreement and hypothetical example are original practice guidance. The Hughes paper was accessed through a publicly available PDF copy; the publication is in the Journal of Change Management, 11(4), pages 451–464, DOI 10.1080/14697017.2011.630506.

Sources: