Back to Blog
Perspective

The Recurrence Problem: Why Your CAPA Is Closing Cases Without Closing Causes

Executive Summary

Recurrence is rarely a discipline failure. It is a diagnosis failure. Most quality metrics reward finishing a CAPA, not proving its cause is gone, so investigations converge on the nearest plausible explanation and the conditions that produced the defect survive the closure.

"Human error" is the industry's default root cause. Over 80% of pharma process deviations are attributed to it, while Deming estimated that 94% of problems belong to the system. The default fix, retraining, addresses roughly one human error in ten.

Recurrence hides itself. A surviving cause returns under a different name each time, passes effectiveness checks that are too short to see it, and gets filed under a new tracking number, so the evidence that would expose it is never read together.

Regulators and auditors read recurrence as a verdict. CAPA was the most-cited FDA Form 483 observation for device makers in FY2025, and problem solving is the top major nonconformity raised against IATF 16949.

You've read this investigation before. Not one like it. It. Same product, same line, same failure mode. Last quarter it arrived as a deviation; someone traced it back through the usual checkpoints, landed on the usual culprit, wrote the usual corrective action, and closed it. The CAPA record is immaculate. The signatures are all there. And here it is again on your desk, wearing a new tracking number like a disguise.

This is the recurrence problem, and it is the quiet tax on every quality organization. It's tempting to read recurrence as a discipline failure: someone got sloppy, someone skipped a step. It's almost never that. Recurrence is a diagnosis failure. A closed CAPA feels like a solved problem, but closing a case and closing a cause are two different acts, and the gap between them is exactly where defects live to fight another day.

In The Closure Bias Problem, we looked at why investigations settle on the most closable cause rather than the most accurate one. This piece picks up where that one ends: what happens when the cause survives the closure, why quality systems are structurally bad at noticing that it has, and what that costs.

The uncomfortable thesis is that most CAPAs are built to recur. Not by intent. By geometry.

The corrective-action loop: Deviation, Investigate, Correct, Close, with the only path back to the start labelled Recurrence
Figure 1: The corrective-action cycle. When the cause survives closure, the only way the loop completes is through recurrence.

The Comfort of Closure

Every quality system measures what it can count: closure time, percentage overdue, CAPAs closed this quarter. These are honest operational metrics, and they all quietly reward the same thing: finishing. Not one of them asks whether the problem is actually gone.

That produces a predictable failure mode. A typical manufacturing site logs just over a thousand deviations a year,1 and when everything becomes a CAPA the queue overflows. The only way to clear it is to triage: reserve real investigation for the catastrophes and give everything else a fast, plausible disposition.2 The deviation gets pinned to the nearest believable cause, the action is written, the box is ticked. The case closes. The cause doesn't.

Closure is a feeling. Recurrence is the fact that checks the feeling.

The Human-Error Illusion

Watch where investigations go to die and you'll find the same headstone over and over: human error. More than 80% of process deviations in pharmaceutical manufacturing are attributed to it,3 and the BioPhorum Operations Group puts the share across biopharma at roughly half.4 It is, by a wide margin, the most popular root cause in the industry. PwC also found that 80% of investigators close with a probable cause rather than a definitive one, and that probable cause is usually a person.

It's also, mostly, fiction. W. Edwards Deming, who spent a career inside the systems that produce defects, estimated that 94% of troubles belong to the system, which only management can change, leaving 6% or less to special causes, the individual worker among them.5 Everything above that line is the system wearing a person's name.

Bar chart comparing deviations attributed to human error (over 80% in pharma, about 50% in biopharma) with Deming's estimate of 6% or less for special causes
Figure 2: Who gets blamed vs. who is actually at fault. The gap between attribution and reality is the reservoir recurrence draws from. Sources: PwC Belgium (2025); BioPhorum via GEN (2025); Deming (1986).

When you write "operator error" on a deviation that was really a system problem, you haven't made a mistake of effort. You've made a mistake of category, and category mistakes recur with perfect reliability, because nothing about the system has changed.

The Retraining Trap

Name human error as the cause and the corrective action writes itself: retrain the operator. It's fast, it's cheap, it closes the record, and it rarely works. By one CDMO's analysis, training is responsible for only about 10% of the human errors that occur, because training addresses gaps in knowledge and skill, not the task design, tooling, and conditions behind the other ninety. The same analysis describes human-error investigations as "generally poor and superficial," with little sustainable corrective action beyond retraining.6

A bar split into roughly 10% fixable by training and 90% rooted in the task, tools, environment, or system
Figure 3: The corrective action that fixes one problem in ten. Source: CARBOGEN AMCIS.

This is why the same issue resurfaces after the training records are signed.7 You can't train a person out of a problem that was never theirs. The moment a CAPA reads "retrain," the recurrence is already on the calendar.

Recurrence Is a Geometry Problem

In Lattice Root Cause we argued that manufacturing failures rarely have roots. They have topologies: configurations in which several individually in-spec conditions align to make failure accessible. A linear investigation finds one node on that lattice, fixes it, and closes. The alignment survives.

That explains why defects come back. It doesn't explain why organizations keep failing to notice that they have. Three properties of a failure topology turn a surviving cause into an invisible one.

The defect changes its name. Each time the conditions realign, the investigation lands on whichever node is most visible that day. In March it's the operator who loaded the part. In June it's a worn tool. In August it's a supplier lot. Each conclusion is partly true, because each is a real node on the lattice, and each fix helps for a while. But the CAPA log now shows three unrelated problems with three closed causes, and nobody has a reason to read them together. It isn't three causes. It's one shape, seen from three angles.

The effectiveness check can't see it. Most CAPA systems verify a fix by watching for recurrence over a fixed window, commonly 30 to 90 days. A topology that aligns only occasionally will pass that check most of the time, whether or not the fix did anything. If a failure mode surfaces on average once every four months, a 30-day check finds no recurrence about 78% of the time even when the corrective action changed nothing. A 90-day check passes about 47% of the time: a coin flip recorded as proof. If one of the aligning conditions is seasonal, such as humidity, a CAPA closed in October can pass any check that ends before summer.

Bar chart: probability that a no-recurrence effectiveness check passes even though the fix changed nothing, for a failure mode recurring every 120 days on average: 78% at 30 days, 61% at 60, 47% at 90, 22% at 180
Figure 4: How often an ineffective fix passes its effectiveness check. Modeled as a Poisson process with a mean recurrence interval of 120 days; the probability of zero recurrences in a window of t days is e^(-t/120).

Every recurrence is filed as a stranger. A new occurrence gets a new tracking number and an investigation that starts from zero. That discards the most valuable evidence a quality system ever receives. A single occurrence is consistent with dozens of explanations. Three occurrences of the same failure mode constrain them sharply, because whatever caused the failure had to be present all three times. The conditions shared across occurrences are where the topology shows itself, and a CAPA system that treats each case as independent never looks at them together.

Illustrative Scenario: The Weld That Kept Changing Its Name

A battery-module line sees intermittent failures on an ultrasonic tab-to-busbar weld. The weld passes at the station but shows elevated resistance after thermal cycling at end-of-line test, on roughly one module in two hundred for a stretch, then nothing for weeks. Over one year it produces three investigations.

March, CAPA-0391. The weld monitor logged "in spec," so the investigation settles on the operator, who must have loaded a tab slightly proud. Corrective action: retrain. The line goes quiet for six weeks and the effectiveness check passes.

June, CAPA-1488. Failures cluster on Station 3, whose horn knurl is worn. The horn is replaced and the replacement interval shortened. The line goes quiet again; the check passes.

August, CAPA-2207. Failed tabs show heavy surface oxide, so the incoming aluminum lot is quarantined and the supplier is issued a corrective action request. Closed.

Read together. Pooled, the three records share conditions that none of them named as a cause. Every failing module was welded on second shift, from tabs that had sat more than a day in an unconditioned staging area, at weld energy in the lower third of the validated window. Staging humidity rose through the summer, and oxide thickened with dwell time. At the low end of the energy window, the horn could break through that oxide only while its knurl was sharp. That is why the new horn appeared to work, and why the quarantined lot looked guilty: the oxidized tabs were simply the ones that had waited longest. Each CAPA touched a real node. None touched the arrangement. The failure stopped when staging dwell time was capped in conditioned storage and the energy setpoint was re-centered, and the team started tracking the failure mode across tracking numbers instead of one case at a time.

None of this is solved by more diligent engineers. It is solved by a different unit of analysis: the failure mode, not the tracking number.

Someone Is Always Reading Your Recurrences

If the operational cost were the whole story, recurrence would still be worth solving. It isn't the whole story. In any industry where a defect can hurt someone or trigger a recall, your pattern of recurrence is the evidence regulators and customers use to decide whether your quality system actually works. It carries a different name in every sector. The diagnosis is identical.

Three statistics: CAPA is the most-cited FDA Form 483 device observation; problem solving is the top major IATF 16949 nonconformity; 27.7 million vehicles affected by U.S. recalls in 2024
Figure 5: Three sectors, one verdict. Sources: FDA FY2025 device observations; IAOB via Amtivo (2026); NHTSA 2024 Annual Recalls Report.

In medical devices, the reader is the FDA. Inadequate CAPA procedures were the single most-cited device observation on Form 483s in FY2025, 279 citations and 10.5% of the total,8 and CAPA deficiencies are often the tipping point that turns an inspection into a Warning Letter.9 In pharmaceuticals, failure to thoroughly investigate discrepancies (21 CFR 211.192) was the second most-cited drug GMP observation in FY2024.10 Our own analysis of 417 FDA Warning Letters found that 48% cited inadequate root cause investigation.11

In automotive, the reader is the IATF 16949 auditor. Problem solving is the most common major nonconformity in IAOB data, typically because problems were closed without establishing the root cause or verifying that the fix worked.12 In aerospace, AS9100 auditors treat corrective action as one of the most-scrutinized processes and expect a root cause that explains recurrence, not just a description of what went wrong.13 Different acronyms. One verdict.

And the audience is never only the regulator. It's the OEM customer who can put you on new-business hold, the registrar who can suspend your certificate, the public that reads the recall notice. U.S. recalls touched 27.7 million vehicles in 2024 alone.14 A recurrence isn't just a defect you pay for twice. It's a sentence in the story all of them are writing about your operation, and that story compounds faster than the cost does.

The Compounding Cost

Recurrence compounds. Every repeat means you pay the failure cost again (scrap, rework, holds, lost batches) and the investigation cost again, while the underlying topology keeps quietly issuing invoices. ASQ estimates that true quality-related costs run as high as 15 to 20% of sales revenue at many organizations, and that the cost of poor quality in even a thriving company is about 10 to 15% of operations.15 Ford set aside about $1,203 per vehicle for warranty repairs in 2023, roughly double its 2019 figure of $591, and spent $4.8 billion that year fixing customers' cars.16 Some share of figures like these is the same defects, billed repeatedly.

And the deviation you dispositioned as "operator error" in January can become the field failure in April: same topology, larger blast radius. Recurrence doesn't stay small. It waits, and it grows.

Related Research

The Investigation Gap: Our analysis of 417 FDA Warning Letters found that 48% cited inadequate root cause investigation. Read more →

Closing the Cause, Not the Case

Breaking the loop doesn't demand more discipline from people who are already working hard. It demands a different unit of work. Instead of finding a plausible cause and stopping, investigate the full failure space: every contributing condition, in spec or not, and the way they align. Treat "human error" as a hypothesis to be tested, never a conclusion to be reached. Test each candidate against the evidence rather than the calendar. Exonerate the operator when the system is at fault, and link every conclusion to the data that supports it, so the record holds up to the auditor and to the next engineer alike. Then measure success the way the defect does: track recurrence by failure mode across tracking numbers, pool the occurrences before concluding anything, and size the effectiveness check to the failure's real recurrence interval rather than to the calendar.

This is the work Lattice is built for. The Conductor, our reasoning engine, investigates the topology rather than the nearest node, tests the full hypothesis space instead of triaging it away, and produces a traceable record that an auditor, or your own team six months from now, can trust. The deeper argument lives in Lattice Root Cause; the practical payoff is simpler. When you close the cause instead of the case, the defect doesn't come back wearing a new tracking number.

A CAPA should close a cause. Most close a case. The distance between the two is the recurrence problem, and it is a solvable one.


References

  1. Bioprocess Online, "Human Error" Deviations: How You Can Stop Creating (Most Of) Them.
  2. SG Systems Global, What Is CAPA? Corrective & Preventive Action in a Modern QMS.
  3. PwC Belgium (2025), Reducing human error in the pharma quality environment.
  4. Genetic Engineering & Biotechnology News (2025), Blaming Just Human Error Can Mask Deeper Bioprocess Problems, citing the BioPhorum Operations Group.
  5. Deming, W. E. (1986). Out of the Crisis. MIT Press.
  6. CARBOGEN AMCIS, Investigating Human Error in Pharmaceutical Manufacturing.
  7. CAI (2025), Human Factors in Pharma Manufacturing: Beyond Error Prevention.
  8. U.S. FDA, FY2025 inspection observations (devices), via MedDeviceGuide, FDA Medical Device Warning Letter Trends 2024–2026.
  9. Hogan Lovells (2025), FDA medical device inspections in 2025.
  10. Regulatory Affairs Professionals Society (2026), FDA official details top GMP violations cited in inspections.
  11. Auxiliary Machines Research (2026), The Investigation Gap: Why 48% of FDA Warning Letters Cite the Same Problem.
  12. Amtivo (2026), Top 5 Major IATF 16949 Nonconformities, citing IAOB data.
  13. SoftwareConnect, The AS9100 Guide for Aerospace Quality Management.
  14. National Highway Traffic Safety Administration (2025), 2024 Annual Recalls Report.
  15. American Society for Quality, Cost of Quality (COQ).
  16. BizzyCar (2024), Automotive Recall Alert, reporting comments by Ford CEO Jim Farley.

Cite this Article

Karam, S. (2026). The Recurrence Problem: Why Your CAPA Is Closing Cases Without Closing Causes. Auxiliary Machines. https://www.runlattice.com/blog/the-recurrence-problem

Start with an investigation your team can explore.

Begin with a tailored dataset and engineer's guide, or bring a closed case your team knows well. See how Lattice builds the evidence trail before deciding whether to use your own live data.