How to Prioritize Failure Modes When the Fix Is Out of Reach

by , | Cartoons

Every failure mode and effects analysis eventually reaches an uncomfortable point: the team has identified an important failure mode, evaluated its effects and causes, and realized that the best risk-reduction action is not within engineering’s authority. The technical fix may require capital, additional staffing, a supplier change, or a management decision. Learning how to prioritize failure modes gets harder when the analysis identifies a risk the team can describe but cannot close on its own.

How Risk Priority Numbers Rank Failure Modes

In many traditional FMEA systems, teams assign ratings for severity, occurrence, and detection, often on 1-to-10 scales, and multiply them to calculate a risk priority number, or RPN. Using 1-to-10 scales, RPN = Severity × Occurrence × Detection, producing a value from 1 to 1,000. The method is still widely recognized, but it is not the only way to prioritize FMEA findings and it should not be treated as a precise measurement of risk.

The three ratings describe different parts of the risk picture. Severity reflects the consequence of the failure effect. Occurrence reflects how likely or frequent the relevant cause or failure is under the scoring method being used. Detection reflects the ability of current controls to detect the cause or failure before the defined effect reaches the next point of concern. In traditional scales, a higher detection rating generally means weaker detection capability.

RPN is a screening and prioritization aid. It is not a probability, and a larger number does not automatically mean that one failure mode deserves action before every failure mode below it.

Severity, Occurrence, and Detection in Practice

The ratings only become useful when the team applies a defined scoring guide consistently. Severity should be tied to the effects that matter in the analysis, which may include safety, regulatory, environmental, customer, production, or asset consequences. Occurrence should follow the method’s defined frequency or likelihood criteria and use available evidence where possible. Detection should be based on the capability of current prevention or detection controls, not on confidence that someone will probably notice a problem.

That discipline matters because the numbers are ordinal ratings, not measurements on a continuous physical scale. Rather than debating whether something “feels like” a seven or an eight, the team should map each rating to the criteria in the FMEA method it has adopted and document the evidence behind the choice. Consistent scoring makes later reviews more meaningful and helps prevent the ranking from becoming a contest of opinions.

How to Prioritize Failure Modes With the RPN Framework

RPN can provide a useful common view when an organization has chosen that method, but sorting the worksheet from highest to lowest and starting at the top is too simplistic. Two failure modes can have the same RPN and very different consequences, controls, and action needs. Teams should look at the full severity-occurrence-detection profile, the defined scoring criteria, and any method-specific prioritization rules before deciding what to address first.

This distinction is especially important because FMEA practice is not uniform across industries. IEC 60812:2018 recognizes prioritization of failure modes and includes alternative approaches to RPN, including a criticality-matrix method. In automotive FMEA, the AIAG & VDA harmonized handbook introduced Action Priority tables to replace RPN as the action-prioritization method. A plant should therefore follow the FMEA method, customer requirement, or internal standard that actually governs its analysis rather than treating RPN as universal.

Weigh Consequence, Not Just the Raw Number

The arithmetic itself shows one of RPN’s limitations. A 5-5-8 rating and a 10-10-2 rating both produce an RPN of 200, but they describe very different profiles. The first combines moderate severity and occurrence with weak detection. The second combines the highest severity and occurrence ratings with comparatively strong detection. Equal products do not make the underlying risks equivalent.

For that reason, teams should not use a single RPN cutoff as the only trigger for action. High-severity effects deserve explicit review under the organization’s chosen FMEA method, particularly where safety, regulatory, environmental, or other critical consequences are involved. If the organization uses escalation rules or severity thresholds, those rules should be defined in its procedure rather than invented during the meeting.

Two failure modes with the same RPN can demand very different responses. The number can help focus the discussion, but it does not replace the analysis behind it.

When Detection Scores Expose a Blind Spot

A high detection rating is useful because it can expose weakness in current controls. If a severe or recurring failure has little chance of being detected before the effect occurs, the team should examine whether prevention or detection controls can be improved. That may lead to condition monitoring, an engineered interlock, a process control, a proof test, an inspection change, or another control appropriate to the failure mechanism and application.

What a poor detection rating does not tell you is who owns the problem. Detection is a technical rating within the FMEA method. It does not, by itself, mean the root problem sits in another department or above the plant floor.

When the Recommended Action Is Outside Engineering Control

The distinction between a failure mode and a management constraint matters. A failure mode describes how an item or process fails to perform its intended function. A seized bearing, loss of pump flow, an incorrect assembly, or failure of a protective function can be failure modes, depending on the scope of the analysis. Chronic underfunding, an unapproved capital request, or a staffing shortage is normally not the equipment or process failure mode. It may be a contributing condition, an organizational risk, or a constraint that prevents the recommended action from being implemented.

That does not make the constraint irrelevant. If the FMEA identifies a legitimate technical action but the team lacks the authority or resources to implement it, the action should still be documented and assigned to an owner who can make the decision. The technical record should remain clear about the failure mode, effect, cause, current controls, priority, recommended action, responsibility, and status required by the organization’s FMEA procedure.

Engineering does not have to own every decision for the FMEA to be useful. Its job is to make the technical risk and the proposed treatment clear enough that the correct decision-maker can act on it.

How to Prioritize Failure Modes You Cannot Engineer Away

When a recommended action depends on capital, headcount, purchasing authority, or another management decision, escalate the action rather than redefining the management issue as the failure mode. Keep the FMEA focused on the technical or process failure being analyzed, then record the resource or authority constraint in the action-tracking or risk-management process used by the site.

That escalation is stronger when the analysis is specific. Leadership should be able to see the failure mode, credible effects, the basis for the ratings or priority, the current controls, the proposed action, and what remains exposed if the action is not approved. An RPN or an occurrence rating should not be presented as a calculated probability unless the underlying method actually provides a quantitative probability model.

If your team needs to sharpen the method before the next review, it is worth stepping back through How to Perform FMEA and checking that functions, failure modes, effects, causes, controls, ratings, and actions are being kept distinct.

Keeping the Analysis Honest

A useful FMEA is maintained as evidence changes. If a failure occurs that the team had rated as unlikely, review the occurrence rating against the scoring criteria and the new evidence. The correct response is not automatically to increase the score after a single event, nor to defend the old rating because it is already approved. Update the rating when the evidence and the adopted scoring method justify the change, and document why.

Prioritizing failure modes is ultimately about making disciplined decisions from imperfect information. RPN can be useful in methods that still use it, but the full failure analysis matters more than the product of three ratings. When an important action sits outside engineering’s authority, preserve the technical logic of the FMEA, assign the action to the right owner, and escalate the decision without turning a budget or staffing problem into a failure mode it is not.

Sources

IEC 60812:2018, Failure modes and effects analysis (FMEA and FMECA). https://webstore.iec.ch/en/publication/26359

AIAG, “It’s Here…Claim Your Copy of the New AIAG & VDA FMEA Handbook Today!” (2019). https://blog.aiag.org/its-here…claim-your-copy-of-the-new-aiag-vda-fmea-handbook-today

AIAG & VDA FMEA Handbook, official AIAG manual page. https://www.aiag.org/training-and-resources/manuals/details/FMEAAV-1

SAE J1739_202101, Potential Failure Mode and Effects Analysis (FMEA). https://www.sae.org/standards/j1739_202101-potential-failure-mode-effects-analysis-fmea-including-design-fmea-supplemental-fmea-msr-process-fmea

 

Authors

  • Reliable Media

    Reliable Media simplifies complex reliability challenges with clear, actionable content for manufacturing professionals.

    View all posts
  • Alison Field

    Alison Field captures the everyday challenges of manufacturing and plant reliability through sharp, relatable cartoons. Follow her on LinkedIn for daily laughs from the factory floor.

    View all posts
SHARE

You May Also Like