Every asset has multiple ways it can fail to perform its required functions. Identifying those scenarios before failure lets a team decide which ones deserve prevention, condition monitoring, failure-finding, redesign, or a deliberate run-to-failure policy. The exercise is valuable because it turns vague risk into specific failure scenarios that can be analyzed before downtime forces the discussion.
What a Failure Mode Actually Is
A failure mode is the manner in which an item or process fails. The exact wording depends on the level of analysis. At the pump-system level, no flow or insufficient flow may be failure modes or functional failures. At a bearing level, seizure or spalling may be failure modes. Inadequate lubrication, by contrast, is usually a cause or contributing condition, not the mode itself. Keeping mode, cause, and effect separate prevents the analysis from mixing different levels.
Write modes at a level that can be tied to effects, causes, and decisions. Seal leaks can be a failure mode; elastomer hardening may be one cause. Bearing seizes can be a failure mode at the bearing level; inadequate lubrication may be one cause. Do not jump straight from a cause to a maintenance interval. Task selection depends on consequence, failure behavior, detectability, and whether the proposed task is technically feasible and worth doing.
A named failure mode is not automatically a preventive-maintenance task. It is a decision point: detect it, prevent it, find a hidden failure, redesign the condition, accept run-to-failure, or document why no proactive task is justified.
How to Identify Failure Modes From Function
A reliable starting point is the required function and its performance standard. A cooling pump function might be to deliver at least 400 gallons per minute at a specified differential pressure and operating condition. Once success is defined, ask how that requirement can be lost: no flow, insufficient flow, excessive flow where relevant, wrong direction, leakage, intermittent delivery, or operation at the wrong time. In RCM language these are often functional failures; in an FMEA, the team may then work down to component-level failure modes that could produce them.
Working from function helps catch total, partial, intermittent, and unintended operation, including failures of protective or secondary functions. A relief valve that fails to open on demand and an indicator that gives a false normal reading are different failure scenarios with different consequences. This functional framing supports a structured FMEA effort and gives the team a sound starting point for how to perform FMEA at the chosen analysis level. Because FMEA can be performed at different levels, the team should state the level before naming modes.
Use prompts, not a rigid four-question rule. Depending on the function, ask whether it can fail to start, fail to stop, operate prematurely, operate late, operate intermittently, produce too little or too much output, operate in the wrong direction, or provide a misleading indication. The applicable list depends on the asset and mission; the point is to challenge comfortable assumptions systematically.
You cannot list failure modes well until you have defined what successful performance means, using measurable limits where practical.
How to Identify Failure Modes From History and Physics
Function identifies what must be achieved; history shows which failures have actually occurred. Work-order records, breakdown logs, inspections, and operator or maintainer experience can provide evidence about relevant modes and causes on this asset or similar assets. Use that evidence, but do not assume the history is complete or correctly coded.
History has limits because it only records events that have already occurred and been captured. Engineering and physics expand the list to plausible failures that have not yet appeared in the record. Understanding physical component failure mechanisms such as fatigue, corrosion, erosion, wear, overload, or electrical damage can help explain how a mode might develop. Material, load, environment, and design details narrow what is physically plausible, but they do not predict an individual failure with certainty.
Use history and physics together. History anchors the analysis in plant experience; engineering reasoning expands it beyond events already recorded. When a new failure occurs, a disciplined root cause failure analysis can update the failure-mode list, causes, controls, and task strategy. The goal is not to promise that a surprise never repeats, but to make the learning persistent.
Operators and maintainers are important sources because they see startup behavior, abnormal sounds, recurring workarounds, and repeat part replacements that may not be coded well in the CMMS. Capture those observations, then verify them against physical evidence and records where possible. Experience is valuable input, not a substitute for analysis.
Ranking the Modes Once You Have Them
A raw list of failure modes is a starting point, not a plan. Prioritization can use consequence severity, likelihood or occurrence, detectability, criticality, or other criteria depending on the method. IEC 60812 describes multiple approaches; a traditional FMEA often scores severity, occurrence, and detection, while FMECA adds criticality considerations. Do not let a simple combined score hide a low-probability mode with an unacceptable safety, environmental, or regulatory consequence.
Prioritization also keeps the exercise manageable. Not every mode needs proactive maintenance. A low-consequence mode may be a valid run-to-failure candidate when that choice is safe, environmentally acceptable, operationally tolerable, and economically justified. High-consequence hidden or protective-function failures may deserve attention even when they are infrequent. The response follows consequence and feasibility, not frequency alone.
If you use numerical scoring, define the scales and decision rules before the workshop and keep them consistent. A simple risk priority number or combined score can help sort discussion, but the arithmetic should not be treated as an absolute measure of risk; different combinations can produce the same number, and high severity may warrant action regardless of the total. Use scores to support, not replace, engineering judgment.
The score is a pointer to a decision. The evidence, assumptions, and rationale behind it are the part worth preserving.
The point of listing failure modes is not to prevent all of them. It is to choose a deliberate, defensible response for the modes that matter.
From Failure Modes to Action
A finished failure-mode analysis should map significant modes to an appropriate strategy. Use condition-based tasks when a detectable potential failure gives enough warning and the task is technically feasible and effective. Use scheduled restoration or replacement only when age-related behavior and a defensible interval support it. Use failure-finding tasks for applicable hidden protective functions. Consider redesign or operating changes when no maintenance task adequately controls an important consequence, and run to failure only when the consequences are acceptable. The analysis is the bridge from this can happen to here is how we will manage it.
Done well, this discipline moves a plant off its back foot. Instead of discovering every known failure scenario through downtime, the team identifies important modes, evaluates their consequences, and chooses controls in advance. That can reduce avoidable repeat failures and improve reliability where the selected controls are effective. It does not make every failure predictable. The benefit is that known risks have an explicit response and new failures feed back into the analysis instead of remaining isolated breakdown stories.
Failure-mode analysis also preserves knowledge. A ranked list with function, mode, effect, cause, existing controls, and chosen action gives new technicians a structured map of how the asset can fail and why tasks exist. It will not replace field experience, but it can shorten the time needed to understand recurring risks and prevent important knowledge from living only in a few people’s heads.
So the practical summary of how to identify failure modes is to start from function and performance, mine history for what has actually happened, use physics and engineering to expand the plausible list, separate modes from causes and effects, and prioritize by consequence and risk. Start with critical assets if the scope is limited, review the analysis after significant changes or new failures, and keep the assumptions current. The result is not a perfect list of everything that could happen; it is a maintained record of failure scenarios the plant has deliberately chosen how to manage.









