Study Guide

CRE Study Guide: Distributions, Models, and Lifecycle Math

A CRE study guide built around model selection: censored data and Weibull shapes, block diagrams, FMEA versus fault trees, accelerated testing, and…

Updated September 202612 min readStudy GuideEnergy Cert Exam
Daniel Morgan — Editorial profile

Editorial profile

Daniel Morgan

Energy Cert Exam Editorial Team

Prepare for the Certified Reliability Engineer credential by drilling model-selection decisions, not formula recall. Work censored datasets by hand, compute block diagrams and k-out-of-n systems on paper, contrast FMEA, FMECA, and fault trees on the same system, and practice accelerated-test extrapolations with stated assumptions. Finish by scoring yourself against the readiness rubric in the final section.

Why an MTBF Figure Alone Cannot Tell You When to Replace Anything

MTBF summarizes the average time between failures for items already in service; it carries no information about the failure pattern. Replacement decisions require a distribution or hazard model, because the same average can describe wear-out, random failure, or infant mortality.

Under the exponential distribution, which is the only case where MTBF fully characterizes behavior, the hazard rate is constant and the distribution is memoryless: a unit that has survived 7,900 hours has exactly the same remaining-life distribution as a brand-new unit. That has a hard consequence. If failures are genuinely random in time, replacing an item at any age does not reduce future failures; age-based preventive replacement only helps when the hazard increases with age, which is a property of the underlying distribution, not of the average.

Worked scenario: a plant observes controller-board failures averaging 8,000 hours between them, and the maintenance planner schedules fleet-wide replacement at 8,000 hours. The plausible mistake is treating MTBF as a wear-out point. The better decision is to plot the individual failure times on Weibull axes first. If the fitted shape parameter comes out around 0.7, early-life defects dominate, and replacement at any age will not help; the effective responses are incoming screening, burn-in, or supplier corrective action. Only a shape above 1 justifies age-based replacement, and the interval should then be set against the cost of failure. The distinction matters because the wrong decision spends budget while leaving the real failure driver untouched.

Reading Censored Data: What Survivors and the Weibull Shape Tell You

Field datasets mix failed units with survivors that were removed or are still running. Median-rank and Kaplan-Meier methods use failures as data points while survivors only shrink the risk set; the fitted Weibull shape classifies the failure pattern.

Censoring is easy to mishandle. A unit that has run 5,000 hours without failing contributes no failure time, but it does contribute to the denominator of the failure probability at every earlier failure time. Averaging only the failed units, or discarding survivors before fitting, biases the estimated mean downward and can flip the apparent shape of the hazard. When you practice, always ask two questions of a dataset: which units failed, and what happened to every other unit.

The shape parameter is the interpretive payoff. A Weibull shape below 1 signals a decreasing hazard, meaning infant mortality, and points toward screening and burn-in rather than replacement intervals. A shape above 1 signals wear-out, where preventive replacement can genuinely reduce failures. A shape of 1 collapses to the exponential case from the previous section. Treat the estimate with humility when few failures exist: with three or four failure times, the confidence range on the shape is wide, and the classification can change with one more data point.

Distribution cheat sheet for practice problems:

DistributionHazard patternTypical reliability useKey caution
ExponentialConstant hazardSteady-state items, simple allocationValid only when the memoryless assumption holds
WeibullBelow 1 decreasing, 1 constant, above 1 increasingField failure data, wear-out modelingShape needs enough failures to estimate
LognormalIncreases then decreasesRepair times, degradation-driven failuresHeavy tails dominate small samples
NormalSymmetric around the meanComponent life with tight spreadThe negative-life tail is a modeling artifact

Series, Parallel, and k-out-of-n: Compute the Block Before the System

Reliability block diagrams turn a system layout into one number. Series blocks multiply; parallel blocks combine through complements; k-out-of-n systems need the binomial sum. Confusing active redundancy with a standby path changes the result substantially.

The three core computations are worth doing by hand until they are automatic. For series blocks with reliabilities R1 through Rn, multiply them. For two active parallel blocks, take one minus the product of the unreliabilities. For k-out-of-n, sum the binomial terms for exactly k, k+1, up to n successes. Compare two designs with component reliability 0.9: two active parallel units give 1 minus 0.01, which is 0.99, while a two-out-of-three arrangement gives 3 times 0.81 times 0.1 plus 0.729, which is 0.972. The two-out-of-three design trades some mission success for the ability to keep operating on one channel while the other is repaired, which is why it appears where false trips are costly.

Worked scenario: two parallel pumps, each with a reliability of 0.90 for a 100-hour mission, give a system mission reliability of 0.99. The plausible mistake is quoting the pumps' steady-state availability, computed as MTBF over the sum of MTBF and MTTR, as though it were mission reliability. These are different quantities. Availability is a long-run fraction of time the system is up, averaged over many repair cycles; reliability for a mission is the probability of no failure through the mission length. A pump with 0.999 availability can still carry a meaningful chance of tripping during a specific 100-hour run, and the decision that matters, whether to add a third pump, depends on which quantity the requirement actually specifies.

FMEA, FMECA, and Fault Trees Answer Different Questions

FMEA works bottom-up, enumerating component failure modes and their effects. FMECA adds severity and criticality ranking. Fault tree analysis works top-down from one undesired event through logic gates. Picking the wrong direction wastes the analysis.

A FMEA is a worksheet discipline: list each component, its failure modes, local and system-level effects, then rate severity, occurrence, and detection, and rank modes by a composite score or by criticality in a FMECA. Its strength is coverage of individual modes; its structural limit is that it does not combine failures. Two rows showing a pump failure and a valve failure do not, by themselves, reveal whether their coincidence is what actually causes system loss.

A fault tree starts from a single top event, such as loss of cooling, and decomposes it through OR and AND gates into basic events, then reduces the tree to minimal cut sets. Worked mini-scenario: the tree for coolant loss shows an AND gate combining pump failure and valve failure, plus an OR branch for a single operator-error path. The one-element cut set, the operator error, becomes the priority because it alone causes the top event, and the AND-gate cut set shows the redundancy that no individual FMEA row would expose. The tools complement each other: FMEA output supplies well-vetted basic events for the tree, while the tree supplies the combination logic the worksheet lacks.

Tool comparison at a glance:

ToolDirectionQuestion it answersTypical output
FMEABottom-upWhat can each part fail into?Ranked failure-mode list
FMECABottom-up with criticalityWhich modes matter most?Severity and criticality ranking
Fault tree analysisTop-downHow can this specific top event occur?Logic gates and minimal cut sets
Reliability block diagramFunction-orientedDoes the configuration meet a reliability target?System reliability from block math

Accelerated Testing: Extrapolate the Mechanism, Not Just the Hours

Acceleration is valid only when elevated stress speeds up the same failure mechanism that operates at use conditions. Temperature models such as Arrhenius convert a stress difference into an acceleration factor, and that factor inherits every assumption in the activation energy used.

The Arrhenius acceleration factor takes the form of an exponential of the activation energy divided by Boltzmann's constant, times the difference of the reciprocals of the use and stress temperatures in kelvin. Two properties deserve attention in practice. First, the factor is extremely sensitive to the activation energy: because the energy enters the exponent, a modest error in it scales the projected life by multiples rather than percentages. Second, a model fitted from a single stress pair cannot confirm itself; two or more stress levels, with times-to-failure that line up under the model, are what give the extrapolation support.

Worked scenario: an adhesive joint is tested at 125 degrees Celsius, and the engineer applies a rough rule of thumb that life halves every 10 degrees to claim a ten-year life at a 55-degree operating point. The plausible mistake is that the generic rule embeds an activation energy that was never measured for this adhesive and this mechanism. The better decision is to test at two or more temperatures, estimate the activation energy by regression against the Arrhenius line, and verify by inspection that the failure mode observed at stress matches the one seen in field returns before quoting any factor. It also matters which stresses dominate: thermal models say nothing about vibration or voltage mechanisms, so each stress needs its own model and its own justification.

Availability, Maintainability, and the Time Terms Behind Them

Inherent availability divides mean operating time between failures by that plus active repair time. Achieved availability adds preventive maintenance downtime, and operational availability adds logistics and administrative delay. Each term answers a different ownership question.

The definitions form a chain worth tracing deliberately. Start with MTBF for the failure side and mean time to repair for the restoration side; their ratio yields inherent availability, which is a property of the design alone. Add scheduled maintenance time and you have achieved availability, a property of the maintenance plan. Add waiting for parts, crews, and paperwork and you have operational availability, which is what the user actually experiences. When a specification states a target, identify which of the three is meant, because a design meeting the inherent figure can still miss the operational one.

Maintainability adds a second dimension: repair time is a distribution, not just a mean, and two designs with the same MTTR can differ sharply in how often repairs run far past it. This creates a classic tradeoff to practice. Adding a second active pump raises mission reliability, as computed earlier, but it also doubles the preventive maintenance workload and adds failure modes to the fault tree. A sound design decision computes both sides: the reliability gain from redundancy and the availability and cost of the added maintenance, then compares them against the requirement rather than optimizing one number in isolation.

Lifecycle Reliability: Allocation, Growth, Derating, and Your Study Sequence

Lifecycle reliability connects early design decisions, allocation, derating, and growth testing, to in-service performance. A study sequence that moves from data handling through modeling to lifecycle tradeoffs mirrors that flow and keeps every session concrete.

Three lifecycle tools deserve deliberate practice. Allocation distributes a system reliability target across subsystems, whether equally or weighted by complexity and criticality, so each design owner gets a number they can verify. Derating runs components well below rated stress so the operating point sits on a flatter part of the hazard curve. Reliability growth models, such as the Duane and Crow approaches, track how mean time between failures improves across test-fix-test cycles, and they rest on the assumption that fixes found in test are actually incorporated and re-verified, not just logged.

Practical exercise with expected observations: take this paper dataset of failure times in hours, 120, 340, 780, 1500, 2600, and 4200, plus three units censored at 5000 hours. Compute median ranks or a Kaplan-Meier estimate, sketch the points on Weibull axes, and estimate the shape parameter. A correct treatment gives a shape below 1, meaning a decreasing hazard, and the conclusion that screening or burn-in, not replacement intervals, addresses the pattern. If your plot suggests a shape above 1, recheck how you handled the censored units in the ranks, because including or discarding them wrongly is exactly the bias described earlier.

Self-check rubric, scored as learning milestones rather than pass predictions: first, for each formula you use, you can state its assumption in one sentence. Second, you can compute series, parallel, and k-out-of-n reliability by hand without notes. Third, you can explain the difference between availability at time t and mission reliability in a single exchange. Fourth, you can name the censoring treatment applied to every non-failed unit in a dataset. Fifth, you can explain why a single stress level cannot confirm an acceleration model.

  • Sessions 1 to 2: fundamentals. Hazard functions, the distributions in the table above, and censoring; rebuild the exercise dataset until rank handling feels routine.
  • Sessions 3 to 4: modeling. Draw block diagrams for systems from your own work and hand-compute series, parallel, and k-out-of-n results before touching any software.
  • Sessions 5 to 6: risk tools. Run one FMEA and one fault tree on the same small system, then write down what each found that the other missed.
  • Sessions 7 to 8: testing and lifecycle. Work Arrhenius and reliability-growth problems, and draft an allocation for a system you know, with derating choices justified.
  • Final pass: mixed paper scenarios from the free practice set, scored against the rubric; rework only the items that miss.

References and further reading

Use these references to explore the concepts and check the latest information from the relevant organizations.

Continue your preparation

FAQ

Frequently Asked Questions

Practical answers to help you apply the guidance for Certified Reliability Engineer (CRE).

Is the exponential distribution a safe default for reliability problems?
No. It is exactly correct only when the hazard is genuinely constant, which is a property to verify from the failure pattern or the mechanism, not to assume. A Weibull fit that shows a shape differing from 1 changes the conclusion about replacement, screening, or burn-in, so check the shape before reaching for the exponential shortcut.
How much statistics does this credential's body of knowledge require beyond MTBF?
Working comfort with estimation from censored data, distribution fitting and interpretation of shape parameters, binomial combinations for redundant configurations, and the distinction between long-run time averages and mission probabilities. Small hand-computed datasets are enough to build all of these; real-scale software is a convenience, not the skill.
Should I practice FMEA or fault tree analysis more heavily?
Both, ideally on the same system. FMEA drills disciplined enumeration of failure modes and effects, while fault trees drill logic decomposition and minimal cut sets, which the worksheet format cannot express. Doing one of each on a shared example shows you exactly which failure insight each method captures and which it misses.
Can I prepare well without access to real failure data?
Yes. Constructed paper datasets with a mix of failure times and censored survivors let you practice ranking, plotting, and interpreting shapes, and scenario-based problems let you practice model selection. The decision skill, matching tool to question, transfers from paper exercises; what paper cannot teach is judgment about data quality, which you can note as a gap for workplace experience.
Where should I confirm administrative details such as requirements, scheduling, and fees?
Use ASQ's official Certified Reliability Engineer page at asq.org/cert/reliability-engineer for eligibility, exam administration, and recertification details. This guide covers the technical body of knowledge and study method; administrative specifics belong to the issuing body and can change, so treat their page as the reference for those items.

Keep Reading

Related Study Guides

Explore related guides and preparation topics.