What Is MTBF (Mean Time Between Failures)?
Also known as: mean time between failures, MTBF hours
Definition
MTBF (Mean Time Between Failures) is the average operating time a repairable asset runs between consecutive failures, calculated as total operating time divided by the number of failures in that period. It measures reliability, not expected lifetime.
MTBF (Mean Time Between Failures) Explained
The calculation uses operating hours in the numerator, not calendar hours, and counts only failures that stop the asset from performing its function. If a machine runs 2,400 hours in a quarter and fails six times, MTBF is 400 hours. Including calendar time when the asset was idle, or counting planned maintenance events as failures, are the two errors that make MTBF numbers incomparable between departments that thought they were using the same definition.
MTBF applies to repairable systems. The equivalent for non-repairable items, such as a bearing or a fuse that is replaced rather than fixed, is MTTF, mean time to failure. Electronics component data sheets typically quote MTTF or a failure rate in FIT, failures in time per billion device-hours. Confusing the two leads to the classic misreading of a component with a 200,000-hour MTBF as one expected to last twenty-three years.
The deepest misconception is treating MTBF as a lifetime. A drive with a one-million-hour MTBF is not expected to run 114 years. MTBF describes the failure rate during the useful-life portion of the bathtub curve, so it means that within a large population operating in that period, roughly one failure occurs per million cumulative operating hours. Individual units still wear out on their own much shorter service life schedule.
MTBF becomes actionable when it is tracked per asset and per failure mode rather than as a plant average. A plant-level MTBF blends a reliable press with a chronically failing conveyor and moves too slowly to guide action. Per-asset trending, especially on constraint equipment, identifies deteriorating machines while there is still time to plan an intervention, and per-failure-mode analysis distinguishes a spare-parts problem from a maintenance-practice problem.
Why It Matters
- Quantifies reliability trend per asset, converting a subjective sense that a machine is getting worse into evidence for a capital or overhaul request.
- MTBF and MTTR together determine availability, the first term in OEE and the input to realistic capacity planning.
- Drives spare-parts stocking decisions, since failure frequency and lead time jointly set the risk of an extended outage.
- Supports supplier and equipment selection by comparing demonstrated field reliability rather than specification-sheet claims.
In Practice
A grinder logs 1,860 operating hours in a half year with nine failures, giving an MTBF of 207 hours. Six of the nine trace to a single coolant pump seal. Removing that failure mode alone would raise MTBF to about 620 hours. This is why a plant average is nearly useless for action: one component, costing 340 dollars, was responsible for two thirds of the reliability loss on a machine that had been on the capital replacement list for two years.
Frequently Asked Questions
Does a 100,000-hour MTBF mean the equipment lasts 100,000 hours?
No. MTBF describes the failure rate during the useful-life phase of a population, not the service life of one unit. A 100,000-hour MTBF means that across many units operating in that phase, roughly one failure occurs per 100,000 cumulative operating hours. Any individual unit still has a much shorter design life and wear-out period.
What is the difference between MTBF and MTTF?
MTBF applies to repairable assets and measures average operating time between failures that are fixed and returned to service. MTTF applies to non-repairable items that are discarded and replaced, such as bearings, lamps, or electronic components, and measures average operating time until the single failure. Using MTBF for consumable components is technically incorrect though very common.
Related Terms
MTTR (Mean Time To Repair)
MTTR (Mean Time To Repair) is the average time required to restore a failed asset to full operating condition, calculated as total repair time divided by the number of repairs. It measures maintainability rather than reliability.
TPM (Total Productive Maintenance)
TPM (Total Productive Maintenance) is a company-wide approach to equipment reliability that shifts routine care to operators, builds structured planned maintenance, and targets the elimination of breakdowns, speed loss, and defects caused by equipment condition.
OEE (Overall Equipment Effectiveness)
OEE (Overall Equipment Effectiveness) is a manufacturing metric that multiplies Availability, Performance, and Quality into a single percentage showing how much of scheduled production time an asset spends making good parts at its rated speed.
Go Deeper
Manufacturing Downtime Cost Calculator
Combine lost contribution margin, idle labor, and absorbed overhead into a defensible monthly and annual downtime cost - with recovery effects modeled.
MES-ERP Integration for Discrete Manufacturing
MES-ERP integration guide for discrete manufacturers: architecture patterns, ISA-95 mapping, SyteLine and LN connectors, costs, timelines, and AI-driven sync.
ERP RFP Template for Discrete Manufacturers
A complete ERP RFP template for discrete manufacturing: requirements matrix, CMMC and ITAR questions, weighted scoring model, and vendor demo scripts.
Working with MTBF (Mean Time Between Failures) in a live environment? Our engineers do this every day - and our AI agents automate most of it.