When a pump station goes silent at 3 a.m., the first question the control room asks is, “When will it be back online?” That’s an availability question. The harder question, the one that comes later in the morning meeting, is “Why did it fail again?” That’s about reliability. In field engineering, these two terms get tossed around like they mean the same thing. They don’t. And confusing them leads to poor maintenance schedules, wasted spare parts, and a false sense of security.
I’ve spent years working with teams that maintain remote systems—offshore platforms, mining conveyors, water treatment plants in the middle of nowhere. The vocabulary shifts depending on whether you’re talking to a maintenance planner in São Paulo or a reliability engineer in Aberdeen. Confiabilidade versus disponibilidade. The physics of failure doesn’t change, but the way we count the hours sure does.
Defining the Terms with Field Math
Reliability is the probability that a system will perform its required function for a given period under stated conditions. It’s a statistical measure, often expressed as Mean Time Between Failures (MTBF). Availability is simpler: it’s the fraction of time the system is actually working. Uptime divided by total time. A system can be highly available but deeply unreliable, and that paradox is where most field arguments begin.
Consider a diesel generator at a remote telecom tower. It runs for 100 hours, fails, and a technician arrives and fixes it in 2 hours. Availability? About 98%. Now imagine it fails every 100 hours like clockwork. The availability still looks fine on the monthly report, but the reliability is terrible. If you’re the technician bouncing down a dirt road every four days to reset it, you feel that difference. The manager staring at the dashboard doesn’t.

Why High Availability Can Mask Poor Reliability
Many field sites hit their availability targets through redundancy and fast repairs, not because the equipment is well-engineered. A pump station with three identical pumps, each with an MTBF of 500 hours, can still show 99% availability if the crew swaps out failed units quickly. The system is available, but the individual pumps are failing constantly. This works until the spare parts container is empty or a common-mode failure takes out all three at once.
I once reviewed a failure log for submersible pumps in a drainage project. The availability metric was above the contractual target, so the supplier was happy. But the pumps were failing every 400 hours—way below design life. The high availability came from having a full-time technician on site and a shipping container full of spare pumps. The client was buying availability with labor and inventory, not with sound engineering. That’s not a win.
Mean Time Metrics and Their Dirty Data
MTBF and Mean Time To Repair (MTTR) are the classic numbers. Availability is often calculated as MTBF / (MTBF + MTTR). The formula is clean. The data behind it is usually a mess. Field failure data is dirty by nature. A pump that trips on overload and gets reset five minutes later—is that a failure? If the operator resets it without a work order, it never enters the CMMS. The MTBF looks great on paper, but the real reliability is hidden in those unlogged resets.
In practice, I prefer to track interventions per operating period, whether they’re officially called failures or not. An intervention is any unplanned human interaction with the equipment. If a technician has to touch the machine, it wasn’t available, and it wasn’t reliable. That simple rule cuts through a lot of data games.

Designing for Availability vs. Designing for Reliability
These two goals pull the design in opposite directions. Designing for reliability means choosing components with long intrinsic life, simplifying the system to reduce failure modes, and running within conservative limits. Designing for availability means adding redundancy, making things easy to repair, and keeping spare parts on the shelf. The best field systems do both, but budgets force trade-offs.
In remote locations, the cost of a repair visit can be ten times the cost of the failed part. There, reliability is king. You want that pump to run for five years without anyone laying a hand on it. In a plant with 24/7 maintenance coverage, availability might be the smarter bet—use cheaper components and swap them out often. The mistake is applying the same philosophy everywhere without thinking through the logistics.
The Portuguese-English Engineering Bridge
Working across Brazilian and North Sea engineering cultures, I’ve noticed a subtle difference in how these ideas are discussed. In Brazilian technical Portuguese, confiabilidade (reliability) and disponibilidade (availability) often get used interchangeably in casual conversation, even among engineers. The formal definitions are in the ABNT standards, but the field language blurs them. In English-speaking contexts, the distinction is sharper, partly because of the strong influence of groups like the Society of Maintenance and Reliability Professionals. Bridging this gap takes more than translation—it takes aligning the operational philosophy.
For example, a Brazilian maintenance plan might lean heavily on calendar-based preventive tasks, aiming to ensure availability. A North Sea plan might lean toward condition-based monitoring, targeting reliability. Both can work, but mixing them without understanding the underlying goal creates a mess. I’ve seen teams replace bearings on a fixed schedule (availability-focused) while also installing vibration sensors (reliability-focused) and then ignoring the sensor data because the schedule said the bearing was fine. Wasted effort on both sides.
Practical Ways to Measure Both in the Field
You can’t manage what you don’t measure, but you also can’t measure everything. Here are the methods I’ve found practical for field systems with limited instrumentation.
1. The Logbook Method for Reliability
For older systems without digital historians, the operator logbook is your best reliability data source. Train operators to record every anomaly, not just trips and alarms. A vibration that “feels different” or a pressure gauge that flickers is early failure data. Review these logs weekly and classify events. Over time, you build a failure pattern without a single sensor. This method is labor-intensive but costs nothing in capital.
2. The Stopwatch Method for Availability
Availability is simpler to measure: when is the system producing output? For a pump, that means flow. For a conveyor, that means material moving. Install a simple hour meter on the motor starter. Compare running hours to elapsed calendar hours. If the system should run continuously, availability is just the ratio. If it runs intermittently, compare running hours to demanded hours. This avoids complex SCADA integration and works on any machine with an electric motor.

3. The Intervention Ratio
Divide the number of unplanned maintenance interventions by total operating hours. This ratio captures both reliability and availability in one number. A system with high reliability and low availability will have a low intervention ratio because it fails infrequently, even if repairs are slow. A system with low reliability but high availability will have a high intervention ratio—frequent failures, fast fixes. This single metric often tells you more than MTBF and MTTR combined.
Common Failure Patterns and What They Reveal
Field systems tend to follow a few characteristic failure patterns. Recognizing them helps you decide whether to invest in reliability or availability improvements.
Infant mortality: Failures occur early in life due to manufacturing defects or installation errors. Improving reliability here means better commissioning and quality control. Increasing availability by stocking spares just masks the problem.
Random failures: Constant failure rate over time, often from external events like power surges or operator errors. Here, availability strategies like redundancy and rapid repair make sense because you can’t design out the root cause.
Wear-out failures: Increasing failure rate at end of life. Reliability engineering focuses on predicting this point and replacing components before failure. Availability strategies like redundancy are less effective because all similar components wear out around the same time.
When High Availability Becomes a Trap
I’ve seen sites where the availability metric was 99.5%, and management celebrated. But the maintenance team was exhausted, the spare parts budget was blown, and the technicians knew the next failure was just around the corner. This is the availability trap: using redundancy and fast repair to hit a number while ignoring the underlying reliability rot. The system is one common-mode failure away from a catastrophic outage, and the team is too burned out to see it coming.
The fix is to temporarily ban the word “availability” from performance reviews. Force the team to discuss failure causes, not just downtime minutes. Ask: “What broke, and why?” instead of “How fast did we fix it?” This shift in language changes behavior. Technicians start looking for root causes instead of quick resets. Engineers start analyzing failure patterns instead of ordering more spares.
Bridging the Gap with Condition-Based Maintenance
Condition-based maintenance (CBM) is the practical bridge between reliability and availability. By monitoring actual equipment condition—vibration, temperature, oil quality—you can predict failures before they happen and plan repairs during scheduled downtime. This improves reliability by catching degradation early and improves availability by reducing unplanned outages.
But CBM isn’t a magic wand. I’ve seen sites invest heavily in sensors and software, then drown in data they can’t interpret. The trick is to start simple: pick one failure mode that costs you the most downtime, install a sensor to detect its precursor, and prove the concept. Once the team trusts the data, expand. A single vibration sensor on a critical pump, paired with a trained technician who checks it weekly, can deliver more value than a full online monitoring system that nobody looks at.
FAQ: Reliability and Availability in Field Systems
Can a system be reliable but not available?
Yes. A system with very long MTBF but extremely long repair times can be reliable but unavailable. For example, a custom-built compressor that runs for 10,000 hours without failure but takes 1,000 hours to repair because parts must be machined from scratch. Its reliability is excellent, but its availability is only 91%. This scenario is common with specialized, low-volume equipment.
Which metric matters more for remote field systems?
Reliability usually matters more because the cost of a repair visit dominates. If sending a technician costs $5,000 in logistics, you want the equipment to run for years without attention. Availability through redundancy helps only if the redundant unit can take over automatically and the failed unit can be repaired during a scheduled visit. Otherwise, you still need to send someone, and the availability gain is minimal.
How do I convince management to invest in reliability instead of just tracking availability?
Show them the total cost of ownership, not just uptime. Calculate the annual cost of unplanned repairs, including labor, travel, spare parts, and lost production. Then compare that to the cost of upgrading to more reliable components or implementing condition monitoring. When the numbers are laid out in currency, not percentages, the argument becomes clear. One site I worked with reduced annual maintenance costs by 40% after switching from a run-to-failure approach to a reliability-centered design upgrade.
What is the simplest way to start measuring reliability without a CMMS?
Use a paper log and a wall chart. Track two things: every time the equipment stops unexpectedly, and every time someone touches it for an unplanned reason. Plot these events on a timeline. Within a few months, you will see patterns—certain pumps fail after rain, certain conveyors trip on Monday mornings. This low-tech approach builds the data discipline needed for a CMMS later and costs nothing but attention.