I remember a late-night call from a water treatment plant in Minas Gerais. The operator was exasperated. “Diego, the backup generator kicked in after three seconds. The PLC never even blinked. But the SCADA screen stayed black for twenty minutes.” He paused, then said something that’s stuck with me ever since: “The system was reliable—but it wasn’t available.”
That one sentence captures a muddle I see all the time in engineering teams, especially those bridging Portuguese and English documentation. We toss around confiabilidade and disponibilidade like they’re the same thing. Out in the field, though, the gap between reliability and availability can mean the difference between a minor hiccup and a full-blown regulatory headache.
This piece is for the engineers and technicians who keep physical systems running—pumping stations, conveyor networks, remote telemetry units, power distribution panels. We’ll strip both metrics down to what they actually measure, see how they trip over each other, and figure out why a bulletproof component can still leave you with a dead system. I’ll pull examples from real Brazilian industrial sites, where the distance to a spare part can be measured in weeks, not hours.
Reliability and Availability, Without the Jargon
Let’s clear the fog first. Even sharp engineers blur these lines when flipping through English manuals or drafting reports for international partners.
Reliability: How Long Before It Quits
Reliability is the probability a system does its job without failing, for a set time, under stated conditions. In the field, we usually boil it down to Mean Time Between Failures (MTBF). If a pressure transmitter carries an MTBF of 50,000 hours, you can reasonably expect it to hum along for about five and a half years before it gives up—assuming it’s installed right and not cooked by conditions outside its spec.
Reliability cares about one thing: how long something runs before it breaks. It doesn’t blink at how fast you patch it up. A cast-iron diesel engine that chugs for a decade with basic upkeep? That’s high reliability.
Availability: Is It Ready When You Need It
Availability is the slice of time a system is actually in working order. The standard formula:
Availability = MTBF / (MTBF + MTTR)
MTTR is Mean Time To Repair. And repair doesn’t just mean wrench time. It swallows everything: spotting the failure, driving to the site, diagnosing the mess, hunting down the spare part, doing the fix, and testing afterward.
Availability obsesses over uptime. A system can be absurdly reliable—failing once every five years—but if that single failure takes three months to sort out because the spare comes from Germany, your availability number tanks. I’ve watched this play out with niche PLC modules and custom valve actuators more times than I can count.
Two Pumps, One Hard Lesson
Let’s make this concrete with a comparison from a mining operation in Pará.
Pump A: The Tank
Pump A is a heavy-duty centrifugal unit, a design that’s been around forever. Its MTBF clocks in at 20,000 hours—over two years of nonstop running. But when it finally throws a bearing, the repair eats 200 hours because the impeller has to be shipped to São Paulo for rebalancing. Run the numbers:
Availability = 20,000 / (20,000 + 200) = 0.990, or 99.0%
Pump B: The Swappable One
Pump B is a modular unit with a seal design that’s, frankly, a bit delicate. Its MTBF is a mere 5,000 hours. But the whole pump cartridge can be swapped in 4 hours by a local tech with a wrench and a coffee. Availability:
Availability = 5,000 / (5,000 + 4) = 0.9992, or 99.92%
Pump B fails four times as often—it’s objectively less reliable—but it’s more available. For a process that can’t stomach downtime, Pump B might be the smarter pick, especially if you keep a spare cartridge on the shelf. This is the kind of trade-off that maintenance managers and field engineers need to wrestle with when writing specs or planning redundancy.
Why the Mix-Up Sticks Around in Multilingual Teams
In Portuguese, we lean on confiabilidade to cover both ideas. An English datasheet might boast “high reliability” when it really means “high availability.” I’ve seen this twist equipment selection in Brazilian projects that reference international standards. A team reads “99.99% reliable” and pictures a component that almost never fails. But the fine print shows that number is availability, propped up by redundant parts and fast swaps—not because each piece is indestructible.
This bites when you’re drafting maintenance plans or negotiating service contracts. If you promise a client 99.9% availability, you need to budget for the MTTR that makes that possible, not just buy components with shiny MTBF figures.
Crunching the Numbers from Real Field Data
Let’s walk through a calculation using data from a remote telemetry unit (RTU) watching over a pipeline.
Step 1: Gather the Messy Details
Over one year (8,760 hours), the RTU tripped three times:
- Failure 1: 2 hours to repair
- Failure 2: 1.5 hours to repair
- Failure 3: 3 hours to repair
Total repair time = 6.5 hours. Total operating time = 8,760 – 6.5 = 8,753.5 hours.
Step 2: Calculate MTBF and MTTR
MTBF = Total operating time / Number of failures = 8,753.5 / 3 ≈ 2,917.8 hours
MTTR = Total repair time / Number of failures = 6.5 / 3 ≈ 2.17 hours
Step 3: Calculate Availability
Availability = MTBF / (MTBF + MTTR) = 2,917.8 / (2,917.8 + 2.17) = 0.99926, or 99.926%
This RTU is nicely available, but its reliability is so-so—you can expect a failure roughly every four months. If this RTU is the brain for a critical valve, you might want redundancy to dodge unplanned shutdowns.
How Your Maintenance Strategy Shapes Both Numbers
Your maintenance approach pulls reliability and availability in different directions. Here’s how three common strategies play out.
Reactive Maintenance (Run-to-Failure)
You fix stuff only after it breaks. This can stretch MTBF because you’re not interrupting operation for preventive poking. But MTTR often balloons—failures are surprises, so you might not have spares or people ready. Availability takes the hit. This works for non-critical assets where downtime is cheap.
Preventive Maintenance (Time-Based)
You swap parts on a calendar, like changing oil every 500 hours. This can actually hurt reliability if you’re yanking components that still have life left—you’re rolling the dice on maintenance-induced failures. But availability can climb because you schedule downtime during planned windows. The trick is to base intervals on your own failure data, not manufacturer defaults that assume the worst.
Predictive Maintenance (Condition-Based)
You watch equipment condition—vibration, temperature, oil analysis—and step in only when signs of trouble show up. This can boost both reliability and availability because you skip unnecessary interventions while catching problems before they snowball. At a sugar mill in São Paulo, we slapped vibration sensors on critical crusher motors. MTBF jumped 40% because we stopped doing calendar-based bearing swaps. Availability rose because we scheduled repairs during planned stops.
Redundancy: The Quick Fix for Availability
Redundancy is the most direct lever to yank up availability without touching reliability. If a single pump gives you 99% availability, two pumps in parallel can hit 99.99%—assuming a flawless switchover. But field reality is messier.
I once audited a system with dual redundant PLCs. On paper, availability was 99.999%. In practice, the switchover mechanism had a bug that caused a 30-second delay every time. That half-minute was enough to trip downstream equipment. The lesson: redundancy piles on complexity, and complexity hides failure modes that nibble away at availability. Always test failover under real load.
Traps I Keep Seeing in Field Systems
Here are mistakes I’ve stumbled across again and again in Brazilian industrial plants.
Forgetting Detection Time
Availability formulas pretend you know the instant a failure happens. At a remote site, a pump can die at 2 a.m. and nobody notices until the morning crew rolls in. That’s 6 hours of invisible downtime. Your real MTTR includes detection time, not just wrench time. Add remote monitoring or at least basic alarms to shrink that gap.
Mixing Up Component and System Metrics
A vendor hands you an MTBF of 100,000 hours for a pressure transmitter. Nice. But your system has 50 of them. The system-level failure rate is 50 times higher—you’ll see a transmitter failure every 2,000 hours on average. System reliability depends on the weakest link and the pile-up of many components.
Ignoring the Human Factor
MTTR assumes a skilled tech is standing by with the right spare part. In remote spots, travel time alone can tack on 4–8 hours. If the part isn’t in local stock, add weeks. I’ve seen plants in the interior of São Paulo wait 15 days for a specialized circuit board from Europe. That single event dragged annual availability from 99.5% down to 95%.
Designing for Both, Not Either-Or
You don’t have to pick one over the other. Smart design can lift both.
Standardize on Modular Components
When every pump in the plant shares the same mechanical seal, you stock fewer spares and techs learn one repair routine cold. MTTR drops. Reliability might even climb because the team gets really good at that specific fix.
Add Condition Monitoring
Sensors on critical assets let you spot degradation early. You can plan maintenance before failure, which helps reliability (you replace before breakdown) and availability (you schedule the repair during a planned window).
Design for Maintainability
Can a tech reach the failed part without pulling three other assemblies? Are diagnostic LEDs visible without cracking open the panel? These tiny design choices hammer MTTR. When specifying equipment, ask vendors for maintainability data, not just MTBF.
Talking About Metrics Across Teams
In Brazilian engineering environments, you often have field techs who speak Portuguese and project managers or vendors who communicate in English. Misunderstandings about reliability vs. availability can set wrong expectations.
I suggest creating a simple one-page cheat sheet for your team:
- Confiabilidade (Reliability): Probabilidade de operar sem falhas por um período. Medida como MTBF.
- Disponibilidade (Availability): Percentual do tempo em que o sistema está operacional. Medida como MTBF/(MTBF+MTTR).
- Exemplo prático: Um compressor com MTBF de 10.000 horas e MTTR de 50 horas tem disponibilidade de 99,5%.
Stick this in the maintenance workshop and fold it into training materials. When everyone speaks the same language, you dodge costly spec errors.
FAQ
Can a system be reliable but not available?
Absolutely. A system with very long stretches between failures (high reliability) can still have low availability if each repair drags on forever. Picture a subsea valve that fails once every 10 years but needs 6 months to retrieve and fix. Reliability: 10 years MTBF. Availability: a limp 95%.
Which metric matters more for a wastewater treatment plant?
Availability usually wins because environmental rules often demand continuous treatment. A plant can get fined for any untreated discharge, even if the equipment itself is rock-solid. In those cases, redundancy and fast repair capability are what keep availability high.
How do I explain the difference to my maintenance team?
Try a car analogy: Reliability is how often the car breaks down. Availability is whether the car is ready to drive when you need it. A car that rarely breaks but takes a month to fix when it does has high reliability but lousy availability. A car that hiccups more often but can be patched in an hour has lower reliability but higher availability.
Pulling It All Together
Reliability and availability aren’t rivals—they’re two sides of the same coin. A well-designed field system needs both: parts that rarely fail and a maintenance setup that restores function fast when they do. Next time you’re staring at a vendor’s spec sheet, look past that single impressive number and ask: “What’s the MTBF? What’s the MTTR? And what’s my real availability if this part quits at 3 a.m. on a Sunday at a remote site?”
That question has saved my clients more downtime than any single gadget ever could.


