Standing in front of a tripped panel, the difference between a system that’s “reliable” and one that’s “available” stops being a textbook debate. It becomes the difference between a quick reset and a long, frustrating night of tracing wires. In the field, we toss these words around like they mean the same thing, but they measure two very different realities. Mixing them up leads to bad designs and even worse maintenance plans.

Defining the Terms on the Shop Floor
Let’s skip the academic jargon and get down to brass tacks. Reliability is the probability a system will do its job for a set time without failing. It’s measured by Mean Time Between Failures (MTBF). A high MTBF simply means the kit doesn’t break very often. It’s a measure of the machine’s inherent toughness, baked in at the factory.
Availability, on the other hand, is the percentage of time the system is actually ready to work. It’s the uptime metric. The formula is straightforward: Availability = MTBF / (MTBF + MTTR). MTTR is the Mean Time To Repair. Here’s the kicker: a system can fail constantly and still be highly available, as long as you can fix it in seconds. Picture a lightbulb in a socket that takes two seconds to swap. It might pop every month, but its availability is near 100% because the repair time is a rounding error.
This isn’t just theory. I’ve walked into plants where management demanded “maximum reliability” and got a system so over-engineered and proprietary that when it finally broke—and everything breaks—the MTTR was measured in weeks, not hours. The availability was abysmal, and the production loss was a punch to the gut.
The Classic Field Example: The Pump Skid
Picture a standard pump skid at a water treatment plant. You’ve got two choices for the main pump:
- Option A: A heavy-duty, slow-speed positive displacement pump. Built like a tank. MTBF is 10 years. But when it fails, you need a crane to lift the thing, and the rebuild kit takes 12 weeks to arrive from Germany. MTTR is two weeks.
- Option B: A standard centrifugal pump. Less tough. The seals might go every 18 months. MTBF is 1.5 years. But you can swap the whole pump out with a spare from the shelf in four hours. MTTR is four hours.
Let’s run the numbers. For Option A, availability is 10 years / (10 years + 2 weeks) = 99.6%. For Option B, availability is 1.5 years / (1.5 years + 4 hours) = 99.97%. The “less reliable” pump is actually more available. In a critical process where downtime costs $10,000 an hour, the choice is a no-brainer. The field engineer’s job is to make this math painfully clear to the project manager who’s only staring at the MTBF number on the spec sheet.

Why MTTR Is the Field Engineer’s Real Metric
Reliability is a design and manufacturing metric. It’s locked into the equipment before it ever hits your loading dock. As a field engineer, you have almost zero control over the MTBF of a pump, a drive, or a transmitter. What you can control is the MTTR. This is where practical, resource-conscious decisions turn a constant headache into a system that’s actually available.
Reducing MTTR isn’t about throwing money at the problem. It’s about smart prep and standardization. Here are the levers you can pull:
- Standardized Spares: Using the same model of VFD across a dozen applications means you stock one spare that covers all of them. This is a fight worth having with procurement, who often want to buy the cheapest option for each individual job.
- Access and Labeling: A reliable machine buried behind a rat’s nest of unlabeled conduit is an unavailable machine. Clear labeling, disconnect switches you can actually reach, and pull space for cables aren’t luxuries; they’re availability multipliers.
- Proactive Diagnostics: A drive that gives you a clear fault code like “Overcurrent Phase B” gets back online faster than one that just blinks a cryptic red LED. Specifying equipment with a good local HMI and diagnostic capabilities is a direct investment in availability.
- Documentation That Lives in the Panel: A laminated single-line diagram and a list of the last five fault codes stored inside the panel door can save an hour of head-scratching. That’s an hour of uptime.
The Hidden Trap: High Reliability Can Mask Poor Availability
There’s a nasty psychological trap I’ve seen snag plenty of teams. When a system is extremely reliable, it can hum along for years without a hiccup. This breeds a false sense of security. The documentation rots. The spare parts budget gets slashed. The old-timers who knew the system’s quirks retire or move on. Then, when the inevitable failure hits, the MTTR is astronomical because nobody’s ready. The system’s high reliability has directly caused its low availability during a crisis.
This is a classic story in municipal water systems and older power plants. A transformer that hasn’t been touched in 30 years finally gives up, and the crew discovers the replacement is a six-month lead-time item that needs different bushings and a new concrete pad. The design was reliable, but the lack of planning for the eventual failure made the system unavailable for months.

Designing for Availability in Brownfield Projects
Most of our work doesn’t start on a clean sheet of paper. It’s in existing facilities with legacy equipment, tight budgets, and a production schedule that doesn’t stop. In these brownfield environments, chasing reliability is often a fool’s errand. You can’t just swap a 1970s motor for a modern, high-efficiency unit without also replacing the coupling, the baseplate, and the starter—and that’s a capital project, not a maintenance task.
Instead, zero in on availability. Here’s a practical approach I’ve used in the field:
- Map the failure modes of the existing system. What actually breaks? Don’t guess. Pull the last two years of work orders and find the top three failure points.
- For each failure mode, ask: can we cut the time to restore? This might mean pre-fabricating jumper cables for a common fault, stocking a specific bearing kit, or simply adding a local HMI so the operator can see the fault without calling for a technician.
- Calculate the cost of downtime per hour. This number is your ammunition. If an hour of downtime costs $5,000, then spending $500 on a spare sensor that reduces MTTR by one hour pays for itself the first time it’s used.
This approach is resource-conscious. It doesn’t beg for a new system. It asks for targeted improvements that directly boost the percentage of uptime.
Bridging the Engineering Culture Gap
In my experience working with both Brazilian and Anglo-American engineering teams, there’s a subtle cultural difference in how these concepts are prioritized. In many Brazilian industrial settings, the focus is intensely practical: “O que faz a planta rodar?” (What keeps the plant running?). There’s a natural inclination toward availability, often driven by the reality of longer supply chains and the need for creative, on-the-spot solutions. The concept of gambiarra—a makeshift fix—is born from this necessity. It’s a solution that prioritizes restoring function now, even if it’s not a permanent, reliable fix.
In more structured engineering environments, the focus often leans heavily toward reliability, with rigorous adherence to manufacturer specifications and planned maintenance schedules. The strength here is consistency and long-term asset life. The weakness can be a lack of flexibility when the planned process fails.
The best field engineering practice bridges these two worlds. It respects the need for a reliable, standards-based design but never loses sight of the practical, on-the-ground reality of keeping a system available. It means designing a panel with a bypass for every critical safety relay, not because you plan to use it, but because you know that at 3 a.m. on a Sunday, someone might need to. It’s about building resilience into the system by planning for failure, not just trying to prevent it.
Practical Steps to Audit Your System’s Availability
You don’t need fancy software to get a handle on this. Here’s a simple, paper-and-pencil audit you can do on your next site visit:
- List your top 5 critical assets. These are the ones whose failure stops the process dead.
- For each asset, answer three questions:
- What is the most likely failure mode? (Be specific: not “pump fails,” but “mechanical seal leaks.”)
- What is the current MTTR for that failure? (Be honest. Include travel time, diagnosis, and waiting for parts.)
- What single change would most reduce that MTTR? (A spare part on the shelf? A better procedure? Training for the local operator?)
- Calculate the current availability for each asset using the formula. Compare it to the reliability number on the datasheet. The gap is your opportunity.
This exercise often reveals that the biggest gains aren’t in buying more reliable equipment, but in making the existing equipment more maintainable. A $50,000 pump with a two-week MTTR might have lower availability than a $15,000 pump with a four-hour MTTR. The field engineer’s value is in presenting this analysis clearly, with numbers, to decision-makers who are often focused only on the purchase price and the MTBF.
Frequently Asked Questions
What is the main difference between reliability and availability?
Reliability measures how long a system can run without failing. It’s about the frequency of failures. Availability measures the percentage of time a system is ready to perform its function. It considers both how often it fails and how quickly it can be repaired. A system can be very reliable but have low availability if repairs take a long time, or it can be less reliable but have high availability if repairs are very fast.
Why is availability often more important than reliability in a production environment?
In a production environment, the cost of downtime is usually the dominant factor. Availability directly measures uptime, which is what generates revenue. A system that fails frequently but can be fixed in minutes may have a higher availability—and thus cause less production loss—than a system that fails rarely but takes weeks to repair. The focus should be on minimizing the total downtime over a year, not just the number of failures.
How can I improve availability without buying new, more reliable equipment?
Focus on reducing the Mean Time To Repair (MTTR). This can be done by: stocking critical spares on-site, improving diagnostic information (like better fault codes on drives), creating clear and accessible documentation, training operators to perform first-line troubleshooting, and standardizing components across multiple machines to simplify spare parts management. These are often low-cost changes that have a high impact on uptime.
What is a realistic availability target for a typical industrial system?
It depends heavily on the process, but many industries target 99.5% to 99.9% availability for critical systems. However, a raw percentage can be misleading. 99.9% availability still allows for almost 9 hours of downtime per year. The target should be set based on the cost of downtime and the practical limits of your maintenance resources. Chasing an extra “nine” of availability (e.g., from 99.9% to 99.99%) often requires exponentially more investment.