How to Debug the Failure You Cannot Reproduce: Treating Field Intermittent Faults Like a Story You Have to Plot Backwards

By | Jul 13, 2026

Last year a client called me about a textile plant in the interior of Minas Gerais. Twelve STM32-based nodes, Modbus RTU over RS-485, roughly two hundred meters of bus running across the factory floor. Fourteen months of clean operation. Then October hit, and the bus started dropping nodes every afternoon between two and four. By five, everything worked again. The maintenance crew had already replaced transceivers on three nodes, swapped the master’s power supply, re-crimped every connector they could reach. Nothing stuck. They called me because they were out of ideas and nearly out of budget.

I will get to what we did, but first: this kind of failure breaks the normal debugging workflow. The tools we reach for first — oscilloscope, logic analyzer — are often the wrong tools for intermittent, environment-dependent faults. Not because they cannot capture the problem. Because they cannot tell you when to capture it, or what the problem actually is.

The Problem With Reproducing the Unreproducible

Most embedded debugging training assumes you can reproduce the fault. Board on the bench. Run the code. It crashes. Attach a debugger. Read the registers. Form a hypothesis. Test it. Fix it. That loop works because the failure is deterministic enough to repeat under controlled conditions.

Field intermittent faults do not cooperate. The RS-485 bus in Minas Gerais worked perfectly at nine in the morning. Worked perfectly at seven in the evening. Failed for two hours every afternoon — but only on weekdays, and only when the plant ran at full production. I could not reproduce this on my bench in São Paulo because my bench does not have a textile plant attached to it.

That is the fundamental problem. The failure is a function of environmental conditions you do not control and cannot fully replicate. You can measure some of them — temperature, humidity, line voltage — but not all of them simultaneously, and certainly not while sitting at a desk with a debugger attached.

So what do you do? What every embedded engineer does eventually: you start guessing. Swap parts. Add capacitors. Change the baud rate. Try different termination resistor values. Each guess is a hypothesis, but an unstructured one, and because the failure takes hours or days to manifest, you cannot test hypotheses quickly. You end up in a circular loop — change something, wait, see if it fails, change something else, wait, see if it fails. Two weeks later you have lost track of what you changed and why, and you are no closer to the root cause.

Why Your Oscilloscope Is Not Enough

First site visit, I brought my oscilloscope. A four-channel Rigol, nothing fancy but adequate for RS-485 signal integrity work. Clipped it onto the bus at the master node and waited for the afternoon failure window.

Here is what I saw. The differential signal looked clean at one in the afternoon. By two-thirty, slightly attenuated — maybe 200 mV less amplitude, nothing dramatic. By three o’clock, degraded enough that some nodes were missing frames. By three-thirty, the bus was effectively dead.

The oscilloscope told me what was happening. It did not tell me why. The signal attenuation was a symptom, not a cause. I could see the effect of something changing on the bus, but I could not see what was driving the change. Temperature? Ground potential shift? EMI from a specific machine turning on at a specific time? All three at once?

The oscilloscope is a snapshot tool. It shows you a moment in time. But intermittent field failures are stories that unfold over hours, and a snapshot cannot capture a narrative. You need a different kind of tool — not a hardware tool, but a thinking tool.

Building a Debugging Plot Document

After the first site visit, I sat down and wrote out the problem as if I were outlining a mystery novel. Not because I was being poetic. Because I realized I needed structure. I had a pocket full of measurements, a head full of half-formed hypotheses, and no coherent way to organize either. The narrative structure forced me to be precise about what I knew, what I did not know, and what I was assuming.

Here is the structure I used — and the one I now use for every intermittent field debug:

The Inciting Incident (Problem Statement): What exactly fails, when, and under what conditions. Not a vague description like “the bus is unstable.” A precise statement: “Nodes 4, 7, and 11 drop off the RS-485 bus between 14:00 and 16:00 on weekdays when ambient temperature exceeds 38°C and production line 3 is operating.”

Rising Action (Evidence Chain): Every observation I have made, in chronological order, with timestamps and environmental conditions noted. This includes the things that did not change as well as the things that did. The absence of a change is evidence too. If the power supply voltage held rock-steady at 5.02V throughout the failure window, that fact eliminates an entire class of hypotheses.

The Climax (Root Cause Hypothesis): The single explanation that accounts for every piece of evidence in the rising action section. Not three possible explanations — one. If you have three, you do not have a root cause yet. You have a list of suspects. The climax is the moment where you commit to a narrative that connects all the evidence.

Resolution (Fix and Verification): What you changed, why you changed it, and how you verified that the fix addressed the root cause and not just the symptom. This section also includes what you deliberately did not change, and why.

This is not a postmortem written after the fact. It is a living document you update during the investigation. The act of writing it forces you to confront gaps in your evidence. When you try to write the rising action section and realize you do not know whether the factory’s HVAC system runs during the failure window, that gap becomes your next data collection task. The document drives the investigation, not the other way around.

The Minas Gerais Bus: A Worked Example

Let me walk through the actual plot document for the textile plant case, because the structure is easier to understand with real content.

Inciting Incident: RS-485 bus failures at a textile plant in Minas Gerais. Nodes 4, 7, and 11 — all in the eastern half of the bus, physically closest to production line 3 — stop responding between approximately 14:00 and 16:00 on weekdays. The master continues to see nodes 1, 2, 3, 5, 6, 8, 9, 10, and 12 throughout the failure window. The system recovers without intervention by 16:30.

Rising Action — Visit 1: Measured differential signal at master node. Signal amplitude at 13:00: 2.8V differential (normal for SN65HVD75 at 5V). At 14:30: 2.6V. At 15:00: 2.1V. At 15:30: 1.4V, with visible distortion on falling edges. At 16:00: 1.1V, multiple nodes not responding. At 16:30: 2.7V, all nodes responding.

Ambient temperature at bus midpoint: 31°C at 13:00, 39°C at 15:00, 37°C at 16:00. Power supply at master: 5.02V throughout. Power supply at node 7: 4.98V throughout. I did not measure ground potential difference between nodes — that gap became the next task.

Rising Action — Visit 2: Measured ground potential between master node and node 7 using a battery-powered multimeter. At 13:00: 0.3V difference. At 15:00: 1.8V difference. At 16:30: 0.4V difference. The ground potential was shifting during the failure window by nearly 2V.

Why? The bus cable ran through a cable tray that also carried three-phase power feeds to production line 3. Line 3 ran continuously from 14:00 to 16:00 as part of the afternoon production schedule. The cable tray was a steel ladder tray, and the RS-485 cable was unshielded twisted pair — the cheapest the contractor could find at the local electrical supply shop.

Climax: The ground potential shift was caused by inductive coupling from the three-phase power cables in the shared cable tray, induced when production line 3 drew full load current during the afternoon shift. The unshielded twisted pair provided inadequate common-mode rejection for the induced noise. The SN65HVD75 transceivers have a common-mode input range of -7V to +12V, but the common-mode noise was causing the receiver differential threshold to be exceeded intermittently as the induced voltage varied with load current fluctuations. Nodes 4, 7, and 11 were closest to the section of cable tray where the power cables and RS-485 cable ran parallel for approximately forty meters.

Resolution: Routed the RS-485 cable through a separate, dedicated conduit with 600mm physical separation from the power cable tray. Used shielded twisted pair (Belden 9841, which the client sourced from a supplier in Belo Horizonte). Connected the shield to earth ground at the master node only. After the change, the ground potential difference between master and node 7 during full production was 0.2V. Signal amplitude held at 2.8V differential throughout the afternoon window. The system has been running for nine months without a recurrence.

Notice what the plot document does that the oscilloscope capture does not. It connects the evidence into a causal chain. The oscilloscope showed signal attenuation. The plot document explains why the signal attenuates at that specific time, in that specific location, under those specific conditions. The narrative is the debugging tool. The oscilloscope is just a sensor.

What the SRE World Already Knows About This

This approach is not something I invented. Large-scale production engineering teams have been using structured narrative documentation for incident response for years. The Google SRE book dedicates entire chapters to effective troubleshooting methodology and postmortem culture, and includes example incident state documents and postmortem templates in its appendices that look remarkably like what I am calling a plot document — structured accounts of what happened, what evidence was gathered, what hypotheses were tested, and what the root cause turned out to be. You can read the full table of contents and see how much of the book is organized around structured narrative documentation of failures at the Google SRE book index. The gap is not that this methodology does not exist — it is that it has not crossed over into embedded field engineering, where engineers are more likely to reach for a soldering iron than a blank document.

The cultural gap between datacenter SRE practices and embedded field engineering is real but bridgeable. SRE teams write postmortems because they learned that without a written narrative, institutional knowledge about failure modes dies with the engineer who diagnosed it. Embedded teams need the same discipline, but our postmortems look different — they include cable routing diagrams, photographs of corroded connectors, and temperature logs from inside sealed enclosures. The format is different. The principle is the same: narrate the failure as a structured document so that the next engineer who faces it does not start from zero.

When You Are Stuck in the Loop

There is a specific failure mode in debugging that I want to address because it wastes the most time: the circular hypothesis loop. You have three suspects. You test one. It does not conclusively fail or pass. You move to the next. Also inconclusive. You go back to the first. Add a variable. Test again. Days pass. You are not converging on a root cause — you are orbiting one.

This happens because your evidence chain has a gap you cannot see from inside the loop. You need an external force to push you toward a narrative you have not considered. When I am stuck like this, I have started doing something that sounds odd but works: I feed my evidence into a story plot generator to force a narrative arc I would not write myself. The system becomes the protagonist, the failure mode becomes the conflict, the stakes become what breaks if I am wrong about the cause. The generated plot is never the answer — but it reframes the evidence from a perspective I was too close to see, and that reframing has surfaced blind spots more than once.

The same principle applies to the Reedsy plot generator, which formalizes narrative structures like the 3-Act and 7-Point frameworks that decompose a story into inciting incident, rising action, climax, and resolution. You can find their plot generator here and try feeding it your debugging scenario as if it were a mystery novel. The exercise is not frivolous — it forces you to articulate your protagonist (the system), the conflict (the failure mode), and the stakes (what breaks if you are wrong) in terms a narrative engine can process. If you cannot articulate these clearly enough for a plot generator to produce a coherent story, you probably cannot articulate them clearly enough to debug the problem either.

That same discipline applies to narrative structure: before publishing, editors need a way to test events, claims, and consequences actually follow one another, which is where a story plot generator that fits the project can function as a planning aid rather than a substitute for domain evidence.

The Discipline of Writing What You Do Not Know

The hardest part of the plot document is the rising action section, because it requires you to write down what you do not know. Most engineering documentation captures what was done and what was found. The plot document must also capture what was not measured, what was not tested, and what assumptions were not validated. These negative spaces are where the root cause hides.

In the Minas Gerais case, the gap that broke the investigation open was the ground potential measurement. I did not make it on the first visit because I was focused on signal integrity — the oscilloscope was already clipped to the data lines, and it felt like the right place to look. It was the right place to look for symptoms. It was the wrong place to look for causes.

When I wrote the rising action section after the first visit, I had to write: “Ground potential difference between nodes: not measured.” That sentence stared at me from the page for two days before I went back with a multimeter and a specific plan to measure it. The document told me what I was missing. My oscilloscope could not do that.

This is why the narrative structure matters more than the tools. A four-channel oscilloscope with auto-measure and FFT is useless if you do not know where to clip the probes. A logic analyzer with protocol decode is useless if you do not know when to trigger. The plot document tells you where and when, because it forces you to map the full territory of the failure before you start measuring.

Practical Takeaways for Your Next Field Debug

Start the plot document before you visit the site. Write the inciting incident from the information you have over the phone or by email. Be precise about timing, environmental conditions, and which specific nodes or subsystems are affected. Vague problem statements produce vague investigations.

During the site visit, update the rising action section in real time. Note timestamps, temperatures, what is running, what is not running, and what you are measuring versus what you are ignoring. The act of writing while measuring changes what you measure — you start looking for things that will fill gaps in the narrative, not just things that are easy to probe.

After each visit, read the rising action section and list every gap explicitly. “Did not measure X.” “Did not check if Y was running.” “Assumed Z was stable but did not verify.” These gaps are your task list for the next visit.

Commit to a single root cause hypothesis in the climax section before you implement a fix. If you cannot write one sentence that explains every piece of evidence, you are not ready to fix anything. You are still collecting evidence. Go back to rising action.

Finally, keep the document after the investigation closes. Store it alongside the schematics, the firmware revision history, and the cable routing diagrams. The next engineer who inherits this system will not have your memory of the investigation. They will have the document. Make sure it tells the whole story — not just the ending.