A packaging manufacturer running plastic extrusion machines called us in with a problem that was costing them millions.

Power would blip. Not go out. Just blip. A fraction of a second. But their equipment had zero backup protection, so every blip meant every extrusion machine on the line stopped mid-run. Plastic sets up fast. Every stoppage meant hours of cleanout, scrapped material, and lost production.

I inspected the power lines more times than I can count. Nothing. We modeled out rerouting their service onto a different circuit entirely. Solid plan. Didn’t solve the actual problem. I brought in our reliability team to walk the customer’s side, look at what protections they could add downstream. We met on-site three times.

Here’s the part that’s easy to forget when you’re the one holding the flashlight: I didn’t have the answer for six weeks. Nobody did. The cause turned out to be a single post-op switch on the circuit, slowly failing. It finally gave out completely. Once it was replaced, the problem was gone.

No dramatic save. No single moment of insight. Just staying with a problem that was expensive and frustrating for the customer, bringing in the right expertise when mine wasn’t enough, and not walking away until the answer showed up.

That’s what reliability actually costs a business. Not just the outage. The weeks of not knowing. It’s why I think about operational data the way I do: the goal isn’t a dashboard that looks right. It’s understanding well enough to know when something’s actually wrong, and staying with it until you do.