Select Page

A panel PC that shuts down during a heatwave. A flash memory chip that wears out silently. A power connector that intermittently loses contact inside a marine console. These aren’t headline-grabbing failures, but they do stop production, cause delays and cost money.

Over the years we have investigated, diagnosed and resolved a wide range of industrial hardware failures. The causes are rarely exotic. They are almost always a mismatch between the equipment and the real-world operating conditions, or a subtle configuration detail that was easy to miss during commissioning.

Here are some of the most important lessons we have learned, and how to avoid making the same mistakes!

Article: Lessons from the Field

Thermal failure in a kiosk in full sun

A customer in South Australia installed an industrial monitor inside an outdoor kiosk. The monitor was rated for demanding environments. But the kiosk was positioned in full sun.

After a few years, the LCD screen developed permanent black patches. When we consulted the factory engineers, their analysis was blunt: the LCD panel would need to reach temperatures above 100 degrees Celsius for that type of damage to occur. Inside a sealed kiosk in direct sunlight, with the heat generated by the monitor and PC adding to the solar load, that threshold had been exceeded.

The factory engineers made several recommendations to improve the installation, including increasing the air space around the monitor and adding forced-air cooling inside the enclosure. But they were not confident enough in these measures to warrant the product against recurrence. The most effective solution was the simplest – install the kiosk under shade to eliminate the direct solar heating that pushed internal temperatures beyond what electronics can survive.

The lesson is not about the monitor. It’s about the difference between the operating temperature of a component and the actual temperature inside its enclosure. A monitor rated to 50 degrees will not survive in an enclosure that reaches 100 degrees, regardless of what the datasheet says. The installation environment is part of the specification.

A command structure difference that killed flash modules

A customer deployed several Advantech ADAM I/O modules on a distributed RS-485 network across their factory. The system was configured, tested and worked perfectly. Then after seven to eight months, modules started failing one after another.

All the failed modules came back from factory repair with the same diagnosis: internal flash memory chips had worn out and needed replacing. Several modules failing at roughly the same time pointed to a systematic cause rather than random component failure.

Our engineers investigated the command protocols being used. The ADAM modules have two very similar commands for writing data to outputs. One writes the output value and also commits it to flash memory. The other writes the output value, without touching flash. The difference in the command structure is subtle. The customer’s control software was using the flash-writing version for every output command, hundreds of times per day. Flash memory chips in these modules are designed for storing configuration settings. They are typically rated for 100,000 to 300,000 write cycles over their lifetime. At hundreds of writes per day, that lifetime was being consumed in months rather than years.

Once identified, the fix was straightforward: change to the non-flash-writing command variant. The issue never recurred. But without diagnosing the root cause, the customer would have been replacing modules indefinitely, at significant cost and with repeated production disruption.

The lesson: component wear is not always gradual or visible. Understanding how your software interacts with hardware at the protocol level matters, particularly for devices with finite write endurance.

Power damage the most common failure we see

Power-related damage is the single most frequent issue on units returned to us, and it takes several forms.

  • Wrong input voltage. Many industrial PCs are available with either a standard 12V DC input or a 9-36V DC wide-range input. The wide-range version includes an internal DC-DC converter and costs more. We have seen multiple cases where an integrator ordered the standard 12V model, then connected it to a 24V industrial power supply. The result is immediate damage. Our sales team now checks the customer’s intended power source before quoting every order, because this mistake is easy to make and expensive to fix.
  • Surge damage from shared power supplies. We have supplied a large number of panel PCs to utility sites across Queensland, installed in MCC switchboard doors running SCADA software. Despite specifying units with 9-36V DC wide-range input, we had a small number of failures where the internal DC-DC converter was damaged by voltage surges. The likely cause was large inductive loads, such as pumps and motors, cycling on and off on the same power supply. A wide-range input handles voltage variation, but it doesn’t necessarily protect against sharp transient spikes from heavy equipment. Additional surge protection at the supply side is worth considering wherever a PC shares power with industrial machinery.
  • Lightning and data line surges. We have also seen damage where neither the power supply nor the data lines were properly protected against lightning-induced surges. At remote or exposed sites, surge protection on both power and communications lines should be part of the standard installation specification.

The lesson: power is the foundation. If the power supply and protection are not right, nothing else matters. It’s also the easiest problem to prevent with the right advice at specification time.

When industrial components don’t behave as expected

We built a batch of 1RU rackmount industrial PCs for installation in operator consoles on marine vessels. The specification was robust: industrial-grade Advantech motherboard, high-quality Zippy industrial power supply, built to handle the demanding conditions of a shipboard environment.

On deployment, the customer reported intermittent boot failures. Investigation traced the problem to the ATX 24-pin power connector from the Zippy power supply not making reliable contact with the motherboard socket. This was surprising given the quality of both components, but connector tolerances can vary even between reputable manufacturers, and in a marine environment subject to vibration and movement, a marginal connection becomes an intermittent failure.

The project timeline ruled out a warranty process with Zippy, so the customer opted to replace the power supplies with units from Bicker, a German industrial power supply manufacturer. The replacement units resolved the problem completely.

The same 1RU marine project produced a second unexpected issue. The tight chassis required a ribbon cable PCIe x16 riser card to connect a graphics card to the motherboard. During testing, we experienced random instability related to the graphics card. Investigation revealed that the ribbon riser cards were designed for PCIe Gen 2, while the motherboard and graphics card supported Gen 3. The higher signalling speed of Gen 3 was unreliable through the ribbon cable. Restricting the motherboard to PCIe Gen 2 in the BIOS eliminated the instability entirely.

The lesson: even high-quality industrial components can produce unexpected failures when combined. Testing the complete assembled system under realistic conditions catches issues that no amount of datasheet review will reveal.

The common thread

These failures have little in common on the surface. A kiosk in the sun, flash memory commands, power supply voltages, connector tolerances, PCIe signalling speeds. But they share an underlying theme: the failure was not caused by defective hardware. It was caused by a gap between the assumptions made during specification or installation and the actual operating conditions.

In every case, the hardware was doing exactly what it was designed to do. It was the environment, the configuration or the system integration that created the problem. And in almost every case, the failure could have been prevented with the right information at the right time.

How ESIS can help

We have seen these failure patterns across many industrial installations over the years. That experience informs how we specify hardware, how we advise on installation and how we test before equipment leaves our workshop.

If you are specifying industrial hardware for a new project, or troubleshooting an issue with existing equipment, we can help you avoid the common pitfalls and get the installation right the first time.

Contact us to discuss your project details with one of our experts.

Call Now Button