Blog
A train door closing a fraction slower, every week, for three months
Predictive maintenance does not start with a model. It starts with measuring something nobody was measuring.
Published
One train door starts taking about two tenths of a second longer to close. No passenger notices. No technician records it, because everything still works.
Three months later the delay has grown until the door will not close at all. The train is held at a platform, passengers are moved, and the whole line runs fifteen minutes behind during the peak.
Had anyone seen the trend in the first month, the repair would have taken forty minutes in the depot overnight.
Two common approaches, and the gap between them
Time-based maintenance replaces parts every so many kilometres or months. It is safe and wasteful, because a great many parts are replaced with plenty of life left.
Run-to-failure waits for the problem and then fixes it. Cheap on paper, but the real cost is the service interruption — and in public transport that cost does not fall on the operator alone. It falls on thousands of passengers.
Predictive maintenance sits between the two: replace when the data says failure is coming, rather than when the calendar says so or after the fact.
What has to exist before a model does
The most common mistake is starting with a model when there is no data.
Predictive maintenance needs three things, in order. First, continuous measurement of something that reflects the equipment's condition: motor current, vibration, temperature, the time taken for one operating cycle.
Second, maintenance history recorded in a form a machine can read. A free-text note saying "replaced drive unit" is worth far less than a structured record naming the asset, the part, the symptom, and the date.
Third, and only third, a model — which will be exactly as good as the first two are.
Many organisations spend an entire first year on the first two. That is the correct use of the year.
Start with whatever fails most
There is no need to instrument everything. Pick one class of equipment that exists in quantity and fails often: train doors, escalators, station air conditioning.
Quantity matters, because dozens of identical assets can be compared against each other. The one behaving differently from its peers is the one to go and look at — and that analysis works before there is any model at all.
Measuring whether it worked
The metric is not prediction accuracy. It is the number of service interruptions, and the proportion of maintenance done in a planned window against work done as an emergency.
Emergency work costs several times what planned work costs — not in parts, but in people, in hours, and in whatever else was planned and had to give way.
What GIPSIC does here
We do IoT work for public transport, from designing hardware that survives vibration and electrical noise through to back-end systems that hold a great deal of telemetry and let it be searched back.
Our hardware and software people sit together, which matters here: deciding what to measure and at what frequency is a decision about circuit design and database design at the same moment.
If you would like to talk about the equipment in your system, get in touch.
About the author
Film — Wisit. A businessman who still does his own BA work more often than he probably should, and writes a fair bit of code, front and back. Runs two or three small businesses. Follows technology and business obsessively, in Thailand and everywhere else. Off the clock: physics, astronomy, DIY, and anything to do with networks. Music always on, though he cannot sing. Plays instruments anyway, badly. Plays a lot of sport, racket sports above all. Not much of a traveller by himself, but happy to take Mint anywhere in the world. A man who fears — sorry, loves — his wife. One flaw: he barely touches video games.
Written with Claude Opus 5