Hardware. Firmware. Software. One engineering partner.Based in India · Working worldwide   Deutsch ↗
Industrial IoT

Predictive maintenance: what it actually takes to get there

Prediction is earned rather than bought. What condition monitoring delivers first, why failure history is the constraint, and how to run a programme people still trust after a year.

In short

  • Prediction requires recorded failure history. Most operations do not have it, because failures are rare and nobody was measuring when they happened.
  • Condition monitoring comes first. It needs only a baseline, delivers value in months, and builds the history prediction later depends on.
  • The P-F interval decides how often you must measure. Monthly checks are useless if warning comes two weeks ahead.
  • Machine learning is often unnecessary. A well-chosen measurement with a sensible threshold catches a great deal and is far easier to trust.
  • Programmes usually fail on false alarms and unread alerts, not on the analysis.

The gap between the request and the reality

The request is almost always some version of “tell us before it breaks”. It is entirely reasonable. Unplanned downtime is expensive and disruptive, scheduled replacement throws away components with life left in them, and the promise of knowing in advance is obviously attractive.

The difficulty is that predicting a specific failure, far enough ahead to act, generally requires having seen that failure develop before and having recorded the signals that preceded it. Most operations have neither, for an understandable reason: failures are rare, and when they happened nobody was measuring the things that would have revealed them.

A programme that promises prediction from data that does not exist will disappoint, and the damage outlasts the project. Once a plant has been sold prediction and received false alarms, the next proposal is very much harder to get approved.

The good news is that the intermediate step is genuinely valuable on its own, and it is the thing that makes prediction possible later.

Four strategies, and what each asks of you

Four maintenance strategies compared: reactive requires nothing but gives no warning; preventive requires a schedule but discards usable component life; condition-based requires measurement and a baseline and is the realistic starting point; predictive additionally requires recorded failure history.
Each step right buys planning ability and costs something in what you have to have in place first.

Reactive means running equipment until it stops. It is not always wrong: for a cheap component that is quick to replace and whose failure causes no knock-on damage, waiting is often the economically correct choice. It becomes indefensible when failure is expensive, dangerous or disruptive.

Preventive means replacing on a schedule, whether or not the part needs it. This trades component life for predictability. Its weakness is well documented: a substantial share of failures are not age-related, so a calendar catches some and misses others while consuming parts that had years remaining. It also introduces risk of its own, since intervening on a working machine occasionally breaks it.

Condition-based means measuring the equipment and acting when its behaviour changes. This is where most operations should start, because it requires only measurement and a baseline, not a history of failures.

Predictive adds an estimate of when, rather than only that something has changed. It requires everything condition-based requires, plus enough recorded history to relate what was observed to what subsequently happened.

The P-F interval: the number that sizes your programme

This concept comes from reliability engineering and it is the single most useful idea for anyone designing a monitoring system.

A curve showing equipment condition declining over time from normal through a point of potential failure to functional failure, with vibration and ultrasonic detecting earliest, then oil analysis, then thermal, then audible noise, then heat and smoke.
Different techniques detect at different points on the decline. The earlier the detection, the more planning time it buys.

Equipment rarely fails instantaneously. Something begins to degrade, and from that point the condition declines until the machine can no longer do its job. The point at which degradation first becomes detectable is called potential failure, and the point at which the equipment stops functioning is functional failure. The gap between them is the P-F interval.

Two consequences follow, and both are practical.

Measurement frequency must be well inside the interval. If a bearing typically gives six weeks of warning in vibration, monthly readings might catch it once and miss it the next time. A common rule of thumb is to measure at least twice within the shortest interval you care about, which for many rotating assets pushes you toward continuous monitoring rather than periodic walk-round surveys.

Technique determines how early you see it. Vibration and ultrasonic methods detect degradation much earlier than heat, noise or anything a person notices in passing. By the time a bearing is audibly noisy, most of the interval has gone. This is why continuous vibration monitoring is worth its cost on assets where the consequence of failure is high, and why relying on operator observation gives you very little time.

What gets measured

The measurements themselves are covered in detail in our guide to retrofitting legacy machines, which also sets out mounting, installation and the practicalities of getting data off equipment that was never designed to provide it. In summary, for rotating equipment the useful signals are vibration for mechanical condition, supply current for load and developing resistance, temperature and its trend relative to load, acoustic signature including ultrasonic ranges, and cycle timing for how the asset is actually used.

Which of these matters depends on the asset and on the failure modes you are trying to catch, which is why the question “what should we measure” cannot be answered before the question “what fails, and what does it cost”.

Techniques beyond vibration

Vibration gets most of the attention, and for rotating machinery it deserves it. But it is one technique among several, and the others catch things it does not. A programme built on vibration alone has blind spots that are entirely avoidable. Which technique suits which machine is set out in condition monitoring for motors, pumps, compressors and fans.

Oil analysis

For gearboxes, hydraulic systems and large rotating equipment with an oil circuit, the lubricant carries a detailed record of what is happening inside. Wear particles indicate which components are degrading and how fast, and their shape and composition can distinguish gear wear from bearing wear. Contamination shows ingress of water or dirt, either of which shortens equipment life considerably. Additive depletion and viscosity change show when the oil itself has stopped doing its job.

The appeal is that it sees inside the machine without opening it, and it often detects degradation at a similar stage to vibration while identifying the mechanism more specifically. The practical constraint is that it is usually a sampling exercise rather than a continuous measurement: someone draws a sample, a laboratory analyses it, and results come back days later. Inline sensors exist for particle counting and water content, and on high-value equipment they are worth considering, but for most assets oil analysis sits alongside continuous monitoring rather than replacing it.

Thermography

An infrared camera sees temperature differences, and temperature differences reveal a specific set of problems well. Loose or corroded electrical connections heat up under load, and they are among the most common causes of both unplanned outages and fires in industrial installations. Overheating bearings, blocked cooling paths, failing insulation, blocked steam traps and uneven heating across a process all show clearly.

For electrical problems thermography is frequently the single best technique available, and an annual survey of switchgear and distribution boards under load is among the highest-value inspections most plants can do. For mechanical degradation it detects later than vibration, since by the time a bearing is measurably hot the damage is well advanced. Use it for what it is good at rather than as a general-purpose condition check.

Ultrasonic

Ultrasonic instruments detect sound above the range of human hearing, and two quite different applications matter industrially.

Airborne ultrasound finds compressed air and gas leaks, which are invisible, continuous and expensive, and it detects electrical arcing and corona discharge in switchgear. For plants running compressed air, a leak survey usually pays for itself quickly, and the losses found are often larger than anyone expected.

Structure-borne ultrasound detects the very early stages of bearing degradation and inadequate lubrication, often earlier than conventional vibration analysis. It is also used to guide regreasing, which matters because over-greasing damages bearings as reliably as under-greasing does.

Techniques compared by what they detect and how early.
Technique Detects How early Continuous or periodic
Vibration Bearing, imbalance, misalignment, looseness, gear faults Early Either; continuous on critical assets
Ultrasonic, structure-borne Very early bearing wear, lubrication problems Earliest for bearings Usually periodic
Ultrasonic, airborne Air and gas leaks, electrical arcing Immediate when present Periodic survey
Oil analysis Wear mechanism, contamination, lubricant condition Early Periodic sampling; inline on high-value assets
Thermography Electrical connections, overheating, blocked cooling Early for electrical, late for mechanical Periodic survey
Motor current Load, developing resistance, rotor and eccentricity faults Moderate Continuous, easily
Process values Efficiency loss, output drift, blockage Varies Continuous where a controller exists

Establishing what normal looks like

This step is routinely underestimated and it is where a great many programmes acquire their credibility problem.

Two nominally identical machines, same model, same building, same duty, will sit at different vibration levels. Mounting differs, foundations differ, alignment history differs, age differs. A threshold taken from a standard or from a supplier’s default will be too sensitive for one and too permissive for the other, producing constant nuisance alarms on the first and silence on the second.

So the baseline has to be per machine, and it has to account for operating context. A machine running a heavy product draws more current and vibrates differently from the same machine running a light one. A baseline that ignores this will flag every product changeover as an anomaly. Capturing load, product and speed alongside the condition signals is what allows the comparison to be like for like.

Establishing a baseline takes time, because it must cover the normal variation: different products, different shifts, different operators, and ideally different seasons, since ambient temperature affects almost everything.

How detection actually works

There is a spectrum of approaches, and the sophisticated end is not automatically the right end.

Detection approaches, in increasing order of complexity.
Approach How it works Needs Best when
Fixed threshold Alert when a value crosses a set level A sensible level The signal is clear and the safe range is known
Per-machine threshold Level set from that machine’s own baseline Baseline period Machines differ, which is nearly always
Trend and rate of change Alert on how fast something is moving, not only where it is History Slow degradation that will cross a limit eventually
Statistical anomaly detection Flag behaviour outside the normal distribution for current conditions Baseline plus context data Normal varies with load or product
Machine learning, unsupervised Learn the shape of normal across many variables and flag departures Substantial normal data Many interacting signals; no failure examples
Machine learning, supervised Learn to recognise specific failure modes Labelled examples of those failures Failures are recorded and recur

Most working programmes spend most of their time in the first four rows. That is not a failure of ambition; it is that a well-chosen measurement with a sensible per-machine level catches a great deal of what matters, can be explained to a sceptical maintenance engineer in one sentence, and does not degrade quietly when conditions change.

Machine learning earns its place where the relationship between signals and condition is genuinely complex, or where normal behaviour shifts with operating context in ways a threshold cannot follow. We discuss that in AI for industrial machinery and, for the engineering of running models on devices, in edge AI on microcontrollers.

Anomaly detection and failure prediction are different problems

These are frequently conflated in vendor material, and the distinction is the most important one in this field.

Anomaly detection answers: is this machine behaving differently from how it normally behaves? It requires only examples of normal, which you can collect starting today. It will flag genuine developing faults, and it will also flag a change of product, a new operator, a replaced component and a sensor that has come loose. It tells you something changed, not what or how serious.

Failure prediction answers: is this specific failure mode developing, and roughly how long do we have? It requires examples of that failure mode developing, recorded against the signals that preceded it. Without them, no technique produces it.

There is a statistical awkwardness worth knowing about. Failures are rare, which means any dataset is heavily imbalanced. A model that simply predicts “no failure” every time will be right the overwhelming majority of the time, and its accuracy figure will look excellent while it is completely useless. Accuracy is the wrong measure for this problem; what matters is how many real developing faults are caught and how many alerts turn out to be nothing.

The data you will actually need

What each capability requires before it is achievable.
Capability Data required Typical time to reach it
Run state, utilisation, downtime duration Current or cycle signal Immediate once installed
Departure from normal Baseline period per machine, covering normal variation Weeks to a few months
Recognising a recurring known fault Several recorded instances with preceding signals Depends entirely on how often it happens
Estimating remaining life Multiple full degradation sequences to failure Often years, and sometimes never for reliable assets
Comparing across an estate Consistent measurement and mounting across machines From the start, if designed in

The bottom-right cell is worth sitting with. For a well-maintained asset that fails once a decade, you may never accumulate enough instances to predict its failure statistically. That is not a defeat. Condition monitoring still catches the degradation when it begins, which is the outcome you actually wanted.

The habit that makes everything else possible

When equipment is investigated, record what was found, when, and against which asset. When something is repaired or replaced, record that too.

This sounds administrative and it is the single highest-value practice in the whole field. Sensor readings without recorded outcomes are a large quantity of numbers about which nothing can be concluded. The same readings with a history of what was found and when become a dataset: this is what the signals looked like in the fortnight before we found a spalled bearing, and here is what they looked like when we investigated and found nothing.

The second of those is as valuable as the first. Knowing what a false alarm looks like is how thresholds improve.

Making it happen is a design problem rather than a discipline problem. If recording a finding takes five minutes and a separate login, it will not happen during a difficult shift. If it takes twenty seconds in the tool the engineer already has open, it will.

Designing alerts people act on

Programmes fail here more often than in the analysis. The technical work can be excellent and the outcome still nil if alerts are ignored.

The trade-off nobody escapes

Every detection system sits somewhere between two failure modes. Set it sensitive and you catch nearly everything, at the price of alerts that turn out to be nothing. Set it conservative and every alert is meaningful, but some real faults slip past.

There is no setting that avoids both, and the right position is a business decision rather than a technical one. It depends on what a missed failure costs against what an unnecessary investigation costs. For a critical asset where failure means a line down for a day, a high rate of false alarms may be entirely acceptable. For a machine with a spare sitting next to it, it is not.

What we would recommend in the first months is deliberately conservative. A system that produces nuisance alarms in its first fortnight is disregarded by week three, and the credibility does not come back easily. Better to miss some early events and retain the belief of the people who have to respond.

What an alert should contain

“Vibration high on Press 2” invites a shrug. The same alert is actionable when it carries the current value, the machine’s own baseline, the trend over recent weeks, when the change began, and what typically causes this pattern. The difference is whether the recipient has to go and investigate before they can even decide whether to investigate.

Where it arrives

Into the system the responding team already uses, which usually means the maintenance management system, so an alert becomes a work order rather than an email in a shared inbox nobody owns. An alert delivered somewhere that is not part of anyone’s routine is not an alert.

Measuring whether the programme is working

Avoided failures cannot be counted directly, which is the central measurement problem: success looks like nothing happening. That makes it important to track the things that can be measured.

Measures that tell you whether a programme is healthy.
Measure What it tells you Watch for
Alerts raised Whether the system is active at all A sudden rise usually means conditions changed, not that machines did
Proportion investigated Whether anyone trusts it A falling rate is the early sign of alert fatigue
Proportion confirmed as real Whether thresholds are set sensibly Very high may mean you are missing early signs; very low means nuisance
Lead time from alert to intervention Whether warning is early enough to be useful Short lead times suggest measuring too infrequently
Unplanned downtime on instrumented assets The outcome that matters Compare against the same assets before, not against other assets
Findings recorded per investigation Whether the dataset is growing This predicts what the programme can do in two years

Which assets justify the cost

Not all of them, and a supplier who suggests otherwise is selling rather than advising.

An asset is a good candidate when failure is expensive, counting lost output, expedited parts, labour and knock-on effects; when the failure mode gives detectable warning; and when there is something you could actually do with advance notice. All three have to hold.

The third is easy to overlook. Warning is worthless if the spare has a twelve-week lead time and none is held, or if the only possible response is the same emergency intervention you would have performed anyway. In that case the useful project might be changing the spares policy rather than instrumenting anything.

Conversely, poor candidates include assets that fail suddenly with no progression, assets that are cheap and quick to swap, and assets that are not a constraint on output. Instrumenting those produces data and no decisions.

Deciding what to monitor, systematically

Choosing assets by instinct works up to a point and then stops scaling. The established approach for doing it methodically comes from reliability engineering, and a simplified version of it is well within reach of any maintenance team.

The idea is to work from failure modes rather than from assets. For each significant piece of equipment, list the ways it can actually fail, not in the abstract but based on what has happened and what the people who maintain it expect. For each of those failure modes, ask four questions.

  • What happens when it fails? Counting lost output, damage to other components, safety implications and knock-on effects downstream.
  • How likely is it? From your own history where you have it, from the maintenance team’s experience where you do not.
  • Does it give warning? And through which signal, over what interval. A failure mode with no detectable progression cannot be monitored, whatever it costs.
  • Could you do anything with the warning? If the spare has a long lead time and none is held, advance notice may change nothing.

Failure modes that score high on consequence, are detectable, and leave a response available are your monitoring candidates. Everything else is better handled by a scheduled task, a design change, a spares policy or an accepted risk.

This exercise has a useful side effect. It frequently reveals that the most valuable intervention is not monitoring at all, but holding a spare, changing a lubricant, improving alignment practice or fixing a root cause that has been quietly generating failures for years.

The response side, which is easy to forget

Detection is half a system. The other half is what happens after the alert, and a programme that improves detection without improving response produces frustration rather than savings.

Three things determine whether warning translates into avoided downtime.

Spares availability. Knowing a bearing will fail in three weeks is only useful if you can obtain one in three weeks. Monitoring frequently exposes that the real constraint is procurement lead time, and the correct response is to change the spares policy for the components monitoring has shown to matter.

Access to the machine. Planned intervention requires a window. If the asset runs continuously and windows are rare, advance warning must be long enough to reach the next one, which affects how early you need to detect and therefore which technique you need.

Skills and capacity. An alert that requires specialist attention is only actionable if that specialist is available. Some organisations discover that better detection mostly reveals how thin their maintenance capacity already was.

None of this argues against monitoring. It argues for considering the whole chain from signal to repair when deciding what to instrument, because the weakest link determines the outcome.

Claims worth questioning

This field attracts confident marketing. A few assertions should prompt questions rather than enthusiasm.

  • “Reduces downtime by a specific percentage.” Measured where, on what equipment, with what maintenance practice beforehand? A figure from someone else’s plant says little about yours.
  • “Predicts failures with high accuracy.” Accuracy is close to meaningless on rare events. A system that always says “healthy” scores extremely well on accuracy. Ask instead how many real developing faults it caught, and how many alerts proved to be nothing.
  • “No baseline period required.” Possible only if it is applying generic thresholds, which is precisely what produces nuisance alarms on some machines and silence on others.
  • “Works out of the box with AI.” A model that has not seen your equipment cannot know what its normal looks like. Something has to learn it, and that takes observation time.
  • “Detects all failure modes.” No technique does. Ask specifically which modes, through which signal, and at what point on the decline.
  • “Our platform is complete.” Ask what happens to your data if you leave, whether it exports, and whether the hardware continues to work without the subscription.

The general principle: a supplier willing to state clearly what their system will not do is usually more reliable about what it will.

How programmes fail

  • Promising prediction in month one. Sets an expectation the data cannot meet and spends credibility that has to be earned back.
  • Nuisance alarms early. The fastest way to make a system invisible.
  • Measuring too infrequently for the P-F interval. The programme detects faults reliably, just not in time to act.
  • Shared thresholds across different machines. Guarantees simultaneous over- and under-sensitivity.
  • Never recording findings. The dataset never grows, so the programme cannot mature past its starting capability.
  • Alerts sent somewhere nobody looks. Common, and entirely avoidable.
  • Instrumenting assets where warning changes nothing. Data with no available response.
  • Using the numbers to judge operators. Data quality collapses the moment measurement feels like surveillance.

A realistic sequence

Establish which assets justify attention, based on what their failures actually cost you. Instrument a small number of representative ones properly, as set out in retrofitting legacy machines. Collect through enough normal variation to know what normal is. Add conservative alerting on departures from each machine’s own baseline. Route alerts into the maintenance system. Record findings every time something is investigated. Review after an agreed period against criteria set in advance. Then extend, on evidence from your own equipment.

Revisit the prediction question when the recorded history can support it, rather than at the start when it cannot. That ordering is the difference between a programme that matures and one that is quietly abandoned after the first year.

Making the case internally

The business case for condition monitoring is usually easier to make than people expect, provided it is built from your own numbers rather than from industry claims.

Start with a single asset and a single event. Take a recent unplanned failure and account for what it genuinely cost: production hours lost, expedited parts and freight, overtime, scrap or rework produced during the disruption, and any effect on delivery commitments. That number is usually larger than the maintenance budget line suggests, because most of it was absorbed elsewhere.

Then ask the maintenance team a simple question: with hindsight, did that failure give warning? For bearings, gearboxes, pumps and belt drives the answer is frequently yes, and often someone will say they had noticed something. That combination, a real cost and a real missed signal, is a more persuasive case than any vendor statistic.

Set the proposal against it: instrument a small number of assets, for a defined period, at a known cost, with agreed criteria for what would justify extending. Committing in advance to what would mean stopping is what makes the request easy to approve, because it bounds the downside.

What to agree before starting

  • Which assets, and why those. Written down, so the reasoning survives a change of personnel.
  • What the pilot must demonstrate. Specific enough to be judged: detections confirmed, lead time achieved, nuisance alert rate below an agreed level.
  • How long the baseline period runs. Long enough to cover normal variation, agreed up front so nobody expects alerts in week two.
  • Who responds, and through which system. Named, not assumed.
  • How findings get recorded. The mechanism, and whose job it is.
  • When the review happens, and who decides. A date in the calendar rather than a vague intention.
  • What happens to the data. Ownership, export, and what remains accessible if the arrangement ends.

A worked example

A plant with a recurring problem: a critical pump set fails roughly twice a year, each time stopping a line for most of a shift while a replacement is fitted. Nobody has measured anything on it beyond a periodic walk-round.

A reasonable programme: instrument the pump and its motor with continuous vibration on the drive-end and non-drive-end bearings, a current clamp at the supply panel, and surface temperature. Add the same instrumentation to two similar pumps that have not been causing trouble, because comparison across similar assets is informative and because two healthy baselines make the unhealthy one easier to interpret.

Collect for a defined baseline period covering the products and duties the line normally runs. Establish per-machine normals. Set conservative alerting on departures, with alerts routed into the maintenance system as work orders carrying the value, the baseline and the trend.

When anything is investigated, record what was found, including the occasions when nothing was wrong. After an agreed period, review: how many alerts, how many investigated, how many real, what lead time, and what happened to unplanned stoppages on those three assets compared with the preceding year.

If the answer is encouraging, extend to the assets identified as next most costly. If it is not, the review should say so plainly, and the money stops. Either outcome is a result; the failure mode to avoid is a programme that continues indefinitely without anyone assessing it.

From pilot to estate

Scaling introduces problems a pilot does not have, and they are worth anticipating.

Consistency. For comparison across machines to mean anything, sensors must be mounted the same way in equivalent positions. A pilot installed carefully by an engineer who understood the intent does not automatically generalise to fifty installations fitted by whoever was available. Document the mounting standard and check it.

Configuration effort. Establishing a baseline per machine is straightforward for three assets and substantial for a hundred. Tooling that automates baseline establishment, rather than requiring manual threshold entry per point, is what makes an estate-scale programme sustainable.

Alert volume. An alert rate that is comfortable across three machines becomes overwhelming across a hundred. Grouping, prioritisation by asset criticality and suppression of related alerts all become necessary rather than optional.

Ownership. Somebody has to own the system: tune thresholds as conditions change, notice when a sensor has failed, and keep the recording habit alive. Programmes decay quietly when the person who cared about them moves on and nobody inherits it.

How we help

  • Scoping. Which assets justify instrumentation, based on failure cost, detectability and whether warning enables a different response.
  • Measurement design. What to sense, where to mount it, and at what rate given the P-F interval you care about. See sensor selection and signal conditioning.
  • Installation and collection. Sensing, gateways, connectivity and storage sized for the data rate. See legacy machine retrofit.
  • Baselining and alerting. Per-machine normals tuned against your own collected data rather than defaults.
  • Integration. Into maintenance and plant systems so alerts become work orders. See SCADA and PLC integration.
  • Analysis as history accumulates. Revisiting prediction when the dataset supports it, honestly.

The service page for this work is predictive maintenance programmes.

Standards worth knowing about

Several published standards give structure to this work. None of them replaces judgement about your own equipment, but they are useful for establishing common vocabulary and for organisations that need to point at something recognised.

  • ISO 20816 series, which supersedes the older ISO 10816, gives guidance on measuring and evaluating mechanical vibration, including indicative severity levels by machine class. Useful for orientation, not a substitute for per-machine baselines.
  • ISO 13374 describes a framework for condition monitoring data processing and presentation, which is helpful when specifying how a system should structure what it produces.
  • ISO 17359 gives general guidelines for condition monitoring and diagnostics of machines, including how to select parameters and set alarm criteria.
  • SAE JA1011 sets out the criteria a process must meet to be called reliability-centred maintenance, which is a useful check on anything described that way.

Whether any of these applies contractually to your work depends on your sector and your customers. Where they do, raise it at scoping, because it shapes what has to be recorded and how. For where maintenance sits among the other things this data supports, see what industrial IoT actually is.

A short glossary

Terms that recur in this field.
Term Meaning
Baseline What a specific machine’s measurements look like when it is healthy, under its own normal operating conditions.
Condition-based maintenance Acting on measured evidence that condition has changed, rather than on a schedule.
Failure mode A specific way a component can fail, as distinct from the asset that contains it.
Functional failure The point at which equipment can no longer do what is required of it, which is not always the point at which it stops entirely.
P-F interval The time between the first detectable sign of a developing fault and functional failure.
Remaining useful life An estimate of how long an asset can continue before intervention is required. Requires substantial history to estimate credibly.
Nuisance alert An alert that is investigated and found to indicate nothing. The main cause of programmes being ignored.
Enveloping A vibration processing technique that isolates the repetitive impacts characteristic of early bearing damage from lower-frequency machine vibration.

Questions we are asked about this

Common questions

What clients ask before starting

Can you predict exactly when a machine will fail?

Not to a date, and anyone offering that should be questioned. What a well-built programme gives you is early warning that a component has begun to degrade, and often a rough sense of urgency. That is usually enough, because the decision it informs is when to schedule work, not what hour to expect a breakdown.

We have no failure history. Can we still start?

Yes, and most operations are in that position. Start with condition monitoring, which learns what normal looks like and flags departures without needing failure examples. Record what is found each time something is investigated, and the history accumulates as a by-product of doing the work.

How long before we see value?

Alerting on departures from a baseline typically becomes useful within months. Prediction of specific failure modes depends on those failures occurring and being recorded, so it is measured in years rather than months for equipment that rarely fails.

What is the P-F interval and why does it matter?

It is the time between the first detectable sign of a developing fault and the point at which the equipment can no longer do its job. It matters because it sets how often you need to measure: checking monthly is useless if the interval is two weeks.

Is machine learning necessary?

Often not. A well-set threshold on a well-chosen measurement catches a great deal, is explainable, and is far easier to trust. Machine learning earns its place where the signal is genuinely complex or where normal behaviour varies with conditions in ways a fixed threshold cannot follow.

How do we stop people ignoring the alerts?

Start conservative, tune against real collected data, give every alert its context and trend rather than just a status, send it into the tool the responding team already uses, and provide a route for them to record what they found. Trust is lost quickly and regained slowly.

What should we measure to judge whether the programme works?

Alerts raised, how many were investigated, how many were confirmed as real, the lead time between alert and intervention, and unplanned downtime on instrumented assets compared with before. Avoided failures cannot be counted directly, which is why the other measures matter.

Which assets should we instrument first?

The ones where failure is expensive and gives warning. An asset that fails catastrophically with no detectable progression is a poor candidate whatever it costs, and one that is cheap and quick to fix rarely justifies instrumentation.

Start a conversation

Which failures hurt most?

Tell us what the equipment is, what unplanned downtime costs you, and whether you currently record anything about machine condition or repair findings.

Prefer email? Write to info@itechgeeks.in