There is a number being put in front of you at every trade show this year. Millions of components analyzed. Thousands of plants. Decades of fault history. The implication is always the same: our artificial intelligence knows more than the others, because it was fed more.

It is an impressive number. It is also the wrong one to compare — and not for a marketing reason. It is how these systems actually work.

Short answer. The general knowledge of vibration analysis — fault frequencies, standards, failure patterns, the four stages of a bearing defect — is published, and every capable language model already carries it. No vendor owns that knowledge and none can sell it to you. What decides whether a diagnosis is right is how much the system knows about your machine — its datasheet, its bearings, its real running speed, its operating modes, its trend, its spectrum, its waveform, the analyst’s notes — and that it has all of it in front of it at once, which is the one thing no person can do: we read one screen at a time. And then whether it is allowed to go and check the things it is unsure about. Ask any vendor to show you both. The size of a training corpus is the easiest claim for them to make and the hardest for you to verify.

Why a bigger training set does not make a better diagnosis

Think about what that number is really selling you: general vibration knowledge. What a bearing defect looks like in an envelope spectrum. Why misalignment shows up axially at 2X. How sidebands around gear mesh reveal an eccentric gear. What ISO 20816 says about a 200 kW machine on a rigid foundation.

None of that is scarce any more. It lives in the textbooks, the standards, the conference papers and forty years of published case studies, and every serious language model has already read all of it. You do not have to buy it from anyone.

What is scarce is something no database of other people’s failures can contain: your machine.

Diagram contrasting the general question, which published knowledge already answers, with the particular question about one specific pump, which only the plant’s own database can answer; below, the tools the assistant can use to go and check
The question on the left is answered by any capable model. The one on the right is answered by your data, and by what the assistant is allowed to do with it.

What every AI already knows, and what it cannot possibly know

Ask any capable model today to explain how to tell static imbalance from parallel misalignment. You will get a correct, well-organized answer.

Now ask it which of the two your fan has. It has nothing to work with — unless you give it something.

That is the whole game. Two very different questions hide inside what looks like one:

The general question

What does an outer race defect look like?

The physics, the standards, the fault frequencies, the typical progression through the four stages, the reason enveloping catches it before velocity does.

This knowledge is published, stable and universal. It is the same in Monterrey and in Hamburg, on a 1998 pump and on a 2026 one. Nobody has a proprietary version of it, and a model that has read the literature carries it already.

The particular question

Does this pump have one, today?

Its bearing part numbers, its real running speed, the three operating modes it actually has, what its trend did over the last eleven months, how axis H compares with axis A, what the analyst wrote in March.

This knowledge exists in exactly one place: your plant. It is not in any vendor’s training corpus, it cannot be, and the size of that corpus does not change it.

A system that answers the general question beautifully and cannot reach the particular one produces something very recognizable: an answer that could have been written about any machine of that type. It is the textbook talking. Experienced analysts spot it in one paragraph, and they stop trusting the tool shortly after.

Context is ninety percent of the answer

Which is why the tools matter as much as they do. They are not a second ingredient standing next to context — they are how the system goes and gets more of it when what it was handed is not enough.

A human analyst does not diagnose by staring at one screen. They zoom into the region around 1X to see whether there really are sidebands there. They switch the same measurement to envelope. They pull up the waveform to check whether the impacts are periodic. They compare the horizontal axis against the vertical one. They open the capture from three months ago and put it next to today’s. They check whether the sensor itself is healthy before believing anything it says.

None of that is knowledge. It is procedure — and an assistant that cannot do it is reduced to commenting on whatever screenshot it was handed.

In EI-Analytic™ the assistant is called Erby, and it has those same moves available as tools. It can:

  • request the trend for any period
  • read the octave bands
  • open a specific capture
  • zoom into a frequency range
  • compare axes, and compare moments in time
  • pull the asset datasheet
  • check sensor health
  • read the plant overview

When it is not sure, it goes and looks, exactly as a person would — and every step it takes is visible to you inside the answer.

The screens below are one real turn on one real pump. Nothing is staged: Erby was asked for a diagnosis, and this is what it did before giving one.

A turn in progress: the assistant has called Compare moments and Phase between axes, states that the energy jumped at 630 to 700 Hz and 1.1 to 1.3 kHz rather than at 1X or 2X, and draws the trend of the motor inboard point with each measurement colored by its operating mode
Each pill is a tool the assistant chose to call, with what it cost to read. The conclusion names frequencies, not adjectives — and the trend below it is colored by operating mode, with the learned limits drawn in.

Notice what the conclusion is made of. Not high vibration on the motor, but this: the energy rose at 630 to 700 Hz and at 1.1 to 1.3 kHz, and not at 1X or 2X. Which means a bearing defect exciting a resonance — not unbalance, and not misalignment.

You can prove that sentence wrong. Go to the spectrum and check the numbers. Which is exactly what it does next, without being asked.

The before and after spectra of the same point overlaid, August in blue and September in green, with the two peaks that appeared marked at 670 Hz and 1283 Hz
The proof it draws for itself: the same measurement point six weeks apart, with the two new peaks marked. The chart title came out in Spanish in this session.

So the comparison that matters is not whose training set is bigger. It is how much of the real situation the system can reach, and how many of an analyst’s moves it can actually make.

The short version

  • The general physics of vibration analysis is published knowledge; a modern model already has it, and no vendor sells exclusive access to it.
  • What decides whether an answer is right is the specific context of the machine — datasheet, bearings, operating modes, trend, spectrum, waveform, notes, open cases.
  • Almost as decisive: whether the assistant can use the analyst’s own tools to go and check, instead of commenting on a single screenshot.
  • A history of other people’s failures is genuinely useful for starting points, and genuinely unable to tell you what your machine is doing today.
  • The diagnosis still has to be auditable and confirmable by a person. Context makes an answer better; it does not make it authoritative.
  • The memory worth having is not the one the model shipped with. It is the one it builds from the analyst — their criteria, their exceptions, their way of working.

Where a failure database really does help

It would be dishonest to leave this out, so let us be precise about it.

A large body of historical failures across many plants is valuable, for a specific and limited set of jobs:

  • sensible starting alarm limits by component type
  • criticality frameworks
  • typical failure modes to consider for a class of asset
  • how long a given defect usually takes to progress

That is real engineering value, and it is the honest core of what those million-component numbers are selling. When you commission a brand-new machine with no history of its own, statistics from similar machines are the best first guess available — and anyone who tells you otherwise is selling something too.

What such a database cannot do is tell you about your machine rather than about machines like it. Two identical pumps — same model, same year, same duty — mounted forty meters apart will have different baselines: different foundation stiffness, different piping strain, different resonances. The population tells you what to expect. It never tells you what is happening.

And there is a quieter issue. When a fault history is proprietary, the reasoning that comes out of it is proprietary too. You are told that similar components failed this way, and you cannot inspect the claim. That is the black box argument arriving through a new door — and we have written before about why we do not accept it.

How EI-Analytic™ builds the answer instead: three layers

It is worth being explicit about which layer does what — because in our software the artificial intelligence is not the one issuing the diagnosis.

The rule engine names the fault. Thirteen fault types, each one a set of conditions with its arithmetic visible: which amplitude, compared against which, with which factor, met or not met. It works from the very first measurement, with no history and no training. And you can edit every rule when it gets your twenty-year-old fan wrong. This is what produces the diagnosis you see on the dashboard.

The auto-diagnosis panel: candidate faults ranked with their rule counts, the arithmetic of the selected rule, and the spectrum below with the peaks that rule measured marked
The diagnosis is a ranked set of hypotheses, each with the arithmetic that supports it.

Machine learning learns your machine — not somebody else’s. Clustering groups each point’s measurements into the operating modes that machine actually has, with nobody labeling anything, and reports when a mode appears that has never been seen before. Alarm limits are learned from a healthy period you select, on that specific point. Both are trained on your data, about your equipment — the only training that transfers to your decisions.

Erby reads the whole picture and explains it. Before it answers, the software assembles a context package and shows it to you first: which layers go in, how large each one is, and an option to anonymize company and area names before anything is sent. Then it works with the tools above and returns a report with its evidence attached — amplitudes per axis, frequencies, ratios against 1X, bearing fault frequencies from the envelope — so every statement traces back to a measured number.

And the last rule is the one we would not trade for any amount of accuracy: the assistant does not act on its own. It can prepare a case, a note, a route or a report, but the window opens filled in and a person presses save. Your data stays in your platform. Every answer shows what it cost.

The memory that matters is not the model’s. It is yours.

A language model arrives with a generic memory: everyone and everything, averaged. That is a fine place to start and a poor place to finish. So we hand a good part of that memory back, and replace it with something worth far more — the memory the system builds from the analyst who uses it.

How you like your alarms set. Which machines you already decided to ignore, and the reason you gave. That the compressor in area 4 always looks alarming at startup and never is. The way you word a case before you hand it to production. The point at which you stop watching and pick up the phone.

Erby keeps all of that, visibly and undoably, and brings it to the next question. It is not learning vibration from you. It already knows vibration. It is learning you.

What to ask any vendor about their AI

If you are evaluating platforms this year, the size of a training corpus is the easiest thing for a vendor to say and the hardest for you to verify. These six questions are harder to dodge:

Ask this What a good answer looks like
What exactly does the AI see about my machine before it answers? A list you can inspect, and ideally a screen that shows it to you before sending
Can it go and check something it is unsure about? Named tools it can use on your data — trends, captures, zoom, axes — not a single fixed snapshot
Where does the diagnosis itself come from? Rules or models you can open, read and modify
What happens to my data? Stays in your database; anonymization available; nothing sent without you seeing it
Can it act without me? It should not. Prepared for you, confirmed by you
If the answer is wrong, what do I have to argue with? The numbers it used, per axis and per frequency, not a confidence percentage alone

The question that actually separates one vibration AI from another

The industry is about to spend a year arguing about whose corpus is larger. It is a comfortable argument for vendors, because it can be asserted and not checked.

The useful question is simpler, and much harder to fake. A vibration AI is only as good as what it knows about the machine in front of it, and what it is allowed to do to find out more. General expertise now comes free with the model. Your machine’s history, its modes, its baseline and its open cases do not — and they are the entire difference between an answer about pumps and an answer about this pump.

We built ours on that assumption. You should make every vendor, including us, show you exactly what their AI is looking at.

See what Erby is looking at

FAQs about AI in vibration analysis

Does a bigger AI training dataset make a vibration diagnosis more accurate?

Not beyond a point. The general knowledge a large corpus is meant to buy — fault frequencies, standards, failure patterns, ISO limits — is already published, and every capable language model has read it. Extra examples of other plants’ failures do not tell the system anything about the speed, mounting, operating modes or history of the machine you are actually looking at. That information exists only in your own database.

What does an AI need to know to diagnose a specific machine?

Its datasheet and bearing part numbers, its real running speed, the operating modes it actually has, its trend over time, the spectrum and waveform of the measurement in question, how one axis compares with another, and what the analysts have written about it before. In EI-Analytic™ that context package is assembled from your own database and shown to you before anything is sent, with the option to anonymize company and area names first.

Can an AI assistant check something it is unsure about?

It can, if it has been given tools to do it with. An assistant that only sees a single screenshot can do nothing but comment on that screenshot. Erby can request a trend for any period, open a specific capture, zoom into a frequency range, switch to envelope, compare axes and compare two moments in time — the same moves a human analyst makes, with each step visible in the answer.

Is a failure database useless for condition monitoring, then?

No. It is genuinely useful for a specific set of jobs: sensible starting alarm limits by component type, criticality frameworks, the failure modes worth considering for a class of asset, and how long a given defect usually takes to progress. On a brand-new machine with no history of its own, statistics from similar machines are the best first guess available. What a population cannot do is tell you what one individual machine is doing today.

Can the history of one machine predict what an identical machine will do?

Not reliably. Two identical pumps — same model, same year, same duty — mounted forty meters apart will show different baselines, because foundation stiffness, piping strain and local resonances are not the same. Population data tells you what to expect from that class of machine; only that machine’s own baseline tells you whether something has changed.


Further reading