Research4 min read
Can AI recognize engine problems through sound?
Exploring how machine learning can identify mechanical issues from engine audio.
By SpotaLabs Team

Good mechanics diagnose by ear all the time. A lifter tick, a failing water pump bearing, an exhaust leak at the manifold: each has a sound that someone with enough experience can pick out in seconds. That raises an obvious question for anyone working with machine learning. If a person can learn to hear it, can a model?
This is a write-up of how we think about that question while building EngineEar: what makes it tractable, what makes it hard, and what we believe a phone-based system can and can't do.
Why sound carries diagnostic information
An internal combustion engine is a machine full of periodic events. Pistons move up and down, valves open and close, the crank and cam rotate at fixed ratios, and accessories like the alternator and water pump spin at speeds tied to the crank through belts.
Because these events are periodic, their sounds are too. A problem with a component usually changes the sound in a way that's linked to that component's timing:
- Once per combustion event: misfires, injector problems and some valve issues.
- Once per crank or cam revolution: rotating-assembly and valvetrain noises.
- Tied to accessory speed: bearings in belt-driven components, often at a pitch that doesn't match the engine's firing frequency.
- Broadband and irregular: leaks, loose heat shields, and rattles that depend on vibration rather than timing.
That structure is what makes the problem learnable. The model isn't looking for "a bad noise". It's looking for patterns in frequency and timing that correlate with specific mechanical behavior.
Turning audio into something a model can read
Most audio models don't work on the raw waveform. A common approach is to convert audio into a spectrogram, an image of how energy is distributed across frequencies over time, often on a perceptual (mel) frequency scale. Periodic engine events show up as regular structures; anomalies show up as disruptions to that structure.
From there, the problem looks a lot like image and sequence classification, and borrows from the same toolbox: convolutional and transformer-based models, trained on labeled examples and evaluated on recordings they've never seen.
The hard parts
If it were that simple, this would be a solved problem. Several things make it genuinely difficult.
Every engine is different
Engine layout, cylinder count, firing order, displacement, aspiration and exhaust all change the baseline sound dramatically. A healthy boxer four sounds "lumpy" in a way that would be alarming from an inline four. A model that learns "irregular = bad" from one engine type will be confidently wrong on another.
Labels are expensive
To learn "this sound means a worn lifter", you need recordings where someone has actually confirmed the diagnosis. Those are slow to collect, and real-world labels are often uncertain: a car goes to a shop, several things get fixed, and nobody knows which one made the noise.
The recording environment
Phones aren't lab microphones. Wind, road noise, other cars, echoes off walls, the phone's own processing and the distance from the engine all change the signal. A model can easily learn to recognize parking garages instead of engine faults if the training data isn't varied enough.
Faults overlap
Real cars rarely have exactly one problem. Overlapping noises and partial failures make a single "the answer is X" classification unrealistic.
What we think works
These constraints shaped a few design principles that we're exploring in EngineEar.
Control the recording, not just the model. Guided scans (idle, cold start, controlled revs) mean recordings are more comparable. A lot of the "intelligence" in a system like this sits in the data collection, not the network.
Compare against the car's own baseline. Instead of relying only on what a healthy engine "should" sound like, compare new recordings against earlier recordings of the same car. Anomaly detection relative to a personal baseline sidesteps some of the engine-to-engine variation problem.
Say when you don't know. A model should be able to reject a recording as unusable, and to report uncertainty rather than forcing a diagnosis. An honest "we couldn't get a clean read, try again somewhere quieter" is more valuable than a confident wrong answer.
Describe, don't diagnose. The output is framed as "this pattern is associated with these issues, and here's what to check", which fits what audio evidence can actually support.
So, can it?
Our current answer: yes, for a meaningful set of problems, within limits. Audio can flag that something has changed and point toward the likely area, especially for faults with a clear timing signature. It can't see inside the engine, and it shouldn't pretend to.
We think the real value isn't replacing the mechanic's ear. It's giving drivers a way to notice change early, to record it, and to walk into a shop with something better than "it makes a sort of ticking noise sometimes".
If you want to see how this research turns into product, read Teaching AI to listen to engines on the blog.
- engineear
- machine learning
- audio
- research