The Fret
I got here by thinking about an oud.
More specifically, I was thinking about why Arabic music can sound as if some of its notes are living in the cracks between ours.
A guitar string is physically capable of producing a continuous range of pitches. Shorten the vibrating length by a tiny amount and the frequency rises by a tiny amount. There is no law of physics insisting that the string has to stop at E, F, F-sharp or anything else we have named.
Then we put frets underneath it.
Suddenly the continuum has preferred landing places.
continuous string position
→ fret
→ permitted pitch
The fret is doing something I had never really thought of as computation.
It is quantising.
My finger can be a little too far forward or a little too far back and, provided I am behind the same fret, the effective string length is fixed at the fret itself. A whole range of imperfect finger positions collapses onto essentially the same result.
That means the fret is doing something else too.
It is correcting error.
Not with software. Not by measuring my finger and moving it. The correction is simply a consequence of the physical structure.
slightly wrong
slightly wrong
almost right
slightly wrong
↓
same fret
↓
same note
An oud doesn't give the player that particular safety rail. There are no frets. The finger can stop the string at positions that a guitar has physically collapsed together.
But that doesn't mean the oud player is wandering around an infinite pitch continuum without structure.
The structure has moved.
Some of it is no longer nailed to the instrument.
It's in the music.
The Next Note
That made me think about what happens when a note is slightly wrong.
Not mechanically wrong. Musically wrong.
If I hear one isolated frequency, there isn't much to go on. It is simply a frequency.
Put it inside a melody and something changes.
The notes before it create expectation.
The final note isn't determined. Music would be fairly pointless if it was.
But the possibilities aren't equally likely either.
Some continuations feel natural. Some feel surprising. Some sound unresolved. Some sound so wrong that even somebody with no musical training knows something has happened.
The surrounding structure is carrying information about the missing note.
Language does this so obviously that we barely notice it.
I put the kettle on to make a cup of t_a.
There is information missing from the word, but there is enough information elsewhere to recover it.
The message contains redundancy.
That word can sound wasteful, but here redundancy is useful. It means the message contains more structure than the bare minimum required to distinguish one symbol from another.
That extra structure makes correction possible.
Music has it too.
A note lives inside a scale, a rhythm, a phrase, a style, a history of previous notes and an expectation of what might happen next.
So the music itself is doing some of what the fret did.
It is constraining possibility.
And once possibility is constrained, prediction becomes possible.
The Residual
If I have heard enough of a piece of music, I can have some expectation about what comes next.
I may not be able to name the note.
I may not even consciously know that I am predicting anything.
But surprise gives the game away.
I can't be surprised by a note unless something in me expected something else.
So underneath the experience there is something like:
history
→ expectation
→ next event
→ comparison
That comparison leaves a remainder.
If the event is close to what was expected, the remainder is small.
If something unexpected happens, the remainder is large.
In engineering language, that remainder is a residual.
residual = observed - predicted
Residuals are everywhere.
We build models of systems, predict what they should do, compare that with what they actually do, and look at the difference.
Usually the ambition is to make that difference disappear.
A good model explains the signal.
What it can't explain gets pushed towards the bucket marked noise.
That was the next little snag.
Why do we assume the bucket is empty of meaning?
The Other Question
Suppose I am measuring the rotational speed of a motor.
The signal I care about is speed.
The sensor also sees tiny variations caused by vibration, electrical interference, temperature, mechanical loading, manufacturing tolerances and whatever else the physical machine is doing.
For my speed measurement, those things are nuisances.
I filter them.
I average them.
I call them noise.
But hand exactly the same data to somebody interested in the bearing and the vibration may be the interesting part.
Give it to somebody interested in electrical loading and another part of the discarded signal may become useful.
Give it to somebody trying to distinguish one nominally identical motor from another and tiny manufacturing differences may become a fingerprint.
Nothing in the data changed.
The question changed.
That makes noise a slightly dangerous word.
There is genuine randomness. Sensors have limits. Thermal noise is real. Information can genuinely be absent or destroyed.
But in ordinary engineering use, noise can also mean something much less fundamental:
things this measurement was not built to care about
That isn't the same as:
things containing no information
Once those two categories came apart, I started seeing sensors differently.
The Common Cause
A physical event rarely announces itself through only one mechanism.
Start an electric motor.
Current changes. The magnetic field changes. The rotor accelerates. The housing vibrates. Sound appears. Bearings move under load. Temperature begins to change.
These aren't separate events.
They're different consequences of the same event.
There is a lovely formulation of this in a 2017 paper by Relja Arandjelović and Andrew Zisserman called Look, Listen and Learn.[1]
They were interested in something apparently quite different: whether useful visual and audio representations could be learned from large numbers of unlabelled videos.
The useful information was already sitting inside the videos.
Fingers move and an instrument makes a sound.
Lips move and speech appears.
Cars move and engines make noise.
The authors make the important point that the visual and auditory events occur together because they have a common cause.
That phrase unlocked something for me.
common physical cause
↓
┌────┼────┐
↓ ↓ ↓
sound motion vibration
If two observations keep changing together because something underneath them is causing both, the correspondence itself contains information.
And the researchers didn't need a human to sit beside millions of video frames writing:
this is a piano
these are fingers
that sound came from this object
They trained networks to decide whether visual information and an audio snippet corresponded.
The relationship between the modalities supplied the supervision.
The world had accidentally labelled itself.
Free Supervision
That idea turns out to have several different forms.
Andrew Owens and colleagues took ambient sound and used it as supervision for visual learning.[2]
The roar of a car tells you something about what might be in the scene.
The sound of water tells you something.
The buzz of a refrigerator tells you something.
They trained a visual network to predict statistical properties of the sound associated with a video frame.
They weren't giving the network explicit object labels.
Yet object-selective units emerged inside the network, and the learned representation proved useful for recognising objects and scenes.
That is a peculiar inversion.
The sound was not originally created to label the image.
It was simply another consequence of whatever was happening in the world.
But because the two observations were related, one became a source of supervision for the other.
Then Pulkit Agrawal, João Carreira and Jitendra Malik showed another version of the same move in Learning to See by Moving.[3]
The extra signal wasn't sound this time.
It was movement.
A mobile agent already has information about its own motion. A robot can get it from motor commands, gyroscopes or accelerometers. An animal has its own machinery for sensing movement and orientation.
The researchers used that egomotion as supervision for learning visual features.
Again, the interesting bit wasn't created as a label.
movement
was not a label
until somebody asked
what it could teach vision
That's very close to the thought I had been circling.
A signal can be irrelevant to the job for which it was originally collected and still contain structure that teaches us something else.
The Sound of Pixels
Then the relationship becomes stranger.
In The Sound of Pixels, Hang Zhao and colleagues used the natural synchronisation between video and audio to learn to locate sound-producing regions and separate mixed sounds without manual labels telling the system which instruments were present, where they were, or what each one sounded like.[4]
The visual signal helped untangle the audio signal.
That matters.
It means the second sensor isn't merely adding another measurement to a pile.
It can constrain the interpretation of the first.
If I hear several instruments mixed together, the audio alone has an ambiguity problem.
If I can also see which objects are producing sound, part of that ambiguity disappears.
messy audio + vision
↓
fewer plausible explanations
That is beginning to look a lot like error correction again.
Not because the modalities contain copies of one another.
They don't.
Because they contain overlapping evidence about the same underlying world.
The redundancy is imperfect.
That's what makes it useful.
The Shared Space
By 2021, this had moved beyond specially designed audio-visual systems.
VATT, the Video-Audio-Text Transformer, was trained from raw video, audio and text using multimodal contrastive learning.[5]
The important word for me is representation.
The model isn't simply being trained to produce one answer.
It learns an internal representation useful enough to transfer to downstream jobs such as video action recognition, audio-event classification, image classification and text-to-video retrieval.
Then ImageBind pushed the idea somewhere much closer to the thing I had in my head.[6]
Six modalities:
images
text
audio
depth
thermal
IMU
One shared embedding space.
And there is a particularly interesting wrinkle.
ImageBind doesn't require a giant dataset in which every one of those sensors observed every event simultaneously.
Images act as the bridge.
Image-text pairs can come from one source.
Video-audio pairs can come from another.
Images and depth from another.
Images and thermal data from another.
Video and inertial measurements from another.
Align each of those to the image representation and relationships emerge between modalities that were never directly paired during training.
The authors call some of the resulting behaviour emergent alignment, something I began studying 20 yrs ago and wrote about here 9 yrs ago Could ants power Web3.0 to new heights? OSPF v’s ANTS
audio ─────┐
depth ─────┤
thermal ───┼→ shared space
IMU ───────┤
text ──────┘
This is not evidence that the model has learned physics in some deep ontological sense.
But it is evidence that very different observations can carry enough shared structure to become aligned in a common learned representation.
And suddenly the idea of several cheap physical sensors watching one machine doesn't look quite so eccentric.
The Body
Speech gives an even more physical example.
I speak and a microphone hears pressure waves travelling through the air.
But the airborne sound isn't the act of speaking.
It's one consequence of it.
The same act also sends mechanical vibration through the body.
The Vibravox dataset makes that wonderfully concrete.[7]
Its researchers recorded 188 people using an ordinary airborne reference microphone alongside five body-conduction sensors: two in-ear microphones, two bone-conduction vibration pickups and a laryngophone.
Each sensor gets a different version of the same speaker.
The body-conducted channels have disadvantages. Tissue filters the signal and restricts bandwidth.
But they have a rather useful advantage too.
They pick up the speaker through the body and are much less exposed to ambient acoustic noise.
So:
air pressure ────────┐
bone vibration ──────┤
ear-canal signal ────┼→ same act of speaking
throat vibration ────┘
Again, none is the truth.
They're different physical projections of the same event.
That immediately suggests applications in robust speech and speaker verification, which Vibravox explicitly investigates.
But it also suggests something more general.
If two sensors observe the same event through different physics, the disagreement between them may be just as informative as the agreement.
The Five Problems
By this point I had independently wandered into a field that already had names for several pieces of the problem.
Baltrušaitis, Ahuja and Morency's survey of multimodal machine learning divides the field into five broad challenges: representation, translation, alignment, fusion and co-learning.[8]
That taxonomy is useful because it stops all of this collapsing into the vague phrase sensor fusion.
These aren't the same problem.
representation
→ how do several modalities describe something?
translation
→ can one modality be mapped into another?
alignment
→ which parts correspond?
fusion
→ what can they tell us together?
co-learning
→ what can one modality teach another?
The survey also uses two words that matter enormously here.
Complementarity and redundancy.
Two sensors can tell us different things about an event.
They can also tell us some of the same things in different ways.
Those sound like opposites.
They're actually the reason the combination can be more useful than either sensor alone.
Redundancy lets one observation constrain another.
Complementarity supplies information the other observation never had.
Now put both properties into a physical machine that repeats its behaviour thousands of times.
That's where my boiler comes back in.
The Boiler
A domestic boiler is a wonderful machine for this thought because it is boring.
It isn't a laboratory.
It isn't a jet engine.
It isn't carrying a rack of instrumentation costing more than the machine itself.
It sits on a wall and repeatedly performs variations of the same physical sequence.
There is demand for heat.
The fan starts.
Ignition occurs.
Combustion establishes.
The pump moves water.
Pipes warm.
Metal expands.
Temperatures rise.
The machine settles into operation.
Eventually it shuts down.
A good heating engineer can sometimes hear that something is wrong before any diagnostic code appears.
That's interesting in itself.
The engineer isn't directly measuring a worn bearing with their ear.
They're recognising that the physical behaviour of the machine has moved away from something learned as normal.
So imagine sticking a cheap box beside the boiler.
microphone
accelerometer
current sensing
temperature
perhaps magnetic field
perhaps pressure
The individual sensors don't need to be particularly impressive.
The microphone doesn't have to diagnose the fan.
The accelerometer doesn't have to measure bearing wear.
The temperature probe doesn't have to know why the temperature changed.
They just have to watch together.
One heating cycle becomes one multimodal physical event.
Then another.
Then another.
Cold days. Warm days. Hot-water demand. Heating demand. Short cycles. Long cycles. A new machine. An ageing machine.
Now there are two different things we could ask the model to learn.
The first is the obvious one.
A fault.
The second is much more interesting.
The machine.
Normal
There is already a branch of machine learning built around not having examples of everything that can go wrong.
Anomaly detection.
Deep SVDD is one example.[9]
The starting assumption is often that most of the training data represents normal behaviour.
Instead of learning a catalogue containing every possible anomaly, the system learns a compact description of normality.
In Deep SVDD, a neural network learns to map normal examples into a compact region around the centre of a hypersphere.
New observations further from that learned normal region receive higher anomaly scores.
learn normal
↓
new observation
↓
how far from normal?
The paper explicitly notes machine monitoring as a setting where training data may be collected during normal operating state.
That fits the boiler rather nicely.
I don't need to own examples of every future boiler fault before I start collecting useful data.
I can learn something about what this machine normally does.
But anomaly detection still leaves an interesting question unanswered.
It can tell me that something has moved away from normal.
It doesn't necessarily tell me what hidden physical quantity moved.
For that, industrial sensing already has another idea.
The Soft Sensor
Industry has been estimating things it can't conveniently measure for a long time.
The name is soft sensor.
Sometimes the quantity we really want is difficult, expensive or slow to measure directly.
But it may be related to things that are cheap and easy to measure.
easy measurement 1 ─┐
easy measurement 2 ─┼→ model → difficult measurement
easy measurement 3 ─┘
A recent survey of machine learning for industrial sensing and control makes the practical problem very clear.[10]
Temperature, pressure, flow and level may be available continuously.
A quality variable may be measured once every eight hours or once every twenty-four.
So the cheap data is plentiful and the expensive answer is scarce.
The survey identifies lack of labelled data as a central problem in soft-sensor development and discusses unsupervised methods including PCA and autoencoders for extracting features from unlabelled process data before relating those features to the desired output.
It also supplies a useful slap back towards reality.
Industrial systems change.
Operating conditions move.
A soft sensor trained under one condition may stop performing well under another.
Models degrade and need maintenance, updating or retraining.
So this isn't:
collect data
→ sprinkle AI
→ know machine forever
The physical world moves underneath the model.
That matters enormously.
But something else in the conventional soft-sensor formulation bothered me.
We still begin by knowing the answer we want.
X₁, X₂, X₃ ... → Y
We know Y.
We know that measuring Y directly is awkward.
So we find cheaper signals that can estimate it.
Perfectly sensible.
But everything I had just read about correspondence, self-supervision and learned representations suggested another possibility.
What if we don't know Y yet?
The Unknown Y
This is where the literature stops giving me the answer.
That's important.
None of these papers demonstrates the strong version of what I'm proposing.
ImageBind doesn't show that a cheap box of sensors can be attached to a boiler and spontaneously discover useful new physical variables.
VATT doesn't.
Vibravox doesn't.
Deep SVDD doesn't.
The industrial soft-sensor literature doesn't.
What they give me are pieces.
correspondence
self-supervision
multimodal representation
redundancy
complementarity
normality
soft sensing
transfer
The step I want to take is this.
Don't begin with:
X₁, X₂, X₃ ... → Y
Begin with:
X₁, X₂, X₃ ... → Z
Let Z be a learned representation of the physical system.
Use synchronisation.
Use recurrence.
Use the fact that one event leaks into several sensors.
Use one modality to constrain another.
Use repeated operating cycles to learn what tends to remain stable and what tends to move.
Don't initially insist that a particular dimension means fan speed, pump wear, combustion quality or heat-exchanger fouling.
Learn the machine first.
Then ask what the representation knows.
The Test
There is a fairly brutal way to test whether this is anything more than an attractive idea.
Build the cheap multimodal sensor.
Collect the data.
Don't give the model the eventual target labels.
sound ─────────┐
vibration ─────┤
current ───────┼→ self-supervised learning → Z
temperature ───┤
other signals ─┘
Then freeze the representation.
Only afterwards introduce measurements that weren't used to train it.
Suppose I temporarily fit a much more expensive instrument that measures some physical quantity accurately.
Can that quantity be recovered from Z with only a small amount of labelled data?
Suppose I introduce a controlled degradation the model has never been told about.
Did Z already move before I named the fault?
Suppose I take the box to another boiler of the same model.
Which parts of the representation transfer?
Then start removing sensors.
all sensors
→ remove microphone
→ remove vibration
→ remove current
→ remove temperature
If the useful information genuinely lives partly in correspondence, the damage should be measurable.
Perhaps the microphone alone knows very little.
Perhaps the accelerometer alone knows very little.
Perhaps together they make some hidden state much easier to infer.
And the hardest version of the experiment is the one that matters most.
Choose the question after the representation has been learned.
If it can only answer questions that shaped its training, we've built another task-specific sensor.
If useful physical quantities can be extracted afterwards, we've built something different.
The Instrument
This changes what an instrument can be.
The traditional instrument has a name on the front.
thermometer
tachometer
pressure gauge
voltmeter
The name tells you what question it was built to answer.
That's normally a virtue.
But perhaps there is another class of instrument.
Not a universal sensor. Physics doesn't permit that.
A microphone cannot recover a physical variable that leaves no information whatsoever in pressure waves.
No amount of machine learning can infer information that never reached the sensors.
The more interesting possibility is a device whose useful measurement isn't fixed when the hardware is installed.
physical system
→ several cheap observations
→ learned representation
→ later questions
That is almost the reverse of normal instrumentation.
Instead of deciding what matters and then collecting the minimum data necessary to measure it, we collect several cheap views of the physical event and preserve enough structure to ask other questions later.
The sensor becomes less like a gauge.
More like a witness.
What Does the Noise Already Know?
And that brings me back to the bit we normally throw away.
The residual.
The vibration that spoiled the clean measurement.
The current transient that disappeared in the average.
The background sound.
The thermal wobble.
The little timing discrepancy between two events that were supposed to happen together.
None of these is automatically useful.
Some really will be noise.
Some will be artefact.
Some will be confounding variables.
Some correlations will disappear the moment the operating environment changes.
Some models will confidently learn the wrong thing.
The industrial literature is quite clear that even deliberately constructed soft sensors can degrade when operating conditions move.
So the proposition isn't:
all noise contains hidden meaning
It's narrower.
unwanted for this question
≠
informationless
That distinction now matters because the economics have changed.
Microphones are cheap.
MEMS accelerometers are cheap.
Temperature sensors are cheap.
Current sensing is cheap.
Storage is cheap.
And we now have learning methods specifically designed to extract structure from unlabelled data, align different modalities, exploit naturally paired observations, learn representations that transfer to later tasks, and model normal behaviour without first enumerating every possible failure.
For decades the sensible engineering move was to know what you wanted to measure, buy the right instrument, remove everything that interfered with the measurement, and throw the rest away.
That still makes sense when the question is known.
But perhaps it is no longer the only sensible move.
old:
ask question
→ choose measurement
→ discard irrelevant signal
possible:
observe
→ preserve correspondence
→ learn structure
→ ask questions later
The interesting thing isn't that AI can find patterns in sensor data.
We already know that.
It's that the physical world is producing overlapping, redundant, complementary traces of itself all the time.
Sound knows something about vision.
Motion knows something about vision.
Vision knows something about sound.
Body vibration knows something about speech.
Temperature, current, vibration and acoustics may each know something about the machine that produced all four.
And perhaps the useful measurement we haven't thought to make yet is already distributed across them.
Not labelled.
Not conveniently isolated.
Not necessarily visible in any one channel.
Just sitting there in the correspondence.
We may simply be throwing it away because, for the question we happened to ask today, we called it noise.
Endnotes
- Relja Arandjelović and Andrew Zisserman, Look, Listen and Learn , Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017, pp. 609–617. DOI: 10.1109/ICCV.2017.73. ↩
- Andrew Owens, Jiajun Wu, Josh H. McDermott, William T. Freeman and Antonio Torralba, Ambient Sound Provides Supervision for Visual Learning , European Conference on Computer Vision (ECCV), 2016, pp. 801–816. DOI: 10.1007/978-3-319-46448-0_48. ↩
- Pulkit Agrawal, João Carreira and Jitendra Malik, Learning to See by Moving , Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2015, pp. 37–45. DOI: 10.1109/ICCV.2015.13. ↩
- Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott and Antonio Torralba, The Sound of Pixels , European Conference on Computer Vision (ECCV), 2018, pp. 570–586. DOI: 10.1007/978-3-030-01246-5_35. ↩
- Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui and Boqing Gong, VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text , Advances in Neural Information Processing Systems 34 (NeurIPS), 2021. ↩
- Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin and Ishan Misra, ImageBind: One Embedding Space To Bind Them All , Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 15180–15190. DOI: 10.1109/CVPR52729.2023.01457. ↩
- Julien Hauret et al., Vibravox: A dataset of French speech captured with body-conduction audio sensors , Speech Communication, vol. 172, 2025, 103238. DOI: 10.1016/j.specom.2025.103238. ↩
- Tadas Baltrušaitis, Chaitanya Ahuja and Louis-Philippe Morency, Multimodal Machine Learning: A Survey and Taxonomy , IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 41, no. 2, 2019, pp. 423–443. DOI: 10.1109/TPAMI.2018.2798607. ↩
- Lukas Ruff, Robert Vandermeulen, Nico Goernitz, Lucas Deecke, Shoaib Ahmed Siddiqui, Alexander Binder, Emmanuel Müller and Marius Kloft, Deep One-Class Classification , Proceedings of the 35th International Conference on Machine Learning (ICML), PMLR 80, 2018, pp. 4393–4402. ↩
- Nathan P. Lawrence, Seshu Kumar Damarla, Jong Woo Kim, Aditya Tulsyan, Faraz Amjad, Kai Wang, Benoit Chachuat, Jong Min Lee, Biao Huang and R. Bhushan Gopaluni, Machine learning for industrial sensing and control: A survey and practical perspective , Control Engineering Practice, vol. 145, 2024, 105841. DOI: 10.1016/j.conengprac.2024.105841. ↩