Skip to content
NatureHQ

Senses and abilitiesmechanism

How scientists listen to animals

A model can tell you who made a sound, and when, and what was happening. There is no training data anywhere for what it meant.

Microphones in the woods and hydrophones in the ocean now record continuously for months, and machine learning finds the calls inside recordings nobody could ever listen to. The models are genuinely good at telling you what made a sound, who made it, and roughly what was happening. There is no route from any of that to what the sound means.

Almost everything on this site about animal communication exists because somebody left a recorder running. A hydrophone moored for six months produces more sound than a person could listen to in a working lifetime, and until recently that was the bottleneck: recording was cheap and listening was not. Machine learning removed the bottleneck. Models now detect calls in continuous audio, sort them by species, identify individuals, and predict the behavioural context — all at rates well above chance and fast enough to process an ocean-basin dataset. That is a real transformation and most of what is known about whale distribution, bat activity and dawn choruses at continental scale depends on it. It is also, structurally, a set of classification problems, and it is worth being exact about why that matters. A supervised model learns a mapping from sound to labels a person supplied: this is a humpback, this is individual 47, this call happened during feeding. Every one of those labels is something a human could in principle write down. Meaning is not, because nobody independently knows what any call means — that is the thing being asked. Machine translation between human languages works because millions of sentence pairs already mean the same thing in both; nothing like that exists for any animal, and it cannot be manufactured, because manufacturing it would require already knowing the answer. What the models can genuinely do is generate hypotheses. The elephant work is the model of how that should go: a classifier found structure suggesting calls were addressed to particular individuals, and then somebody played the calls back to the elephants and watched them respond. The machine found the candidate; the animals confirmed it.

Developed coverage · 75% complete · reviewed 2026-08-31

What this page covers

The methods apply to anything that makes a sound: whales, bats, birds, frogs, insects and fish. What differs between them is the frequency range and how far the sound travels.

Often confused with: Translation

Quick facts

What the machines do well
Detect, classify, identify individuals, predict context
What they cannot do
Recover meaning — there is no training signal for it
The core instrument
The spectrogram — sound drawn as a picture of frequency against time
Used properly
To generate hypotheses that a playback experiment then tests

Seeing sound

Most discoveries in this field were made by looking at sound rather than listening to it.

A spectrogram draws a recording as an image: time along the bottom, frequency up the side, and loudness as brightness. It sounds like a convenience and it is closer to a prerequisite. Humpback whale song has a structure that unfolds over many minutes — units into phrases, phrases into themes, themes into a song — and no listener can hold that span in mind. Laid out as a picture, the repetition is obvious at a glance. The 1971 paper that founded the subject did not discover a new sound; it discovered the shape of a sound everybody had already heard.

Diagram

From a wiggling line to a picture of a call

A waveform shows loudness against time and hides everything else. The same signal as a spectrogram shows which frequencies are present at each moment, which is where the structure lives.

The same dolphin whistle, twiceAs a waveform — loudness against timeA sound happened, and it was loud in the middle.That is everything this view can tell you.As a spectrogram — frequency / timeThe rising-and-falling contour is the shape —and the shape identifies the individual.Almost every discovery in this field was made by looking rather than listening.Humpback song runs for many minutes — units into phrases, phrases into themes. No ear holdsthat span. Drawn as an image it is obvious, which is how the structure was found in 1971.Seeing structure in a picture is not evidence that the animal uses it.
The same explanation in words

Two panels of the same recording. The upper panel is a waveform: a single line of amplitude against time, showing that a sound occurred and roughly how loud it was. The lower panel is a spectrogram of the same signal, with time along the horizontal axis and frequency vertical, showing a rising then falling contour — the shape that identifies which dolphin is whistling. The waveform cannot show that contour at all.

Two consequences follow. The first is that much of what animals say is inaudible to us: elephant rumbles carry most of their energy below human hearing and were not a research subject until equipment could reach down there, and bat social calls sit above it. The second is that seeing structure is seductive. A pattern visible in a spectrogram is a pattern in the sound, and whether it is a pattern the animal uses is a completely separate question requiring a completely different experiment.

What is AI actually doing here?

Four tasks, none of which is translation.

The short answer

Has AI translated animal communication?

No, and not because the models are not good enough. There is no training data for meaning and no way to make any: translation between human languages is learned from sentences that already mean the same thing in both, and no such pairing exists for any animal.

The four things these systems do are detection — finding calls inside months of continuous audio; classification — sorting them by species or call type; identification — working out which individual produced one; and context prediction — guessing what the animals were doing at the time. Each is a mapping from a sound to a label a human wrote down, and each is genuinely useful. None is a question about meaning. Unsupervised methods can go further and find structure nobody labelled, which is essentially what the 2024 sperm whale analysis did: it discovered that codas vary continuously in tempo and in an added click, and that the variation depends on the surrounding exchange. That is a real and surprising result about the shape of the signal. It identifies no referent for anything, and no whale has been tested on whether it treats those variants as different. The honest description is that machine learning has become an excellent instrument for finding structure in animal sound, and that finding structure has always been the easy half.

Check it for yourself

Machine learning is very good at working out who made a sound and what kind of sound it is. It has no route to what the sound means.

Established

Specialists would state this without hedging. Multiple independent lines of evidence agree.

Supervised learning applied to bioacoustic data performs detection, classification, individual identification and context prediction at rates well above chance across many taxa. These are mappings from acoustic features to human-supplied labels. No training signal for semantic content exists, because no independently established set of meanings is available to learn from.

Who this applies to
A statement about the methods, applying wherever supervised acoustic classification is used on animal sound.
Studied in
Animalia

You may have heard

“AI can now translate what animals are saying.”

Translation between two human languages is learned from pairs of texts that already mean the same thing. Nothing equivalent exists for any animal: there is no parallel corpus, because nobody knows what any call means to start with. What the models are trained on is sound paired with labels a person wrote — species, individual, context — so what they return is the label, not the meaning.

Why we rate it this way, and what the caveats are
EstablishedHigh confidence

The performance results are extensively replicated; the limitation is structural rather than empirical — a supervised model can only learn the labels it is given, and meanings are not among them.

How far it can be extended

The same architecture and the same limitation apply across birds, bats, cetaceans, primates and insects.

Caveats

  • This is not a criticism of the methods. Passive acoustic monitoring at continental scale is only possible because of them.
  • Unsupervised methods can find structure without labels, but finding structure is still not finding meaning.
  • A model can be a hypothesis generator; the elephant work is the model of how that should go.

Still unanswered

  • Whether any training signal for meaning could be constructed — behavioural response is the only obvious candidate and is expensive to collect.
  • How far models trained in one place transfer to another, which is currently a serious practical limit.

Last reviewed 2026-08-31

The evidence (3 studies)

Diagram

What happens between the microphone and the finding

Every step to the left of the dashed line is about sound. Only the playback, on the right, tests the animal.

Everything left of the dashed line is about the recordingRecordmonths of continuous audioDetectfind the calls inside itClassifyspecies, call type, individualAnalyselook for structure across callsPlayback to a live animalthe only step that asks the animalanything at allA model learns to map sound onto labels a person wrote: species, individual, context.Meaning is not among them, because nobody knows what any call means.
The same explanation in words

A left-to-right workflow. A recorder captures continuous audio. A detector finds candidate calls within it. A classifier sorts those calls by type, species or individual. Analysis then looks for structure across the sorted calls. A dashed line separates all of that from the final step: a playback experiment, which reproduces a signal to a live animal and measures the response. Only the final step produces evidence about what the animal does with a signal; everything before it produces evidence about the recording.

What it is actually for

The uncelebrated half, which is doing most of the work.

Away from the language headlines, passive acoustic monitoring has quietly become one of the most important tools in conservation. A moored hydrophone hears whales that no survey vessel would encounter, and hears them at night and in bad weather. A recorder on a post counts bat passes every night for a season. Automated detection turns those recordings into population trends, and population trends are what shipping-lane closures and habitat protections are argued from.

  • Models trained at one site frequently degrade at another — different equipment, different background noise, different seasons.
  • Labelled training data is scarce for everything except a handful of well-studied species, which biases what can be detected at all.
  • A confident false positive is expensive: a detector that reports a rare whale where there is none can misdirect a conservation decision.
  • Rising ocean noise degrades both the animals’ communication and our ability to record it, and the two are hard to separate in a trend.

Follow it further

Claims about this, checked

Things people have heard, and what the evidence actually supports.

The research behind this page

8 studies, newest first. Each one has a page explaining what it found and what it could not show.

2024Nature Communications

Contextual and combinatorial structure in sperm whale vocalisations

Coda structure varies systematically along dimensions the authors name rubato — smooth variation in overall duration — and ornamentation, an extra click added at the end of a coda.

2024Nature Ecology & Evolution

African elephants address one another with individually specific name-like calls

The model identified intended receivers better than chance, and elephants responded more strongly and more quickly to calls originally addressed to them.

2022PeerJ

Computational bioacoustics with deep learning: a review and roadmap

Deep learning has substantially improved detection and classification of animal sounds, particularly for large passive-acoustic datasets.

2016Royal Society Open Science

Individual, unit and vocal clan level identity cues in sperm whale codas

Different coda types carry information at different levels.

2016Scientific Reports

Everyday bat vocalizations contain information about emitter, addressee, context, and behavior

Classifiers recovered caller identity well above chance, and could identify the addressee and the behavioural context — food, mating, sleeping position, perch — from the call alone.

2008Proceedings of the Royal Society B

The Anna's hummingbird chirps with its tail: a new mechanism of sonation in birds

The loud chirp at the bottom of the dive was produced by aeroelastic flutter of the outer tail feathers rather than by the syrinx; removing the relevant feathers removed the sound, and isolated feathers in a wind tunnel reproduced it.

1983Behaviour

The dawn chorus in the great tit (Parus major): proximate and ultimate causes

Singing began earlier at higher light levels and the amount of dawn singing was affected by food supply, consistent with dawn song being timed to a period when foraging returns are poor.

1979The American Naturalist

A quantitative analysis of the dawn chorus: temporal selection for communicatory optimization

Sound transmits substantially further and with less distortion under early-morning atmospheric conditions than later in the day, giving a modelled advantage to singing at dawn that is largest for the frequencies typical of song.

Where to go from here

Each of these follows from something on this page — a relationship in the evidence, a claim people ask about, or the next mechanism along.

How complete this page is, and what it is still missing

NatureHQ publishes its own gaps. This page is at 75% completeness against what we would call a finished subject, and was last reviewed on 2026-08-31. It carries 4 claims and answers 2 mapped search questions.

  • Acoustic localisation — working out where a calling animal is from arrival times at several sensors — is mentioned but not explained.
  • Terrestrial soundscape ecology, which measures whole habitats rather than species, is not covered.
  • The engineering of hydrophone arrays and long-duration recorders is outside scope here.