Skip to content
NatureHQ

Claim check

Has AI translated what whales are saying?

Not supportedNo good evidence supports this, or the evidence points the other way.

No. The analysis behind the headlines is real and careful, and it found more structure in sperm whale clicks than anyone had catalogued. It identified no meanings, and no whale was tested. Translation needs pairs of things that already mean the same in both systems, and for animals no such pairs exist or can be made.

The claim as it circulates

“Artificial intelligence has decoded whale communication — researchers have found a sperm whale alphabet and we are close to talking to them.”

Where you may have met it: Google’s own search suggestions around AI and animal language; Technology and science reporting of a 2024 analysis of sperm whale codas; Fundraising and popular framing around long-term whale communication projects

What was claimed
That machine learning has decoded, or is close to decoding, the meaning of whale communication — often reported as the discovery of a phonetic alphabet.
What was actually observed
Around nine thousand codas recorded from one sperm whale clan in the Eastern Caribbean were analysed with machine-learning methods. Two features — a tempo change across a coda and an extra click added to it — vary continuously and depend on the surrounding exchange, producing a far larger space of distinguishable vocalisations than the discrete catalogue described.
What the evidence supports
That sperm whale vocalisation is finer-grained and more context-dependent than previously documented, and that its expressive capacity — the number of distinguishable signals available — is much larger than assumed. That is a genuine and surprising result about the shape of the signal.
What it does not support
That any coda has an identified meaning. No referent was established for anything, no receiver was tested, and it is therefore not known whether whales themselves treat these fine distinctions as distinct at all. Separately, some coda variation is known to encode who is calling rather than what is being said, and that identity component has not been subtracted before the remainder is read as content.

The structural reason this is not close is worth understanding, because it does not go away with better models. Machine translation between two human languages is learned from parallel text: millions of sentence pairs that already mean the same thing in both. The model never learns what anything means; it learns which strings correspond. For any animal, nothing of the kind exists — and it cannot be assembled, because assembling it would require already knowing what the calls mean, which is the question.

What the models genuinely do is four things: find calls inside months of continuous recording, sort them by species or type, work out which individual produced one, and predict what the animals were doing at the time. Each of those is a mapping from sound to a label a person wrote down, and each is enormously useful — most of what is known about whale distribution now comes from automated detection. None of them is a question about meaning.

The word "alphabet" is doing the damage. An alphabet is a set of symbols that combine to encode meaning. What was found is that codas vary along two dimensions more finely than the old catalogue recorded, and that the variation depends on the surrounding exchange. Calling that an alphabet imports the part nobody has evidence for — that the units encode something — into the name of the finding.

There is a way to use these methods that does work, and the elephant study is the model of it. A classifier found structure suggesting particular calls were aimed at particular animals; the researchers then played those calls back in the field and watched the addressed elephants respond more strongly. The machine generated the hypothesis; the animals tested it. Nothing equivalent has been done for whale codas.

The claims underneath

Each one carries its own evidence, scope and caveats. Expand any of them to reach the studies.

Machine learning is very good at working out who made a sound and what kind of sound it is. It has no route to what the sound means.

Established

Specialists would state this without hedging. Multiple independent lines of evidence agree.

Supervised learning applied to bioacoustic data performs detection, classification, individual identification and context prediction at rates well above chance across many taxa. These are mappings from acoustic features to human-supplied labels. No training signal for semantic content exists, because no independently established set of meanings is available to learn from.

Who this applies to
A statement about the methods, applying wherever supervised acoustic classification is used on animal sound.
Studied in
Animalia

You may have heard

“AI can now translate what animals are saying.”

Translation between two human languages is learned from pairs of texts that already mean the same thing. Nothing equivalent exists for any animal: there is no parallel corpus, because nobody knows what any call means to start with. What the models are trained on is sound paired with labels a person wrote — species, individual, context — so what they return is the label, not the meaning.

Why we rate it this way, and what the caveats are
EstablishedHigh confidence

The performance results are extensively replicated; the limitation is structural rather than empirical — a supervised model can only learn the labels it is given, and meanings are not among them.

How far it can be extended

The same architecture and the same limitation apply across birds, bats, cetaceans, primates and insects.

Caveats

  • This is not a criticism of the methods. Passive acoustic monitoring at continental scale is only possible because of them.
  • Unsupervised methods can find structure without labels, but finding structure is still not finding meaning.
  • A model can be a hypothesis generator; the elephant work is the model of how that should go.

Still unanswered

  • Whether any training signal for meaning could be constructed — behavioural response is the only obvious candidate and is expensive to collect.
  • How far models trained in one place transfer to another, which is currently a serious practical limit.

Last reviewed 2026-08-31

The evidence (3 studies)

Machine analysis found more structure in sperm whale clicks than anyone had catalogued. What it means, nobody knows — and no whale has been tested.

Preliminary

Early results only. Treat as a lead rather than a conclusion.

Analysis of a large coda archive from a single Caribbean clan identified continuous variation in tempo and in the addition of extra clicks, varying with the context of the exchange. This yields a substantially larger space of distinguishable vocalisations than the discrete coda-type catalogue described. No semantic content has been identified and no receiver response to these distinctions has been tested.

Who this applies to
One clan, in one region of the Eastern Caribbean, from one recording archive.Do not extend this beyond the taxa listed — the popular version over-reaches.
Studied in
Physeter macrocephalus

You may have heard

“AI has decoded the sperm whale alphabet.”

An alphabet is a set of symbols that combine to encode meaning. What was found is that codas vary along two dimensions more finely than the old catalogue recorded, and that the variation depends on the surrounding exchange. Calling that an alphabet imports the part nobody has evidence for — that the units encode anything — into the name of the finding.

Why we rate it this way, and what the caveats are
PreliminaryModerate confidence

The analysis is careful and the dataset large for this species. It is preliminary in the sense that matters here: the finding concerns what can be distinguished in a recording, and nothing yet connects those distinctions to whale behaviour.

How far it can be extended

The structure was found in a single clan’s repertoire and has not been sought in others.

Caveats

  • Distinguishable by an analysis is not distinguished by a whale.
  • The comparison to human phonetics in the coverage of this work is an analogy and carries no evidential weight.
  • No playback has tested whether whales respond differently to these variants.

Still unanswered

  • Whether whales themselves treat the fine tempo and extra-click distinctions as distinct.
  • How much of the structure is identity, and how much is left over once identity is accounted for.

Last reviewed 2026-08-31

The evidence (2 studies)

How scientists listen to animals

A model can tell you who made a sound, and when, and what was happening. There is no training data anywhere for what it meant.

Last reviewed 2026-08-31