Machine learning is very good at working out who made a sound and what kind of sound it is. It has no route to what the sound means.
EstablishedSpecialists would state this without hedging. Multiple independent lines of evidence agree.
Supervised learning applied to bioacoustic data performs detection, classification, individual identification and context prediction at rates well above chance across many taxa. These are mappings from acoustic features to human-supplied labels. No training signal for semantic content exists, because no independently established set of meanings is available to learn from.
- Who this applies to
- A statement about the methods, applying wherever supervised acoustic classification is used on animal sound.
- Studied in
- Animalia
You may have heard
“AI can now translate what animals are saying.”
Translation between two human languages is learned from pairs of texts that already mean the same thing. Nothing equivalent exists for any animal: there is no parallel corpus, because nobody knows what any call means to start with. What the models are trained on is sound paired with labels a person wrote — species, individual, context — so what they return is the label, not the meaning.
Why we rate it this way, and what the caveats are
The performance results are extensively replicated; the limitation is structural rather than empirical — a supervised model can only learn the labels it is given, and meanings are not among them.
How far it can be extended
The same architecture and the same limitation apply across birds, bats, cetaceans, primates and insects.
Caveats
- This is not a criticism of the methods. Passive acoustic monitoring at continental scale is only possible because of them.
- Unsupervised methods can find structure without labels, but finding structure is still not finding meaning.
- A model can be a hypothesis generator; the elephant work is the model of how that should go.
Still unanswered
- Whether any training signal for meaning could be constructed — behavioural response is the only obvious candidate and is expensive to collect.
- How far models trained in one place transfer to another, which is currently a serious practical limit.
Last reviewed 2026-08-31
The evidence (3 studies)
Supports · primary
Computational bioacoustics with deep learning: a review and roadmap
Stowell, 2022 · PeerJ
Reviews the field and its tasks. Every one is detection, classification or identification; none concerns meaning.
Supports · supporting
Everyday bat vocalizations contain information about emitter, addressee, context, and behavior
Prat et al., 2016 · Scientific Reports
A classifier recovered caller, addressee and context from bat calls — information present in the sound, with no test of whether bats use it.
Qualifies · supporting
African elephants address one another with individually specific name-like calls
Pardo et al., 2024 · Nature Ecology & Evolution
Shows the productive way to use these methods: the model found candidate structure, and a playback experiment then tested it on the animals. The model alone would have proved nothing.