Stop thinking about audio. Start thinking about geometry.
Here's a counterintuitive idea at the heart of modern bioacoustics: the best representation of an animal call is often not the sound itself, but a set of numbers — a vector — that captures its structure. Embed thousands of calls this way and each becomes a point in a high-dimensional space. Similar calls cluster together; different ones sit far apart.
Why this is powerful
Once calls are points in a space, questions that were impossible become routine:
- Which calls are variations of the same "word"? They cluster.
- Does this population have a distinct dialect? Its cluster sits apart from others.
- Is this a rare or novel vocalization? It lands in an empty region.
This is how WAVE's Voronoi visualizations work — they partition the space of vocalizations so behavioral correlations become visible at a glance.
The unsupervised advantage
Crucially, you don't need labels to build this geometry. Foundation models learn it from the audio directly. That's what makes cross-species insight possible: the shape of communication has common structure, even when the species — and the meanings — are wildly different.



