The problem: sound is dense, noisy, and unlabeled
An animal call in the wild never arrives clean. It's buried under wind, waves, engines, and other species. And unlike human language, there's no dictionary — no one has told us what a given whale upsweep or wolf howl means. WAVE exists to turn that raw, noisy, unlabeled sound into something a researcher can search, compare, and reason about.
It does that in three layers we call the Technical Triad.
1. Capture — bio-logging sensors and edge AI
WAVE starts where the animals are. Hydrophones and acoustic loggers record continuously, and lightweight models run on the device itself (edge AI). Sub-2MB TinyML models mean a sensor on a remote reef or in a rainforest canopy can detect species presence in real time without shipping gigabytes of raw audio back to a server. This is what makes continuous, large-scale monitoring practical.
2. Resolve — the Superlet Transform
Standard spectrograms blur fine detail. The Superlet Transform, adapted from neuroscience, produces ultra-high-resolution time-frequency images of a call. This resolution matters: Superlet analysis is how researchers showed the Chagos pygmy blue whale song is actually pulsed, not continuous — a structural fact invisible to coarser methods.
3. Understand — foundation models
Finally, WAVE builds on open bioacoustic foundation models like NatureLM-audio and Perch 2.0. These learn shared acoustic structure across thousands of species without supervision — which is why a model trained on one species can help organize another's vocalizations. Perch 2.0 alone classifies close to 15,000 species and set state-of-the-art results on the BEANS and BirdSet benchmarks.
Why the stack matters
Any one layer alone is limited. Together they form a pipeline: capture signal in the field, resolve it into high-fidelity structure, and understand it against everything the models have learned. That's what lets WAVE move from "we heard something" to "here is a comparable, searchable vocal pattern."



