4 ms·
Props for mentioning BirdNET as a potentially more accessible starting point for less technical folks. There are a couple relative advantages of your approach
by refibrillator 2y ago
Props for mentioning BirdNET as a potentially more accessible starting point for less technical folks.
There are a couple relative advantages of your approach that I feel are notable though:
Squeezed wav2vec2 (SEW) architecture leverages Transformer layers and operates directly on time series inputs. But BirdNET converts audio to a spectrogram first and then uses 2D convolution layers (ResNet-like backbone).
This over-representation of inputs to BirdNET implies that SEW will be much more computationally efficient for a given audio classification task (all else held equal).
Plus, simply using a pre-trained SEW model and then training a linear classifier on the embeddings would almost certainly produce strong baseline results. No GPU would be necessary for that.
P.S. Minor typo - precision and recall are confused here:
> “precision” (how many of the animal calls it notices) and its “recall” (the rate at which it makes accurate predictions).
- sdenton4 2y agoFor bioacoustics there are significant problems with domain shift, label shift, label imbalance, and sample bias in training data. You need models to generalize to new data with very different noise profiles than the available training data, and handle significant intraclass variation. The gold standard for input features is a PCEN melspectrogram, largely because it gives useful generalizable features, through compression, normalization, and approximate log scaling of frequency features. Learned frontends tend to overfit training distributions badly - someone finally wrote this up recently, but I'm struggling to find the paper on my phone...
- sdenton4 2y agoFound that paper on frontends: https://www.sciencedirect.com/science/article/pii/S1574954124001158 https://www.sciencedirect.com/science/article/pii/S157495412...
- selimthegrim 2y agoAppreciate the post!