๐Ÿ BuzzASR โ€” A Swarm of 100+ Monolingual Speech Recognition Models

One specialist model per language. BuzzASR is a suite of 102 monolingual ASR models, each a fine-tune of Whisper-large-v3 specialized to a single language.

The work has been accepted at EMNLP 2026 ยท built by Lemn Lab in collaboration with EleutherAI.

The idea

Big multilingual models spread themselves thin across hundreds of languages. We asked whether a single specialist, the same size as Whisper, could beat the giant generalists. It can to a considerable extent :)

Highlights

A few standouts (combined FLEURS + Common Voice, normalized CER %)

Language BuzzASR CER Whisper zero-shot Note
Cantonese 13.0 36.4 SOTA โ€” 2x better than the next-best system
Punjabi 8.8 42.3 SOTA
Mongolian 5.2 38.2 SOTA
Korean 4.6 5.7 SOTA
Amharic 9.1 191 >20x reduction over Whisper

Find your language

Browse all 102 models in the Models tab above, or go to huggingface.co/BuzzASR/<language>.

Usage

from transformers import WhisperForConditionalGeneration, WhisperProcessor
model = WhisperForConditionalGeneration.from_pretrained("BuzzASR/mongolian")
proc  = WhisperProcessor.from_pretrained("BuzzASR/mongolian")
# the language/task prompt is baked in โ€” just call model.generate(input_features)

Links