← Founder Notes
Archive

The language map just grew past the census. meta's speech model now covers 1,600 languages,…

Yethikrishna ROriginal on Threads

the language map just grew past the census. meta's speech model now covers 1,600 languages, including ones with almost no written corpus, per oct 6 release.

the long tail just became a training set.

Context

Meta's Omnilingual ASR is a suite of speech recognition models covering more than 1,600 languages, announced Nov 10, 2025, with an open-source GitHub repo, an arXiv paper and Hugging Face models.

The system targets languages with little written data.

How it compares

The 1,600 plus languages matches Meta's post. It is speech recognition, not a general speech model, and was released in Nov 2025, not Oct 6. Coverage of languages with almost no written corpus was not seen in the excerpts read, so unsupported here, not refuted.

'the long tail just became a training set' is the author's opinion.

Related work

Watch next

  • Check whether Meta shipped an Oct 6 update.

Sources

  1. Meta AI, Nov 10, 2025ai.meta.com
  2. arXiv 2511.09690arxiv.org
  3. Hugging Face: omniASR-LLM-7Bhuggingface.co

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 10 October 2026 at 19:52 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/the-language-map-just-grew-past-the-census-DeUSZw-iNi4" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="The language map just grew past the census. meta's speech model now covers 1,600 languages,…"></iframe>

More notes