glotsuite
open tools, corpora and benchmarks for minority languages.
GlotSuite is the family of open tools, corpora and benchmarks I work on for low-resource and minority languages, from identifying the language and script of a text to collecting, indexing and evaluating data for them.
GlotLID
Language identification for 2,000+ labels
An open-source fastText language identifier covering more than 2,000 labels, built for noisy web text and low-resource languages.
Glot500
A language model and corpus for 500+ languages
Glot500-c, a corpus for 511 mostly low-resource languages, and Glot500-m, a multilingual language model trained on it that improves strongly over XLM-R.
GlotScript
Writing system identification
A resource and tool for identifying writing systems (ISO 15924) for thousands of languages, and for checking which scripts a text is written in.
GlotCC
Open CommonCrawl corpus for 1,000+ languages
A clean, document-level corpus built from CommonCrawl for more than 1,000 languages, together with the open pipeline that produced it.
GlotWeb
Web indexing for 400+ minority languages
A web index of verified pages in 400+ languages, many of them missing from major multilingual datasets, with an interactive search demo.
GlotOCR Bench
OCR benchmark across 100+ Unicode scripts
A benchmark showing that current OCR models, including frontier models, still struggle beyond a handful of Unicode scripts.
GlotStoryBook
Children's storybooks in 180 languages
A parallel collection of children's storybooks in 180 languages, useful for evaluation and for training in truly low-resource settings.
GlotSparse
News corpora for under-resourced languages
Collected news text for languages with very little data available online.
Related work built on or around the suite: MaskLID (code-switching language identification with GlotLID) and FineWeb2 (which uses GlotLID for language identification). All publications are on the publications page.