glotsuite

open tools, corpora and benchmarks for minority languages.

GlotSuite logo

GlotSuite is the family of open tools, corpora and benchmarks I work on for low-resource and minority languages, from identifying the language and script of a text to collecting, indexing and evaluating data for them.

GlotLID

Language identification for 2,000+ labels
An open-source fastText language identifier covering more than 2,000 labels, built for noisy web text and low-resource languages.
GitHub stars for cisnlp/GlotLID

Glot500

A language model and corpus for 500+ languages
Glot500-c, a corpus for 511 mostly low-resource languages, and Glot500-m, a multilingual language model trained on it that improves strongly over XLM-R.
GitHub stars for cisnlp/Glot500

GlotScript

Writing system identification
A resource and tool for identifying writing systems (ISO 15924) for thousands of languages, and for checking which scripts a text is written in.
GitHub stars for cisnlp/GlotScript

GlotCC

Open CommonCrawl corpus for 1,000+ languages
A clean, document-level corpus built from CommonCrawl for more than 1,000 languages, together with the open pipeline that produced it.
GitHub stars for cisnlp/GlotCC

GlotWeb

Web indexing for 400+ minority languages
A web index of verified pages in 400+ languages, many of them missing from major multilingual datasets, with an interactive search demo.
GitHub stars for cisnlp/GlotWeb

GlotOCR Bench

OCR benchmark across 100+ Unicode scripts
A benchmark showing that current OCR models, including frontier models, still struggle beyond a handful of Unicode scripts.
GitHub stars for cisnlp/glotocr-bench

GlotStoryBook

Children's storybooks in 180 languages
A parallel collection of children's storybooks in 180 languages, useful for evaluation and for training in truly low-resource settings.
GitHub stars for cisnlp/GlotStoryBook

GlotSparse

News corpora for under-resourced languages
Collected news text for languages with very little data available online.

Related work built on or around the suite: MaskLID (code-switching language identification with GlotLID) and FineWeb2 (which uses GlotLID for language identification). All publications are on the publications page.