Research
4 min read
All Lab Adds 62 New Languages Added on Common Crawl
We added 1,083 URLs to Common Crawl's web languages project, bringing 62 new languages into the repository.
Words by
African Languages Lab

African Languages Lab has contributed a curated set of web addresses to Common Crawl's web languages project, one of the most widely used open sources of web data for language technology. Our submission added 62 languages that were not represented in the repository before.
Why this matters
Common Crawl is one of the largest public records of the web, and it shapes which languages modern language technology can see. English makes up over 40% of it, while the 24 African languages it tracks, including Swahili, Hausa, Yoruba and Zulu, together account for less than 0.05%. When a language is missing from that record, it is far harder to build search, tools and AI that work for the people who speak it.
What we added
Our team submitted 1,083 new URLs after searching for content in 71 African languages. The sources include cultural websites, local news outlets, community content and language-specific portals, the places where everyday language is actually written.
What changed
The repository grew from 4,452 URLs across 131 languages to 5,535 URLs across 193 languages.
62 languages are now represented for the first time.
1,083 new URLs were added in total.
How we did it
This was careful, manual work. Team members searched the web language by language, cross-checked each other's findings, removed duplicates and broken links, and verified the language of every entry before it was submitted. Quality mattered more than volume: a smaller set of accurate, well-labelled sources is more useful to researchers than a large, noisy one.
What comes next
Every language deserves a place in the global digital record. More African language content in shared resources means better multilingual AI research, and it helps preserve cultural and linguistic heritage online. This contribution sits alongside All Voices, our community platform for speech and text data, and our published research on low-resource African NLP. We will keep adding sources as we find them, and we welcome suggestions from communities who know where their languages live online.

