All Lab Adds 62 New Languages to Common Crawl
OPEN DATA · COMMON CRAWL · WEB LANGUAGES
Common Crawl is the open web archive sitting underneath a large share of the AI systems people use every day. Until recently, its list of underrepresented languages covered 131 of them. African Languages Lab submitted 1,083 new URLs, and that number is now 193. Sixty-two languages went from absent to present in the record, and we did it by hand.
Before: 4,452 URLs across 131 languages · After: 5,535 URLs across 193 languages · Added: 1,083 URLs · Newly represented: 62 languages · Languages searched: 71 African languages
Why an Archive Nobody Talks About Decides Who AI Understands
Ask why an AI model handles French fluently and struggles with Wolof, and the answer is rarely about the model. It is about what the model read. Most large systems learn from crawled web data, and Common Crawl is the largest open source of it, a free archive built over more than 15 years and measured in petabytes.
Which means a plain rule governs this entire field. If a language is barely on the web, it is barely in the archive. If it is barely in the archive, the models never learn it. And if the models never learn it, the speakers of that language get told, politely and repeatedly, that the technology is not for them.
Common Crawl's Web Languages repository on GitHub exists to interrupt that cycle. It is a community project collecting URLs for underrepresented and low-resource languages so they can be found and indexed. Community blogs. Cultural sites. Regional news portals. The ordinary web of a language, which is exactly what a crawler cannot find if nobody points at it.
How We Actually Did It
There is no clever shortcut here, and that is the point. Our team went looking, by hand, for web content in 71 African languages.
- We searched. Members of the African Languages Lab spent hours hunting down live URLs with genuine content in each language, the kind of sites that never surface in an English-language search.
- We checked each other's work. Every contributor cross-checked another member's list. Working with 71 languages means nobody is fluent in all of them, so peer review was not optional.
- We cleaned the set. The combined list was deduplicated, sorted into the correct language categories and stripped of dead or inaccessible links.
What we submitted was consistent, verified and genuinely representative of the languages involved. A long list of broken links would have made the archive worse, not better.
A language does not disappear from AI because anyone decided to exclude it. It disappears because nobody went looking.
What Changes Because of This
- Visibility for low-resource languages. Sixty-two languages that were missing from the dataset are now part of the record that researchers and model builders draw from worldwide.
- Better multilingual AI. Every system trained on Common Crawl now has access to material it previously did not know existed. This is upstream work, and it benefits everyone building in this space, including people who will never hear our name.
- Cultural preservation with a practical edge. Cultural sites, local news and community content are not just data. They are the living written record of a language, and getting them indexed keeps them discoverable.
This Should Not Be Remarkable
We believe every language carries value, holding culture, identity and shared history. The uncomfortable part is how small the effort was relative to the result. A handful of people, several hours of searching, and 62 languages entered the global digital record. That gap between effort and impact tells you how neglected this work has been.
The repository is open and the work is not finished. If you speak an underrepresented language and know where its web lives, that knowledge is genuinely scarce and genuinely useful.
Help us curate URLs and strengthen the presence of underrepresented languages online. Get in touch with our team.


%20(1).png)

