Quick Facts
- Mozilla Data Collective secured $10 million from Mozilla Foundation and spun out as a UK entity in April 2026
- The platform hosts over 300 datasets in 286 languages with 187 vetted organizations approved to share data
- Data contributors receive 100% of compensation while the platform charges downloaders a 5% fee to cover infrastructure costs
Mozilla Data Collective spun out as an independent UK entity in April 2026 with $10 million in funding from Mozilla Foundation. The platform aims to create an alternative to current AI data practices that founder and CEO E.M. Lewis-Jong calls extractive and opaque.
The collective launched its alpha version in September 2025 and now features over 300 datasets accessible in 286 languages. As of April 2026, 187 organizations have been vetted to share datasets on the platform.
“Data is power, and that power should belong to people and organizations who are creating that data,” Lewis-Jong said. “We are building a future where data contributors are valued, where communities can choose who benefits from their data.”
The platform charges data downloaders a 5% fee to cover storage, infrastructure and API maintenance costs. Data contributors receive 100% of any compensation for their datasets with no intermediary fees.
Mozilla Data Collective builds on Mozilla’s Common Voice project, which became the world’s largest open speech dataset with 30,000 hours of speech data across 300 languages. The collective started with voice data collected from over one million contributors worldwide.
Data creators retain full ownership and control over their datasets. They can set custom licenses, access rules, compensation requirements and governance structures. The platform combines legal and technical safeguards including custom licenses, authentication checks, watermarking and community moderation.
“We need clean, abundant, contextualized, consentful datasets to build AI models worth having,” Lewis-Jong said in an email interview.
The platform represents Mozilla Foundation’s first social enterprise incubator project. Lewis-Jong previously directed Common Voice and led language data research programs backed by the National Science Foundation, NVIDIA and Gates Foundation.
Mozilla Data Collective positions itself as addressing what it sees as a market failure where data creators are treated as resources to be mined rather than partners. The platform targets developers seeking to connect with communities historically overlooked by mainstream data markets.
Read more: Mozilla Data Collective seeks to build AI’s data economy around trust
This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.
