
The system
Speech recognition research needs recordings with known consent and a licence a company can build on, and Mozilla's Common Voice project was built to supply that. Its GitHub repository, retrieved 16 September 2026, describes it as "a platform for collecting speech donations in order to create public domain datasets for training voice recognition-related tools." Contributors read provided sentences through a browser interface and donate the clip; others then validate or reject it on the same platform. It is a general multilingual speech corpus, not a singing-specific one.
What the documents establish
The 2019 Common Voice paper, submitted 13 December 2019, states its most recent release then covered 29 languages, with 38 collecting, over 50,000 contributors, and 2,500 hours of audio — figures the authors call "the largest audio corpus in the public domain for speech recognition" by hours and language count, as of that paper. It explains validation: clips are voted on by other contributors, and a simple majority determines whether a clip is validated, invalidated, or left as unvalidated "other" data. Clips are released as mono, 16-bit MPEG-3 files at 48kHz, a format the paper attributes to browser-based recording. It states plainly that "all of the speech data is released under a Creative Commons CC0 license," and the repository confirms most sentence text is also CC0, while the web app's own code is licensed under the Mozilla Public License 2.0.
Craft and rights
The consent model is closer to affirmative donation than a scrape: a contributor records their own voice to give it away under a stated public-domain destination, a clearer record than repurposing existing broadcast audio. Still, the repository describes intended use as "voice recognition-related tools," written with transcription in mind. Whether a donor anticipated their clip training a generative singing or vocal-style model is a separate question a CC0 licence does not resolve — it removes the legal barrier to reuse, but does not establish that every downstream use matches what a donor pictured.
Outcomes and open questions
The 29-language, 2,500-hour figures come from a paper published in December 2019; Common Voice has kept collecting since, so current totals should be checked against the live project rather than restated as current. What has not changed, per both documents, is the CC0 commitment.
- Does a project using Common Voice for generative voice or music work disclose that reuse to the donating community?
- What are the current language and hour counts against the 2019 baseline this entry cites?
- Does the validation majority-vote process catch mislabelled language or accent metadata at a rate the paper quantifies?
Common Voice's licence question is unusually settled among datasets in this lane; the open question is the scope of consent, not legal permission.
Sources & reading trail
States language and hour counts, the community validation process, audio format and the CC0 licence as of December 2019.
Source published: 13 December 2019 · Retrieved: 16 September 2026
Describes the platform's public-domain purpose and confirms CC0 licensing of sentence text alongside the MPL-2.0 code licence.
Source published: Not established · Retrieved: 16 September 2026
Papers, reports and standards establish the entry; the craft-and-rights reading is Soundcraft AI editorial analysis. This retrospective draft does not imply the site published on the event date.