FMA: A Dataset For Music Analysis
- Document
- 6 December 2016
- Event
- 6 December 2016
- Retrieved
- 16 September 2026
The system
Machine listening research needs full tracks, not thirty-second previews, and in December 2016 a team led by Michaël Defferrard addressed that gap by packaging music from the Free Music Archive into a standard research resource. The paper, "FMA: A Dataset For Music Analysis," was submitted to arXiv on 6 December 2016 and later carried an ISMIR 2017 camera-ready credit. It describes the Free Music Archive as a platform hosting Creative Commons-licensed audio, and the paper's contribution is turning that library into a structured dataset for music information retrieval, a field concerned with browsing, searching and organising large music collections.
What the documents establish
The paper states the dataset comprises 917 gibibytes and 343 days of audio drawn from 106,574 tracks by 16,341 artists across 14,854 albums, arranged in a hierarchical taxonomy of 161 genres. It provides full-length, high-quality audio alongside pre-computed features and track- and user-level metadata, including tags and free-form biographical text, and it proposes a train, validation and test split along with three subsets sized for different research needs. The authors report baseline results for genre recognition to demonstrate the dataset's use. Checked against the Free Music Archive's current site, retrieved 16 September 2026, the platform today presents itself under Tribe of Noise branding as a source of "royalty free music," with a separate discovery-only tier — a commercial framing distinct from the paper's description of a Creative Commons research archive, and a reminder that the platform's own presentation of its catalogue has changed since 2016.
Craft and rights
"Creative Commons-licensed" is not one licence. The Creative Commons family includes terms that require attribution only, terms that forbid commercial use, terms that forbid derivative works, and combinations of all three, and the paper does not claim every one of its 106,574 tracks carries the same permissions. A team using FMA to train a model intended for commercial release, rather than academic evaluation, has to check the specific licence attached to each track it draws on; a dataset-level "Creative Commons" label is a starting point for that check, not a substitute for it. This is an editorial point the paper itself does not need to make, because it was built for research evaluation, where the distinction matters less.
Outcomes and open questions
Neither document establishes whether the specific tracks catalogued in the 2016 paper remain available under the same terms on today's platform, now operating under different branding and a "royalty free" commercial pitch rather than a Creative Commons archive framing. That would require checking individual tracks, not inferring continuity from either source.
- Does a specific FMA track used in training carry an attribution-only licence, a non-commercial restriction, or a no-derivatives term?
- Has ownership or platform policy changed since 2016 in ways that affect tracks already indexed by the paper's dataset release?
- Do genre-recognition baselines reported in 2017 still hold as a reference point for newer MIR models trained on the same splits?
The dataset paper is precise about scale and structure; it is the licence-by-licence detail underneath "Creative Commons" that a commercial user still has to do the work of checking.
Sources & reading trail
Gives track count, artist count, genre taxonomy, licensing framing and intended MIR tasks for the dataset.
Source published: 6 December 2016 · Retrieved: 16 September 2026
Shows the platform's current Tribe of Noise-branded royalty-free framing, distinct from the paper's 2016 description.
Source published: Not established · Retrieved: 16 September 2026
Papers, reports and standards establish the entry; the craft-and-rights reading is Soundcraft AI editorial analysis. This retrospective draft does not imply the site published on the event date.