
The system
Automatic music tagging research needs tracks with known genre, instrument and mood labels at a scale no lab can annotate by hand, so Barcelona's Music Technology Group built the MTG-Jamendo Dataset from music already tagged by its uploaders. The audio comes from Jamendo, a platform where artists post tracks under Creative Commons licences, and the dataset was introduced at the Machine Learning for Music Discovery workshop at ICML 2019. The project's GitHub repository, retrieved 16 September 2026, remains its maintained documentation.
What the documents establish
The repository states the dataset "contains over 55,000 full audio tracks with 195 tags from genre, instrument, and mood/theme categories," drawn from tags uploaders themselves attached on Jamendo. After cleaning, the primary autotagging file covers roughly 55,525 tracks with 87 genre, 40 instrument and 56 mood-theme tags, with prepared splits. Licensing is layered: the code is Apache 2.0, the metadata CC BY-NC-SA 4.0, and each audio file carries its own individual Creative Commons licence, catalogued separately. The README states the dataset "is made available solely for non-commercial research and academic use," and commercial use "requires prior written authorization from Jamendo S.A.," naming a contact. Checked against Jamendo's commercial licensing site, also retrieved 16 September 2026, Jamendo runs a distinct "royalty-free music licensing" product on a larger catalogue, confirming the two are separate systems.
Craft and rights
This is an unusually explicit rights statement: the maintainers themselves have drawn the line between research and commercial reuse in writing, naming who to contact to cross it. A team building a commercial product on MTG-Jamendo without that authorization would work outside its publishers' own terms, regardless of the tracks' Creative Commons licences — a CC licence from an artist and a research-only clause from the compiler are separate permissions, and both apply.
Outcomes and open questions
Because each file carries its own Creative Commons licence rather than one blanket term, even a team with Jamendo's written authorization would still need to check the licence attached to any track it draws on. The repository also lists derivatives, including the Song Describer captioning set built on a subset of tracks, which inherits the same non-commercial framing unless licensed separately.
- Has an AI music product referencing Jamendo-sourced tags or audio obtained the written authorization the repository requires?
- Does the specific per-track licence allow the exact reuse being planned, beyond Jamendo's institutional sign-off?
- Do derivative sets like Song Describer carry terms beyond the base dataset's non-commercial clause?
MTG-Jamendo's documentation leaves little ambiguity about scope; the work left for a downstream user is checking the licence file, not the headline description.
Sources & reading trail
States track and tag counts, licensing layers, and the non-commercial research-use restriction with its authorization contact.
Source published: Not established · Retrieved: 16 September 2026
Shows Jamendo's separate commercial royalty-free licensing product, distinct from the free Creative Commons catalogue the dataset draws on.
Source published: Not established · Retrieved: 16 September 2026
Papers, reports and standards establish the entry; the craft-and-rights reading is Soundcraft AI editorial analysis. This retrospective draft does not imply the site published on the event date.