MARSYAS Data Sets (GTZAN Genre Collection)
- Document
- undated document
- Event
- no single event
- Retrieved
- 16 September 2026
The system
GTZAN is a 1,000-track genre-classification dataset, split into ten genres of one hundred 30-second clips each, built by George Tzanetakis for the paper “Musical genre classification of audio signals,” published with Perry Cook in IEEE Transactions on Speech and Audio Processing in 2002. It became, and largely remains, the most cited benchmark for automatic music-genre classification research, a status confirmed independently by the 2013 audit discussed below, which counted its appearance in “at least 100 published works.”
What the documents establish
Tzanetakis' own MARSYAS dataset page describes how the collection was actually assembled: “the database was collected gradually and very early on in my research,” with files gathered in 2000-2001 “from a variety of sources including personal CDs, radio, microphone recordings, in order to represent a variety of recording conditions,” and the page notes there are “no titles” or copyright permissions on file for the individual tracks. A decade later, Bob L. Sturm's peer-reviewed audit, “The GTZAN dataset: Its contents, its faults, their effects on evaluation, and its future use”, documented “repetitions, mislabelings, and distortions” within the set and showed that these faults affect different genre-classification systems unevenly, directly disproving the assumption that shared faults keep cross-system comparisons fair.
Craft and rights
For anyone building or evaluating a genre-classification or generative-audio system today, GTZAN is a caution about benchmark provenance as much as a resource: its own creator's page discloses that the tracks carry no cleared licensing information, and its label sourcing was informal and undocumented at the level modern dataset-disclosure norms would expect. That does not make GTZAN unusable, but it does mean a rights-conscious reader should treat “trained or evaluated on GTZAN” as a statement about an uncleared, historically assembled convenience sample, not a professionally licensed reference corpus.
Outcomes and open questions
Sturm's paper is explicit that its conclusion is not to discard GTZAN outright but “to use it with consideration of its contents,” since the faults it documents do not affect every system identically, which undermines any claim that two papers' GTZAN accuracy figures are automatically comparable. A reader should check whether a given paper used the same train/test split and the same fault-aware precautions Sturm's audit recommends, and watch for newer papers that still cite raw GTZAN accuracy without acknowledging the known duplication and mislabelling issues.
- Does the paper being cited use the exact same GTZAN train/test split as the one it is being compared against?
- Does the paper acknowledge GTZAN's documented repetitions and mislabellings, or report a raw accuracy figure as if the set were clean?
- Is GTZAN being used here as a genuinely comparable benchmark, or only as a familiar name that saves the authors from building a cleaner evaluation set?
GTZAN's popularity outran its documentation for years, and Sturm's audit is the reminder that a benchmark's ubiquity is not the same as its reliability.
Sources & reading trail
Tzanetakis' own description of how the GTZAN dataset was collected in 2000-2001 and its lack of cleared titles or copyright permissions.
Source published: Not established · Retrieved: 16 September 2026
Peer-reviewed audit documenting GTZAN's repetitions, mislabellings and distortions and their uneven effect on system evaluation.
Source published: 6 June 2013 · Retrieved: 16 September 2026
Papers, reports and standards establish the entry; the craft-and-rights reading is Soundcraft AI editorial analysis. This retrospective draft does not imply the site published on the event date.