
The system
Colin Raffel's Lakh MIDI Dataset page, retrieved 16 September 2026, describes a collection of 176,581 unique MIDI files, assembled to support music information retrieval research — symbolic work using the MIDI files alone, and audio-content research using MIDI-derived annotations matched to real recordings. Raffel built it as part of his doctoral work on aligning audio to MIDI. The name, by his own account, plays on the Million Song Dataset it connects to: a lakh is 100,000 in the Indian numbering system.
What the documents establish
The page states the full collection, LMD-full, was deduplicated by MD5 checksum but that "no attempt was made to remove invalid MIDI files," so it holds "a few thousand files which are likely corrupt." A subset of 45,129 files, LMD-matched, was linked to Million Song Dataset entries using dynamic time warping-based audio-to-MIDI alignment, described as "extremely reliable" for confirming a match, though "somewhat invariant to differences in instrumentation," so similar recordings — house-music remixes, the page notes — can be matched incorrectly. The files, the page states, "were scraped from publicly-available sources on the internet." The matching code is published separately in a companion repository, which points back to the dataset page as the distribution point. Raffel releases the page and his compilation under CC-BY 4.0, and states he "did not transcribe any of the MIDI files," and that attributing each file to a transcriber is not feasible since copyright metadata is used inconsistently.
Craft and rights
The CC-BY licence covers Raffel's own work: the compilation, the matching pipeline and the documentation. It says nothing about the copyright status of the roughly 176,000 transcriptions inside, which the page itself says came from unattributed, scraped sources. A MIDI transcription of a copyrighted song is generally a derivative of the underlying composition, regardless of the file's own metadata. Treating "the dataset page is CC-BY" and "the data inside is rights-cleared" as one claim would be a mistake this page does not make — it separates the two — and any system card citing Lakh should be read the same way.
Outcomes and open questions
The alignment method's stated limitation — that instrumentation-invariant matching can pair a MIDI file with the wrong recording — means inherited annotations carry an error rate the page does not quantify beyond calling scores "reliable." Because Lakh remains widely reused for symbolic-music modelling, by the page's own framing as a research tool, watch for whether commercial documentation names Lakh among training data without addressing the compilation-versus-content distinction.
- Does a paper citing Lakh specify which subset — full, matched or aligned — and does that change the rights question?
- Has an entity training on Lakh addressed the transcription-provenance gap its own author describes?
- How large is the mismatch rate for genres the page flags as prone to false matches?
Lakh is transparent about its own limits; the caution is in how confidently a downstream user restates them.
Sources & reading trail
States the dataset's size, matching method, scraped provenance, CC-BY licence and intended research use.
Source published: Not established · Retrieved: 16 September 2026
Publishes the matching and alignment code used to build the dataset and confirms the dataset page as the distribution point.
Source published: Not established · Retrieved: 16 September 2026
Papers, reports and standards establish the entry; the craft-and-rights reading is Soundcraft AI editorial analysis. This retrospective draft does not imply the site published on the event date.