RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The journal · 100 retrospective records ↗
Soundcraft Journal

The journal / Training data & rights

Training data & rights / From the journal · 5 March 2017 event · prepared 16 September 2026

AudioSet ships YouTube references, not the audio itself

Google's AudioSet documents 2 million labelled 10-second clips as YouTube references and precomputed features, not redistributed audio files.

research.google.comprimary record

AudioSet

Document
5 March 2017
Event
5 March 2017
Retrieved
16 September 2026
No visual was published with this record, so its primary document stands in its place.

The system

Building an audio event classifier requires more labelled examples than any lab can record, so in 2017 Google's Sound Understanding group released AudioSet: not a library of audio files, but an ontology and labelled references into YouTube. The AudioSet homepage, retrieved 16 September 2026, describes "an expanding ontology of 632 audio event classes" arranged as a hierarchical graph, covering human and animal sounds, musical instruments and genres, and everyday sounds, introduced in an ICASSP 2017 paper.

What the documents establish

The homepage states the dataset comprises 2,084,320 human-labelled 10-second clips, drawn from YouTube videos nominated by metadata and content-based search, totalling roughly 5.8 thousand hours across 527 classes used for evaluation. The download page spells out what is distributed: CSV files listing, per segment, a YouTube video ID, start and end time, and labels — plus, separately, precomputed 128-dimensional features extracted with a VGG-inspired model, "PCA-ed and quantized to be compatible with" YouTube-8M features. No raw audio is part of either download; an example file the page reproduces records its own creation as "Sun Mar 5 10:54:25 2017." Labels are released under CC BY 4.0 and the ontology under CC BY-SA 4.0, both by Google. The page also discloses an internal quality process: an assessment found "a substantial number of sound classes had poor accuracy," prompting a rerating effort described as "about 50% complete."

Craft and rights

The reference-only design carries a rights consequence: Google's CC BY 4.0 licence covers the labels and ontology it produced, not the underlying YouTube recordings, which remain governed by each uploader's own rights and can be removed or made private at any time. A model trained on AudioSet-derived features is training on a machine-computed representation, and the licence on that representation is distinct from any claim about rights to the original recordings. Because the ontology includes musical genres and instruments as categories, any disclosure listing AudioSet among training data should be read with this reference-versus-recording distinction in mind, not as one blanket clearance.

Outcomes and open questions

The download page's own disclosure that class-level accuracy was initially poor, with rerating only about half complete when written, is a live caveat, not a resolved one. Because the dataset is only as complete as the videos it still points to, its practical size shrinks as source videos disappear — the same limit affecting MusicCaps, which draws its clips from this pool.

  • What share of AudioSet's original 2,084,320 referenced segments are still retrievable from YouTube today?
  • Has the rerating process, described as roughly half complete, been finished, and for which classes?
  • Does a training-data disclosure distinguish AudioSet's precomputed features from a claim of rights to the recordings?

AudioSet's own documentation is unusually candid about both its reference-only structure and its labelling limits; both deserve to travel with any citation of it.

Sources & reading trail

AudioSet ↗

States the ontology size, clip count, hours of audio and evaluation class count.

Source published: Not established · Retrieved: 16 September 2026

AudioSet Download ↗

Describes the CSV-and-features distribution format, the CC licences, and the internal quality/rerating disclosure.

Source published: Not established · Retrieved: 16 September 2026

Papers, reports and standards establish the entry; the craft-and-rights reading is Soundcraft AI editorial analysis. This retrospective draft does not imply the site published on the event date.