RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The journal · 100 retrospective records ↗
Soundcraft Journal

The journal / Standards & provenance

Standards & provenance / Systems note · Entry note · prepared 16 September 2026

IRCAM's descriptors became scaffolding for MPEG-7

IRCAM's CUIDADO-era research supplied a descriptor taxonomy that MPEG-7 adopted, work that predates and underlies today's audio-recognition tools.

Visual for this record: IRCAM's descriptors became scaffolding for MPEG-7
Visual published by storage.ressources.ircam.fr, shown for identification of the record. Credit: storage.ressources.ircam.fr · source page ↗ Rights: owner-review-pending.

The system

IRCAM, the Paris music-technology research institute, contributed audio-description research that fed into MPEG-7, the ISO/IEC multimedia-metadata standard for describing content rather than encoding it. Between 2000 and 2003, an IRCAM-led team ran the CUIDADO project, developing descriptors meant to let software characterize sound by features such as timbre, energy and pitch, work that an IRCAM-hosted 2024 account of the project describes as aimed at building 'a taxonomy of descriptors for the new MPEG-7 standard.' The resulting descriptor set predates today's AI-generated-content detectors by two decades, but the underlying idea, describing audio numerically so software can compare or classify it, underlies audio fingerprinting and content-description tools used now.

What the documents establish

IRCAM researcher Geoffroy Peeters' own technical report, A Large Set of Audio Features for Sound Description, dated by the document itself to 23 April 2004, lays out a taxonomy of low-level audio descriptors, including several named directly after their MPEG-7 identifiers, such as LogAttackTime, AudioSpectrumCentroid and TemporalCentroid, alongside parallel descriptors from IRCAM's own CUIDADO naming scheme. The report groups descriptors into temporal, energy, spectral, harmonic and perceptual categories, consistent with the categories the 2024 IRCAM account attributes to the CUIDADO taxonomy. Together the two documents establish that IRCAM's contribution was descriptor-level: defining what numeric features a sound-description system should extract, rather than a full detection or identification product. Neither document describes IRCAM's separate, present-day AI-detection offerings as part of this earlier standards history, and this entry does not extend that connection beyond what the documents state.

Craft and rights

For a working engineer, the practical legacy of this research is indirect. MPEG-7 style descriptors do not themselves identify who owns a recording or whether it was used with consent; they describe acoustic properties that later systems, including fingerprinting and recognition services, build on to match or classify audio. The rights question this raises is about what description infrastructure makes possible: a standard that lets software compare sounds efficiently also makes it easier to detect unlicensed reuse, moderate uploads at scale, or, in principle, screen for AI-generated material, but the standard itself assigns no rights and settles no ownership question. Crediting IRCAM's specific technical contribution, rather than treating 'AI detection' as one continuous lineage from 2000 to now, is an editorial distinction this entry treats as important.

Outcomes and open questions

MPEG-7's descriptor framework continues to inform later audio-fingerprinting and music-information-retrieval research, but the specific pathway from CUIDADO-era descriptors to any named contemporary AI-detection product is not documented in the sources reviewed here, and should not be assumed without checking a given vendor's own technical materials.

  • Does a given audio-description or fingerprinting tool trace its descriptors to a documented standard, or to an undisclosed proprietary method?
  • Does describing a sound's acoustic features tell you anything about who is entitled to use it?
  • When a vendor claims a lineage back to an academic standard, does the vendor's own documentation actually support that claim?

Standards work like this rarely makes headlines on its own, but it is the scaffolding later recognition and provenance tools are built on, which is exactly why its scope is worth stating precisely.

Sources & reading trail

A Large Set of Audio Features for Sound Description (Classification and Similarity) ↗

IRCAM's own technical report defining the audio-descriptor taxonomy, including descriptors matched to MPEG-7 identifiers.

Source published: 23 April 2004 · Retrieved: 16 September 2026

'What's in a name?' IRCAM, MPEG-7, and the Standardization of Audio Description ↗

IRCAM's own media archive describing the CUIDADO project's 2000-03 timeline and its aim of building an MPEG-7 descriptor taxonomy.

Source published: 22 March 2024 · Retrieved: 16 September 2026

Papers, reports and standards establish the entry; the craft-and-rights reading is Soundcraft AI editorial analysis. This retrospective draft does not imply the site published on the event date.