
The system
Fréchet Audio Distance (FAD) is an automatic evaluation metric proposed by Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek and Matthew Sharifi, researchers at Google, in a paper titled “Fréchet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms”, first submitted to arXiv on 20 December 2018 and later presented, per Google's own publication record, at Interspeech 2019. It adapts the Fréchet Inception Distance used to score generative image models, applying the same statistical idea to audio embeddings so that a system's output can be scored without needing a matched, clean reference recording for every clip.
What the documents establish
The paper itself states the metric is “reference-free,” meaning it compares the statistics of a set of embeddings from generated or enhanced audio against a set from real music, rather than comparing individual output files to individual paired originals. Validated against “a wide variety of artificial distortions,” the authors report FAD reached a correlation coefficient of 0.52 with human perceptual ratings, compared with 0.39 for signal-to-distortion ratio, -0.15 for cosine distance and -0.01 for magnitude L2 distance on the same test set. Those are the paper's own reported numbers, from its own distortion benchmark, not results independently replicated by a third party in this entry.
Craft and rights
FAD's practical value for a producer or evaluator is that it does not require a one-to-one clean reference for every generated clip, which is exactly the situation a text-to-music system creates: there is no ground-truth recording a novel generation is supposed to match. But that same property is a limitation stated by the paper's own design: FAD scores reflect distributional similarity to a chosen reference set of real music, so a low FAD number says a model's outputs resemble that reference set statistically, not that any individual generated track is well-formed, musically coherent, or clear of any particular artist's stylistic fingerprint. Editorially, a FAD score answers a quality-distribution question, not a rights or attribution question.
Outcomes and open questions
Because FAD depends on an embedding model and a chosen reference set, scores from two papers are only meaningfully comparable when both used the same embedding model and the same evaluation set, a condition the original paper does not guarantee will hold across later, unrelated uses of the metric. Readers should watch for papers that report a FAD number without naming the embedding model or reference set used, since that omission makes a headline comparison across systems unverifiable.
- Did the FAD scores being compared use the same embedding model and the same reference audio set?
- Is a reported FAD number being used to claim a generated track sounds musically coherent, or only that its distribution resembles real recordings?
- Has the specific correlation-to-perception figure from the original 2018 paper been re-tested on generative, rather than enhancement, systems?
FAD's original paper is careful about what its own number means, and that caution, more than the metric's growing popularity, is the part worth carrying into any later citation of it.
Sources & reading trail
Defines the reference-free FAD metric, its method, and its reported 0.52 correlation coefficient with human perception.
Source published: 20 December 2018 · Retrieved: 16 September 2026
Confirms Google authorship and states the paper's peer-reviewed venue as Interspeech 2019.
Source published: Not established · Retrieved: 16 September 2026
Papers, reports and standards establish the entry; the craft-and-rights reading is Soundcraft AI editorial analysis. This retrospective draft does not imply the site published on the event date.