RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The journal · 100 retrospective records ↗
Soundcraft Journal

The journal / Model systems

Model systems / From the journal · 8 September 2016 event · prepared 16 September 2026

WaveNet's listening tests closed half the gap to human speech

DeepMind's 2016 paper and blog report specific listening-test gains from raw-audio generation, not any later product's training data.

Visual for this record: WaveNet's listening tests closed half the gap to human speech
Visual published by lh3.googleusercontent.com, shown for identification of the record. Credit: lh3.googleusercontent.com · source page ↗ Rights: owner-review-pending.

The system

WaveNet is a deep generative model from DeepMind that produces raw audio waveforms one sample at a time, rather than assembling pre-recorded fragments or generating an intermediate spectrogram. DeepMind's blog post, published 8 September 2016, describes it as "a fully convolutional neural network, where the convolutional layers have various dilation factors," letting the model's receptive field grow across thousands of audio samples while remaining fast enough to train. Because it predicts each sample from everything before it, the same architecture can be conditioned on text for speech, on a speaker identity to switch voices, or on nothing at all to generate open-ended audio.

What the documents establish

The accompanying paper, posted to arXiv on 12 September 2016, describes WaveNet as "fully probabilistic and autoregressive," and reports that human listeners rated its speech synthesis as significantly more natural than the best parametric and concatenative text-to-speech systems then in use, in both English and Mandarin. DeepMind's blog quantifies that gap with mean-opinion-score listening tests: 4.21 for WaveNet against 4.55 for actual human speech in US English, versus 3.86 and lower for prior systems, which DeepMind describes as closing more than half the remaining gap to human speech. The blog also reports a separate experiment training WaveNet on a classical piano dataset with no score to condition on, where the model generated novel piano passages purely from what it had learned about the audio itself, distinct from its speech-conditioned use.

Craft and rights

WaveNet's raw-audio, sample-by-sample method is the architectural ancestor of many later neural vocoders and singing-synthesis systems, but the 2016 documents describe a research result measured against text-to-speech benchmarks, not a released consumer product, and they say nothing about the licensing of whatever speech or piano recordings were used to train it. That gap matters for anyone tracing a later commercial voice tool back to WaveNet: the paper establishes that raw-audio generation could sound close to human speech, which is a capability claim, while what training material any downstream product used and under what consent remains a separate question the 2016 sources do not address.

Outcomes and open questions

WaveNet's listening-test scores describe two specific languages and a research-lab evaluation setup from 2016; they do not describe the training data, licensing or evaluation method of any later commercial system that has since adopted a similar raw-audio approach.

  • Does a later product's "WaveNet-based" claim describe shared architecture, shared training data, or neither?
  • What voice recordings trained the text-to-speech systems evaluated in the original listening tests?
  • How does a mean-opinion-score test from 2016 compare with the listening methods used to evaluate audio models now?

Read on its own terms, WaveNet is a well-documented architectural leap in synthesis quality, not evidence about any specific later product's training data or rights posture, and the two claims should not be collapsed into one.

Sources & reading trail

WaveNet: A Generative Model for Raw Audio ↗

Describes the dilated causal convolution architecture and reports specific mean-opinion-score listening-test results against prior systems.

Source published: 8 September 2016 · Retrieved: 16 September 2026

WaveNet: A Generative Model for Raw Audio ↗

The paper's abstract describes the autoregressive, fully probabilistic model and its text-to-speech and music-generation applications.

Source published: 12 September 2016 · Retrieved: 16 September 2026

Papers, reports and standards establish the entry; the craft-and-rights reading is Soundcraft AI editorial analysis. This retrospective draft does not imply the site published on the event date.