
The system
Stable Audio Open is an open-weights text-to-audio diffusion model from Stability AI, documented in a research paper submitted 19 July 2024 and a companion announcement. It generates variable-length stereo audio up to 47 seconds at 44.1kHz from a text prompt, using a latent diffusion-transformer architecture the paper describes as a variant of the commercial Stable Audio 2.0 model with one substitution: T5 text conditioning in place of the CLAP conditioning the commercial line uses.
What the documents establish
The paper states the full pipeline's parameter counts — a 156-million-parameter autoencoder, a 109-million-parameter T5 text encoder, and a 1,057-million-parameter diffusion transformer — and, unusually for this lane, an exact accounting of training data: 486,492 recordings totalling 7,300 hours, of which 472,618 recordings (6,330 hours) come from Freesound and 13,874 (970 hours) from the Free Music Archive, all licensed under CC-0, CC-BY or CC-Sampling+. The paper describes a specific filtering step, identifying probable music clips with an audio tagger and sending flagged material to a content-detection service to remove anything that might be copyrighted music hiding inside a field recording. The Stability blog post confirms the model's weights are released under Stability AI's own Community Licence, which permits non-commercial use and commercial use below one million dollars in annual revenue.
Craft and rights
Stable Audio Open's training-data accounting is the most granular in this batch of entries, down to exact recording counts by source, and that transparency is itself the notable craft fact: a producer can check whether Creative Commons field recordings and sound-effect libraries suit a given sound-design task, which a model trained on unnamed pop catalogues does not let anyone verify. The rights question that remains is about the commercial Stable Audio product line: the open weights carry the revenue-threshold Community Licence described above, but Stability's separate, subscription-based Stable Audio 2.0 service at stableaudio.com operates under its own commercial terms, and the two should not be treated as interchangeable licences for the same underlying capability.
Outcomes and open questions
The paper reports its FDopenl3 scores as competitive with the state of the art but not uniformly ahead of closed commercial systems, a limitation the authors attribute partly to the deliberately narrower, rights-cleared training set. Readers should watch whether this CC-licensed-only approach becomes a repeatable model for open audio models or remains a one-off research demonstration.
- Does a Creative Commons licence type (CC-0, CC-BY, CC-Sampling+) carry different downstream obligations for your use?
- Is the revenue-threshold model licence still current, or has Stability AI since revised its terms?
- Would a smaller, rights-cleared training set change the range of sounds you can reliably generate?
Stable Audio Open shows that exact, source-attributed training-data accounting is achievable in a public paper; what it costs in coverage against a larger, less-documented corpus is the open trade-off the authors themselves report.
Sources & reading trail
States the architecture parameters and the exact Freesound/FMA recording and hour counts by Creative Commons licence type.
Source published: 19 July 2024 · Retrieved: 16 September 2026
Confirms the open weights are released under Stability AI's Community Licence with its revenue threshold, as the page reads on 16 September 2026.
Source published: Not established · Retrieved: 16 September 2026
Describes the separate, subscription-based commercial Stable Audio product line at stableaudio.com.
Source published: Not established · Retrieved: 16 September 2026
Papers, reports and standards establish the entry; the craft-and-rights reading is Soundcraft AI editorial analysis. This retrospective draft does not imply the site published on the event date.