RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The journal · 100 retrospective records ↗
Soundcraft Journal

The journal / Model systems

Model systems / From the journal · 9 January 2024 event · prepared 16 September 2026

MAGNeT traded autoregression for speed, not new training data

Meta's masked-decoding paper cuts latency by up to tenfold while its own hours figure for licensed data drifts across documents.

arxiv.orgprimary record

Masked Audio Generation using a Single Non-Autoregressive Transformer

Document
9 January 2024
Event
9 January 2024
Retrieved
16 September 2026
No visual was published with this record, so its primary document stands in its place.

The system

MAGNeT is a Meta audio-generation model, described in a paper submitted 9 January 2024 and released through the same AudioCraft toolkit that hosts MusicGen. Where MusicGen predicts tokens one step at a time, MAGNeT is non-autoregressive: it masks and predicts spans of audio tokens in parallel across a small number of decoding steps, a design the paper says targets the latency that made MusicGen unsuitable for interactive use in a digital audio workstation.

What the documents establish

The paper reports MAGNeT reaches results comparable to MusicGen's autoregressive baseline while running up to seven times faster overall, and up to ten times lower latency at small batch sizes, attributing the gain to masking tokens in 60-millisecond spans and rescoring candidate spans with an external pretrained model before the next decoding step. On training data, the paper's appendix states MAGNeT follows, in its own words, the same setup as MusicGen, and cites 20,000 hours of licensed music. The AudioCraft repository's current MAGNeT documentation, however, states the figure as 16,000 hours from the same three sources — an internal 10,000-track library plus the ShutterStock and Pond5 catalogues — a discrepancy this entry flags rather than resolves, since neither document explains the difference.

Craft and rights

For a producer working inside a DAW, MAGNeT's real contribution is speed, not a new creative capability: the paper is explicit that its output quality is comparable to, not better than, the autoregressive alternative it is faster than. Because MAGNeT reuses MusicGen's training pool rather than assembling a new one, its rights profile is inherited rather than independently documented, which means any licensing caveat that applies to MusicGen's stock-catalogue tracks applies here too. The unresolved hours figure is a reminder that even a well-documented lineage can drift in its own aftercare.

Outcomes and open questions

AudioCraft's repository now also lists JASCO, a newer chord-, melody- and drum-conditioned model added after MAGNeT, so MAGNeT should not be treated as Meta's current flagship text-to-music system; it is one generation in an actively maintained toolkit. A reader should check the repository's present model list before assuming any single paper describes the newest available option.

  • When two of a company's own documents disagree on a training-data figure, which should you trust, and why?
  • Does a faster non-autoregressive model change what a licensed training pool actually permits downstream?
  • Has a newer model in the same toolkit superseded the one you are evaluating?

MAGNeT is best read as a latency paper layered on MusicGen's existing rights profile, not as a fresh data-sourcing story of its own.

Sources & reading trail

Masked Audio Generation using a Single Non-Autoregressive Transformer ↗

States the span-masking and rescoring method, the latency comparison to MusicGen, and the appendix's 20,000-hour licensed-data claim.

Source published: 9 January 2024 · Retrieved: 16 September 2026

AudioCraft GitHub repository (MAGNeT documentation) ↗

States a 16,000-hour licensed-data figure for MAGNeT and confirms JASCO's later addition to the same toolkit, as the repository reads on 16 September 2026.

Source published: Not established · Retrieved: 16 September 2026

Papers, reports and standards establish the entry; the craft-and-rights reading is Soundcraft AI editorial analysis. This retrospective draft does not imply the site published on the event date.