Skip to main content
Research

Research

Detecting AI-generated music is an unsettled technical problem, and we do not pretend otherwise. This section lays out how the detector is tested, what the testing has shown so far, what is still on the to-do list, and the specific ways the method can go wrong.

Updated 2026-08-30 · engine build historical-ensemble-1

Every figure below is either measured or marked as missing — we do not paper over a gap with a guess. Two studies are finished. The one most visitors actually want, a single published accuracy number, is not, and we would rather admit that than hand out a percentage we could not defend under questioning.

Measured

The studies below are finished. Each states its question before its answer, and each was designed so it could be run without waiting on a licence-clear labelled corpus.

Done

Human electronic music false positives

Question. How did a retired internal research build respond to quantised, synthesiser-only procedural references?

Finding. In a historical internal-build study, three of twelve references produced scores above 70%; 83 of 84 transformed files were inconclusive. These findings do not evaluate the current public provider.

Done

Compression and transcoding sensitivity

Question. How far did encoding move the historical research build’s numeric output?

Finding. The historical build moved by 2.4 percentage points on average across the MP3 ladder; ten-second excerpts moved it by 4.4 points. This measured robustness, not accuracy.

Methodology

The full evaluation protocol — objective, dataset design, the holdout rule, the transformation matrix, and precise definitions for every metric — is posted at the benchmark protocol ahead of any results. That order is intentional: writing the method down first stops us reverse-engineering a story once we know what the data says.

For how the product itself works, day to day, see how AI music detection works. For what a result actually means once you have one, plus the failure modes we already know about, see accuracy and limitations.

Planned

These studies are designed but not yet run, so they carry no findings — and won't until the work is actually done.

What do precision, recall, specificity and the false-positive rate look like on verified human and verified AI-generated recordings?

How does the evaluated classifier perform on a generator excluded from protocol development?

Clip duration sensitivity

Planned

At what clip length does a result stop moving around enough to be worth reporting?

Post-production resistance

Planned

How much ordinary mixing, mastering and re-recording does it take before generated audio starts reading as human?

Vocal versus instrumental material

Planned

How much of a measured result is really just vocal artefacts doing the work?

Sources and further reading

  • Benchmark protocol — objective, corpus design, metric definitions, transformation matrix
  • benchmark.json — machine-readable facts, CC BY 4.0, null where unmeasured
  • Editorial standards — how these pages are written and corrected
  • Contact — corrections, replication attempts and corpus contributions