Research
Detecting AI-generated music is an unsettled technical problem, and we do not pretend otherwise. This section lays out how the detector is tested, what the testing has shown so far, what is still on the to-do list, and the specific ways the method can go wrong.
Updated 2026-08-30 · engine build historical-ensemble-1
Every figure below is either measured or marked as missing — we do not paper over a gap with a guess. Two studies are finished. The one most visitors actually want, a single published accuracy number, is not, and we would rather admit that than hand out a percentage we could not defend under questioning.
Measured
The studies below are finished. Each states its question before its answer, and each was designed so it could be run without waiting on a licence-clear labelled corpus.
Human electronic music false positives
Question. How did a retired internal research build respond to quantised, synthesiser-only procedural references?
Finding. In a historical internal-build study, three of twelve references produced scores above 70%; 83 of 84 transformed files were inconclusive. These findings do not evaluate the current public provider.
Compression and transcoding sensitivity
Question. How far did encoding move the historical research build’s numeric output?
Finding. The historical build moved by 2.4 percentage points on average across the MP3 ladder; ten-second excerpts moved it by 4.4 points. This measured robustness, not accuracy.
Methodology
The full evaluation protocol — objective, dataset design, the holdout rule, the transformation matrix, and precise definitions for every metric — is posted at the benchmark protocol ahead of any results. That order is intentional: writing the method down first stops us reverse-engineering a story once we know what the data says.
For how the product itself works, day to day, see how AI music detection works. For what a result actually means once you have one, plus the failure modes we already know about, see accuracy and limitations.
Planned
These studies are designed but not yet run, so they carry no findings — and won't until the work is actually done.
What do precision, recall, specificity and the false-positive rate look like on verified human and verified AI-generated recordings?
Held-out generator arm
PlannedHow does the evaluated classifier perform on a generator excluded from protocol development?
Clip duration sensitivity
PlannedAt what clip length does a result stop moving around enough to be worth reporting?
Post-production resistance
PlannedHow much ordinary mixing, mastering and re-recording does it take before generated audio starts reading as human?
Vocal versus instrumental material
PlannedHow much of a measured result is really just vocal artefacts doing the work?
Sources and further reading
- Benchmark protocol — objective, corpus design, metric definitions, transformation matrix
- benchmark.json — machine-readable facts, CC BY 4.0, null where unmeasured
- Editorial standards — how these pages are written and corrected
- Contact — corrections, replication attempts and corpus contributions