Skip to main content

The AI music detection benchmark

A versioned protocol for evaluating the current public classifier, shown beside archived robustness findings from a separate browser research build. The two systems are not interchangeable.

Benchmark v1.1 · engine build historical-ensemble-1 · updated 2026-08-24 · CC BY 4.0

Status

The historical research-build behaviour study below is done, as of 2026-08-24. Accuracy for the current public provider will show not measured until the labelled corpus is built and evaluated. We will not fill those cells with estimates, modelled numbers, or figures borrowed from a vendor — any cell that later gets a number will carry the date and engine build that produced it.

Archived study — browser research build (measured)

Before the current provider integration, a separate browser research build was tested for repeatability and sensitivity to encoding, downmixing and excerpt length. That run covered 168 evaluations across 12 reference tracks under 7 conditions on build historical-ensemble-1.

Corpus. 12 reference tracks (20–23 s, 44.1 kHz stereo) generated procedurally by the site's own sample generator across 12 genre presets. No third-party or copyrighted audio is used. Ladder. lossless WAV · MP3 320 · 192 · 128 · 64 kbps · mono downmix · 10 s excerpt.

Swipe the table sideways to see all columns.

Measured engine behaviour results
MeasurementResult
Repeat determinism0.0 ppEach of the 84 files was analysed twice in the same research run. All 84 pairs returned an identical probability, so that historical build introduced no observed run-to-run variation.
Encode stability (mean |Δp| vs lossless)2.4 ppAcross the four MP3 bitrates, the mean shift was 1.8 pp at 320 kbps, 1.8 pp at 192 kbps, 2.3 pp at 128 kbps, and 3.5 pp at 64 kbps. The largest single shift observed was 26 pp, at 64 kbps.
Mono downmix sensitivity1.2 ppCollapsing stereo to mono moved the score by 1.2 pp on average, with a worst case of 11 pp, even though stereo-field features become unavailable once the channels are combined.
Short-excerpt sensitivity4.4 ppA 10 s excerpt of the same master moved the score by 4.4 pp on average, and by as much as 31 pp. This demonstrates excerpt sensitivity in that historical build; it is not a rule for the current provider.
Abstention behaviour83 of 84 files inconclusiveOn single-render synthetic reference audio, the historical build declined to issue a verdict in almost every case. That build was configured to abstain rather than guess, and it is a reminder that an inconclusive result is a normal outcome.
Synthesiser-only false-positive risk3 of 12 tracks read above 70%Three of the twelve renders — all synthesiser-only, quantized, single-take material — produced historical scores above 70%. This small procedural set demonstrates a possible confound, and this study reports that risk rather than hiding it.
Analysis time406 ms medianThe median was 406 ms and the 90th percentile was 532 ms per track, as measured on the harness, excluding upload time. These harness timings do not describe the current third-party public service.

What this study does not show

  • This is a behavior and robustness study, not an accuracy study. It contains no AI-generated tracks and no verified human commercial recordings, so it produces no detection rate and no false-positive rate.
  • The corpus is generated procedurally, which makes it reproducible but not representative of released music.
  • All figures apply only to the retired historical-ensemble-1 research build. They do not evaluate the current public provider.

Archived research-build parameters

These parameters document the retired browser model used for the archived study. They remain public for reproducibility, but the live detector now uses a server-side third-party service and does not load this ensemble for classification.

Swipe the table sideways to see all columns.

Archived parameters of the historical research build
ParameterValue
Models in the ensemble4Acoustic, spectral-artifact, extended-feature and embedding detectors, combined by weighted vote.
Vote weightsspectral-artefact 2.2 · embedding 1.6 · acoustic 1.0 · extended-features 0.9Archived in public/config/ensemble.json; not used by the public detector.
Reported probability range15–85%Historical browser-build display range; not current public verdict logic.
Shrinkage toward 50%0.94Historical browser-build calibration; not current public processing.
Confidence bandsmoderate ≥ 0.50 · high ≥ 0.75Historical browser-build bands; current public confidence is provider supplied as Low, Medium or High.
Disagreement penaltythreshold 0.20 · penalty 1.1×Historical browser-build rule; not current provider behavior.
Relationship to public detectorSeparate historical research buildThese parameters describe an archived browser model, not the server-side third-party classification used publicly.

Objective

Measure how the detector behaves on audio of known origin — recordings verified as human-made and recordings verified as AI-generated — spanning multiple generators, genres, durations and production conditions. A finished run produces a confusion matrix for each arm, not a single headline figure.

The archived study shows how its own browser build moved under transformations. It is background methodology, not evidence about the current provider’s accuracy or stability.

The holdout principle

Any audio used to choose protocol settings or interpretive rules counts as development data and never enters the evaluation set. Evaluate a classifier on the same material it was developed against and you learn how well it memorised that material, nothing more. One generator's output is also withheld from every design decision and revealed only at publication, so at least one arm measures a model the system was never fitted to.

Transformation matrix

Real-world uploads pass through encoders, loudness processing, and sometimes a phone speaker and microphone on the way in. Each transformation below changes, and often wrecks, the acoustic detail detection relies on — so every arm is scored per condition instead of averaged into one number.

Swipe the table sideways to see all columns.

Proposed transformation conditions and their expected effect
ConditionWhat it changesMeasured effect
Original lossless (WAV/FLAC)Reference condition — nothing removed.Baseline in the behaviour study (12 tracks)
MP3 320 kbpsMild high-band loss; psychoacoustic masking artefacts.1.8 pp mean |Δp| vs lossless (behaviour study)
MP3 192 kbpsVisible spectral ceiling; stereo joint coding.1.8 pp mean |Δp| vs lossless (behaviour study)
MP3 128 kbpsHard low-pass; smeared transients.2.3 pp mean |Δp| vs lossless (behaviour study)
MP3 64 kbpsSevere band loss; the condition most likely to mimic generated audio.3.5 pp mean |Δp| vs lossless, worst single case 26 pp
AAC (where the decoder supports it)Different masking model to MP3; distinct artefact signature.Not yet measured
Loudness normalisationLevel only; no spectral change.Not yet measured
Broad EQ (±3 dB shelves)Moves tonal balance and spectral centroid.Not yet measured
Dynamic-range compressionReduces crest factor and micro-dynamics.Not yet measured
Brickwall limiting to −8 LUFSRemoves peak structure; mimics commercial mastering.Not yet measured
Mild clipping (0.5 dB over)Adds broadband harmonic distortion.Not yet measured
Resample 44.1 → 22.05 → 44.1 kHzDestroys the upper octave permanently.Not yet measured
Mono downmixRemoves all stereo-field evidence.1.2 pp mean |Δp|, worst case 11 pp (behaviour study)
Added noise floor (−60 dBFS)Masks low-level detail; imitates an analogue chain.Not yet measured
10-second excerptReduces the number of independent analysis windows.4.4 pp mean |Δp|, up to 31 pp (behaviour study)
Acoustic re-recording (speaker → microphone)Adds room, mic and re-encode artefacts to generated audio.Not yet measured

Metric definitions

Most published AI-detection accuracy claims are worthless because nobody says what the metric actually is. Below are the five figures this benchmark reports, each defined precisely enough to quote.

Detection rate (TPR)
Share of known AI-generated tracks assigned an AI category by the evaluated classifier, counted per generator rather than pooled across generators.
False-positive rate (FPR)
Share of verified human-produced tracks assigned an AI category. We treat this as the primary metric, because a false accusation costs a musician more than a missed detection does.
Inconclusive rate
Share of tracks assigned an inconclusive outcome. Abstentions are reported on their own, not folded into the accuracy figures.
Calibration error
Calibration measure defined for any future classifier output that can support it; it will not be calculated by treating component indicators as an overall probability.
Encode stability
Mean absolute change in a comparable numeric output between the lossless reading of a master and the same master rendered at 320, 192, 128 and 64 kbps.

Corpus design

690 tracks total, split across generated arms and human control arms. Sampling is stratified by genre and production style rather than pulled from whatever was easiest to find, because both classes mix easy and hard cases — a corpus stacked with easy ones would produce a flattering number that means nothing.

Generated arms

Swipe the table sideways to see all columns.

Generated arms of the benchmark corpus
GeneratorVersionTracksDetection rateInconclusive
Sunov4 / v4.560not measurednot measured
Udiov1.560not measurednot measured
Stable Audio2.040not measurednot measured
ElevenLabs Musiccurrent40not measurednot measured
RiffusionFUZZ40not measurednot measured
Seed Musiccurrent30not measurednot measured
MiniMax Musiccurrent30not measurednot measured
Murekacurrent30not measurednot measured
Held-out generatorundisclosed until publication40not measurednot measured

Human control arms

The control side is stacked, on purpose, toward the cases where a detector is most likely to trip itself up: heavily processed commercial masters and fully human-made electronic music.

Swipe the table sideways to see all columns.

Human control arms of the benchmark corpus
SourceTracksFalse positives
Commercially released studio recordingsHeavily processed human music is the most common false-positive scenario.80not measured
Independent bedroom productionsIn-the-box production with stock plugins looks superficially synthetic.80not measured
Live acoustic and single-room recordingsThe easiest control case; a failure here would be disqualifying.60not measured
Fully synthetic human-authored electronic musicQuantised, synthesiser-only human work is the hardest control case.60not measured
Hybrid human/AI tracks (AI stems in a human arrangement)Reported separately; there is no correct binary answer for these.40not measured

Evaluation procedure

Two full write-ups cover the behaviour figures above: the electronic-music false-positive study, on the three synthesiser-only renders that scored above 70%, and the compression study, on the per-bitrate score shifts.

  1. Every track is held as a lossless master and rendered into the fixed encode ladder described in the compression study, so results can be reported per source condition instead of pooled.
  2. A future run will pin the provider and integration versions. A configuration change starts a new run rather than being merged into the old one.
  3. The provider’s returned category, component indicators, confidence and windows are recorded without creating a new combined score.
  4. One generator is held out entirely and disclosed only at publication, to measure behaviour on material excluded from protocol development.
  5. Abstentions are counted, never dropped. Headline figures always appear next to the inconclusive rate for the same arm.

Results

No controlled benchmark results exist yet. The accuracy cells on this page read not measured because the labelled corpus described above has not been built or run, and we will not paper over that with estimates, vendor numbers, or figures borrowed from other detectors. Once a run finishes, each cell will carry the date, engine build, and generator versions behind it.

Two studies that did not require a labelled corpus are already finished and published: the electronic-music false-positive study and the compression study.

Limitations

Cite these limitations alongside any number pulled from this page — each one has a stable anchor link for exactly that purpose.

  • Lossy compression is the dominant error source. Lossy encoding changes the submitted audio. Future results will be reported per source condition rather than pooled across bitrates.
  • Generator versions move faster than benchmarks. A figure measured against one model release says little about the next release. Any future number will be stamped with the generator version and evaluated provider version.
  • Hybrid tracks have no ground truth. A human arrangement built around a generated stem fits neither class. These tracks are reported as a separate arm and excluded from the headline figures.
  • The benchmark is not adversarial. Deliberate evasion — re-recording, heavy remastering, aggressive time and pitch manipulation — belongs to a separate study. These figures should not be read as robustness against an attacker.
  • No claim about individual tracks. A population-level rate does not carry over to a single file. No figure on this page should be used as evidence about a specific song or person.

How to cite

Published under CC BY 4.0. Cite the version number along with the URL — figures move whenever the engine or the generator versions do.

Swipe the table sideways to see all columns.

AI Music Detector. "The AI Music Detection Benchmark" (v1.1),
2026-08-24. https://aimusicdetector.com/research/benchmark

The same figures are served as JSON at aimusicdetector.com/benchmark.json, with a null value wherever nothing has been measured yet. We welcome corrections, replication attempts and corpus contributions through the contact page; see also accuracy and limitations and the methodology.

Frequently asked questions

  • A percentage with no documented corpus, metric definition or engine build attached cannot be cited by anyone. Locking in the protocol first also keeps us from shaping an explanation around whatever the eventual data shows.