Skip to main content

Accuracy and limitations

This is the one place we say plainly what's known about how well this detector performs, how to read a result, and every way it's known to fall short. You won't find an accuracy percentage here, because we haven't run the evaluation that would justify quoting one.

Last updated August 2026 · current public-service review

Quick answer

How accurate is AI music detection?

Nobody can honestly reduce this to a single figure, so we don't publish one. Accuracy shifts with the generator, its version, the encode, the genre and how much human production came afterward. Instead we publish the evaluation protocol with per-generator and per-encode arms, and every accuracy cell stays empty until it's actually measured rather than guessed at.

Read the benchmark protocol

Why there is no percentage here

Classification on this site is carried out by a specialist third-party AI music detection service. We don't train or own a detection model, and we haven't run a controlled evaluation of the current system against a labelled corpus. Any accuracy figure we quoted today would be marketing, not measurement.

That is why the report separates its categorical verdict, vocal and instrumental indicators, confidence band and timeline. An Inconclusive outcome means the provider did not return a decisive category; none of these fields supplies a measured error rate.

What would have to be true before we published one

  • A fixed evaluation dataset with documented sampling criteria
  • Separate development and test sets, with no leakage between them
  • Generator-held-out testing — evaluating on generators absent from tuning
  • Genre and language diversity, including instrumental-only material
  • Audio-quality variation: lossless, 320 kbps, 128 kbps, 64 kbps, resampled
  • A published confusion matrix with precision, recall and both error rates
  • A reproducibility note so others can check the claim

The full design is public at the benchmark protocol, and the status of every planned and completed study is tracked on the research hub.

What has and hasn't been measured

  • Measured only on a retired browser research build: repeatability and output movement under encoding, mono downmixing, excerpting, and a small synthesiser-only reference set. These archived observations appear in the compression study and electronic-music study; they are not measurements of the current third-party classifier.
  • Not measured: precision, recall, specificity, the false-positive rate on verified human recordings, the detection rate per generator, and behaviour on a held-out generator. These read not measured on the benchmark, and we don't estimate them anywhere on the site.

What can be interpreted without an accuracy claim

The primary outcome is a category returned by the specialist service. Vocal and instrumental percentages are supporting component indicators, not inputs to a site-defined overall score. Confidence is reported independently as Low, Medium or High. None of these fields supplies a measured error rate, and none identifies an author or generator. The report guide explains each field on what your result means.

What the tool cannot do at all

  • It cannot prove how a track was made. Acoustic analysis observes properties of a signal. Authorship is a fact about a process, and no measurement of the output recovers it.
  • It cannot attribute a generator. There is no Suno, Udio or ElevenLabs classifier here. When the evidence leans generated, the report says “unknown AI generator” and stops there.
  • It cannot separate stems. A synthetic vocal over human instrumentation is measured as one mixed signal, not two separate components.
  • It cannot look a song up. Nothing is matched against a database of known AI releases, so a widely circulated generated track gets no special treatment.
  • It cannot produce forensic evidence. The output is an estimate and isn't admissible as proof in a copyright, disciplinary or employment process. See the disclaimer.

Conditions that complicate interpretation

Encoding, remastering, re-recording and other transformations change the signal presented to a classifier. Hybrid tracks combine sources with different histories, while generator updates can introduce material unlike earlier examples. The current provider has not supplied this site with condition-specific error rates, so these cases call for more caution rather than a numeric correction.

The accepted duration is five seconds to 15 minutes, but acceptance is not a promise that any particular excerpt contains enough information for a reliable conclusion. Speech, crowd noise, field recordings and near-silence are outside this music detector's intended use.

False positives and false negatives

A false positive can damage a real musician’s reputation, while a false negative can let generated material pass a screen. The current service’s rates for either error are not yet measured by this site. That uncertainty is why no report should decide a copyright, employment, academic or enforcement matter on its own.

Why independent validation matters

This page, the explainer and the benchmark protocol were written by the people who operate the detector. Self-evaluation carries an obvious incentive problem, and no amount of careful wording removes it. The machine-readable facts are published at /benchmark.json under CC BY 4.0 specifically so someone else can check them, and any independent criticism will be published alongside our response.

How to get the most reliable reading

  • Use the highest-quality copy of the file you can obtain, ideally lossless.
  • Use a representative section of music rather than silence, speech or incidental noise.
  • Treat an inconclusive result as information, not as a failed attempt.
  • Weigh provenance above acoustics every time — the full process is on is this song AI generated.

How results should be used

  • As one input among several, never as the deciding factor
  • Alongside provenance: session files, stems, release history, artist conversation
  • Never as evidence in a copyright, employment, academic or enforcement decision
  • Never as the basis for a public accusation

See the detection-results disclaimer for the formal statement, read how the detection works, or go back and analyse a track.

Cite this page

Quotation with attribution is welcome. Please keep the wording of factual claims intact and link back to the source page.

AI Music Detector. “Accuracy and limitations of AI music detection.” Updated August 2026. https://aimusicdetector.com/accuracy