Skip to main content

What an archived synthesiser stress test showed

This page preserves a 2026 test of a retired browser research build. Three of twelve procedural synthesiser references exceeded 70% on its output scale. The current public detector uses a different third-party classifier, so these numbers do not describe today’s service.

Measured 24 August 2026 · historical build historical-ensemble-1 · 168 runs

Reference tracks

12

Synthesiser-only, procedurally rendered

Detections

168

12 tracks × 7 encode conditions × 2 runs

Historical score above 70%

3 of 12

Retired research build; no generator involved

Declined to judge

83 of 84

Inconclusive rather than a verdict

The question

The archived model used handcrafted audio features. This study asked whether procedural, grid-aligned synthesiser music could trigger that feature set even though no generative model appeared anywhere in the production chain.

The corpus intentionally removed acoustic performance and room cues. It provides a reproducible stress case for that historical model, but twelve references cannot establish how often any classifier will err on released electronic music.

What was measured

This uses the same corpus behind the engine behaviour study: twelve tracks, each 20–23 seconds at 44.1 kHz stereo, rendered procedurally by this site’s own sample generator across twelve genre presets. Every track is synthesiser-only, quantised, single-take material — no recorded performance, no third-party audio anywhere in it.

Each track ran through seven conditions — lossless WAV, MP3 at 320, 192, 128 and 64 kbps, a mono downmix, and a 10-second excerpt — twice apiece, for 168 detections on build historical-ensemble-1. Because none of this material used a generative model, elevated values conflict with the documented production process.

Results

Three of twelve references exceeded 70% on the historical output scale. Three of the twelve renders — all synthesiser-only, quantized, single-take material — produced historical scores above 70%. This small procedural set demonstrates a possible confound, and this study reports that risk rather than hiding it.

Elsewhere, the historical build mostly declined to answer. On single-render synthetic reference audio, the historical build declined to issue a verdict in almost every case. That build was configured to abstain rather than guess, and it is a reminder that an inconclusive result is a normal outcome.

Read both findings as properties of the retired build. They neither validate nor invalidate the current provider. Their lasting lesson is procedural: include difficult human controls and preserve an inconclusiveoutcome instead of forcing every file into a binary answer.

How to read this

  • An elevated model output is not authorship evidence. The features driving it are the same ones that describe perfectly ordinary in-the-box production.
  • Inconclusive is the normal outcome here, not a bug. It is not a fault and it is not a soft "yes."
  • Corpus design matters. A benchmark must include quantised, synthesiser-only human material instead of relying only on easy acoustic controls.
  • No file-only measurement can settle authorship. Provenance — project files, stems, version history, someone who can walk through their own arrangement — beats every number on this page.

Limitations

  • This is not a false-positive rate. That would require verified, commercially released human recordings; this corpus is procedurally generated, so it is reproducible but not representative of the wider world.
  • There are no AI-generated tracks in this corpus, so it says nothing about detection rate.
  • Renders run 20–23 seconds. The historical results therefore describe that exact duration range and do not establish a minimum duration beyond the product’s verified five-second input limit.
  • All figures apply only to historical build historical-ensemble-1. The current provider is a different system, so these values are not carried forward — see the changelog.

What we do about it

The planned proposed benchmark includes a dedicated human electronic control arm. Its results remain blank until the current public classifier is evaluated on a verified corpus.

Questions

  • A retired internal research build returned scores above 70% for three of twelve procedural synthesiser references. That result documents a possible confound; it does not measure the current provider.

Cite this page

Quotation with attribution is welcome. Please keep the wording of factual claims intact and link back to the source page.

AI Music Detector. “What an archived synthesiser stress test showed.” Updated 24 August 2026. https://aimusicdetector.com/research/electronic

Related