Every window is checked independently for AI vocals and AI instrumentals — a track can flag on both
at once. A dash means that detector had nothing to score in the window; a low score means it checked
and the content looks real. Heavily processed or high-production songs may still false-positive.
| Window | Verdict | Vocal AI | Instrumental AI |
Whole-clip scores over a fixed set of 42 sound-event labels — crowd reactions, animals,
instruments, household sounds. Each score is an independent probability, so several events can be
detected at once. Plain speech has no label of its own: a talk-only clip scores low everywhere.
Events at 50% or higher count as detected.
The model turns each clip into a compact voice fingerprint and scores how similar the
two fingerprints are, from −1 to 1. It compares how a voice sounds, not what is said —
the same speaker saying different things still matches. The verdict is set by this
showcase, not the API; treat scores near a boundary as inconclusive. Recording
conditions matter: two speakers on the same call score closer than two speakers from
unrelated recordings.