Models preview

Velma Triage

ms ms
Detect speaker accent
Clip 1
vs
Clip 2
Single-speaker clips, at least 5 seconds each
0:00
0:00
0:00
0:00
0:00 0:00

Summary

Speakers

SpeakerSpeech patternSpeaking time

Behaviors

Behavior detections

Behaviors Velma was configured to listen for on this call, with reasoning and evidence

SpeakerBehaviorReasoning and evidenceConfidence

Transcript

    Speech Summary

    Speakers

    Speaker Emotion pattern Speaking time Language

    Topics & Sentiment

    Behaviors

    Speaker Behavior Model reasoning Confidence

    Transcript

    Upload an audio file or start streaming to see the transcript.
    TimestampVerdictConfidence
    TimestampMusicSpeech

    Every window is checked independently for AI vocals and AI instrumentals — a track can flag on both at once. A dash means that detector had nothing to score in the window; a low score means it checked and the content looks real. Heavily processed or high-production songs may still false-positive.

    WindowVerdictVocal AIInstrumental AI
    WindowEmotion
    WindowAccent

    Whole-clip scores over a fixed set of 42 sound-event labels — crowd reactions, animals, instruments, household sounds. Each score is an independent probability, so several events can be detected at once. Plain speech has no label of its own: a talk-only clip scores low everywhere. Events at 50% or higher count as detected.

    EventScore

    The model turns each clip into a compact voice fingerprint and scores how similar the two fingerprints are, from −1 to 1. It compares how a voice sounds, not what is said — the same speaker saying different things still matches. The verdict is set by this showcase, not the API; treat scores near a boundary as inconclusive. Recording conditions matter: two speakers on the same call score closer than two speakers from unrelated recordings.

    Upload an audio file to see the redacted transcript.
    Raw JSON