Where scores fall
3,762 scored segments · all languages
One score. Many ways a voice can sound.
Explore speaker consistency through real, cleaned English audio.
300 completed IPTV jobs.
Across languages and historical uploads. Approximate, not a full-corpus census.
Distribution among segments with a recorded score.
3,762 scored segments · all languages
Share scoring at or above your threshold.
score at least 65%
The charts describe a random page sample of completed IPTV jobs, not the 18 handpicked clips. Segments without a score are excluded from the charts.
18 English clips. Six score ranges. All production-cleaned audio.
Watch the spectrogram, follow the playhead,
and compare what you hear.
No cleaned English clips in this range met our selection criteria among 3,810 English-channel manifests checked. This does not mean none exists in the full dataset.
Same spectrogram scale. No added loudness normalization. Clip scores are shown as percentages; they are similarities, not probabilities.
Use these clips to build intuition about a metric, not to mistake a number for a listening test.
Download the study dataThe recorded metric compares speaker embeddings from the two halves of a segment. We display cosine similarities × 100 as percentages. It measures speaker consistency, not speech quality, intelligibility, or the probability that a speaker is correct.
300 randomly selected rows from a PostgreSQL SYSTEM page sample of completed, processed IPTV jobs. It spans historical uploads and languages. We found 8,381 segments: 3,762 scored, 4,383 too short to score, and 236 skipped. Each scored segment has equal weight; the sample is neither language-balanced nor a uniform random sample of all segments.
“Cleaned only” restricts the same sample to segments marked enhanced by the cleaner. Its denominator includes only enhanced segments with a score. The chosen listening clips come from a separate search and are not used to calculate the distributions.
Three clips per available score range, each from a different job and channel. We required a production-enhanced segment, English-only channel metadata and English transcript tags, then reviewed the selected transcript text. We deliberately spread clips across ranges; this is not a representative sample. No good-audio quality filter was applied.
Every selected clip uses SuperSonique step350000. No qualifying example at 95% or above was found among 3,810 English-channel manifests inspected.
Clips were cropped from checksum-verified, production-cleaned Opus files on the original timeline. Playback preserves the recorded gain. Each video displays a fixed-amplitude waveform, a spectrogram with a shared −90 to 0 dB scale, and a synchronized playhead. The spectrum shows 0–12 kHz; the audio remains full-band, mono 48 kHz.
All 18 clips passed audio alignment checks with zero measured sample lag. Transcripts are production ASR output and may contain errors. The displayed speaker score is the recorded manifest value, not a new score calculated from the video.