AN AUDIO EXPLORATION

Hear what the
numbers mean.

One score. Many ways a voice can sound.
Explore speaker consistency through real, cleaned English audio.

Explore the clips
AVERAGE CONSISTENCY60.4%Across scored segments
MEDIAN CONSISTENCY62.1%Half score above, half below
SCORES AVAILABLE44.9%3,762 of 8,381 sampled segments
THE SAMPLE

300 completed IPTV jobs.
Across languages and historical uploads. Approximate, not a full-corpus census.

01 / THE NUMBERS

The shape of consistency.

Distribution among segments with a recorded score.

Where scores fall

3,762 scored segments · all languages

SHARE OF SCORED SEGMENTS
Choose a range to explore its audio below.CONSISTENCY SCORE (%)

How much makes the cut?

Share scoring at or above your threshold.

41.1%

score at least 65%

65%

The charts describe a random page sample of completed IPTV jobs, not the 18 handpicked clips. Segments without a score are excluded from the charts.

02 / THE AUDIO

Less guessing.
More listening.

18 English clips. Six score ranges. All production-cleaned audio.

CLEANED AUDIO ONLY

Watch the spectrogram, follow the playhead,
and compare what you hear.

Same spectrogram scale. No added loudness normalization. Clip scores are shown as percentages; they are similarities, not probabilities.

03 / THE FINE PRINT

A little context
goes a long way.

Use these clips to build intuition about a metric, not to mistake a number for a listening test.

Download the study data
What does speaker consistency measure?

The recorded metric compares speaker embeddings from the two halves of a segment. We display cosine similarities × 100 as percentages. It measures speaker consistency, not speech quality, intelligibility, or the probability that a speaker is correct.

What population do the charts describe?

300 randomly selected rows from a PostgreSQL SYSTEM page sample of completed, processed IPTV jobs. It spans historical uploads and languages. We found 8,381 segments: 3,762 scored, 4,383 too short to score, and 236 skipped. Each scored segment has equal weight; the sample is neither language-balanced nor a uniform random sample of all segments.

“Cleaned only” restricts the same sample to segments marked enhanced by the cleaner. Its denominator includes only enhanced segments with a score. The chosen listening clips come from a separate search and are not used to calculate the distributions.

How were the listening clips selected?

Three clips per available score range, each from a different job and channel. We required a production-enhanced segment, English-only channel metadata and English transcript tags, then reviewed the selected transcript text. We deliberately spread clips across ranges; this is not a representative sample. No good-audio quality filter was applied.

Every selected clip uses SuperSonique step350000. No qualifying example at 95% or above was found among 3,810 English-channel manifests inspected.

What audio and visuals am I seeing?

Clips were cropped from checksum-verified, production-cleaned Opus files on the original timeline. Playback preserves the recorded gain. Each video displays a fixed-amplitude waveform, a spectrogram with a shared −90 to 0 dB scale, and a synchronized playhead. The spectrum shows 0–12 kHz; the audio remains full-band, mono 48 kHz.

All 18 clips passed audio alignment checks with zero measured sample lag. Transcripts are production ASR output and may contain errors. The displayed speaker score is the recorded manifest value, not a new score calculated from the video.

THE FULL LISTENING SET 18 clips · 3:15

Cleaned English audio, ordered from lower to higher speaker consistency.