Methodology
How the corpus was built
Guess Who Said It? is a small experiment about voice. Political streamers, podcasters and pundits talk for hours every week. Strip away the face, the set and the intro music, and how much of what makes each of them recognisable survives in the words alone?
The corpus currently holds 2,492 quotes from 126 speakers: the hosts of the channels the quotes were drawn from, plus the guests, politicians and public figures who appeared on them.
How the quotes were collected
- Sources. Public recordings of political streams, podcasts and shows, the kind anyone can watch, turned into transcripts. Everything in the corpus was said on air, in public.
- Sampling. Passages were sampled at random from across each channel's output rather than from the most-clipped moments, so the corpus reflects how people actually talk, not just their highlights.
- Extraction. A language model read each sampled passage and proposed the most interesting, self-contained lines in it, then ranked them.
- Review. A person read every candidate, dropped the weak ones, and confirmed the speaker by listening to the recording. The text was checked against the audio so that a word appears only if it was actually spoken; the only things removed were filler sounds such as "uh" and "um". Where a line needed context to stand on its own, it was added in square brackets. Each quote is tied to the second of the source video.
- Balance. Speakers with only a handful of usable quotes are still included, but the game draws on people in proportion to how much they speak, so the corpus is not dominated by any one channel.
Two ways to play
- Standard Quiz. The regulars: hosts and headline guests whose voices you could reasonably know.
- Ultra Quiz. The whole corpus, every speaker. Harder, and a lot more faces.
- Don't know either face? Say so. The quote is swapped for another and the swap is recorded. Nobody is removed from the corpus. Each mode allows a limited number of swaps before the game ends.
- 15, 30 or 50 quotes per game, your choice. Your score is compared only with other completed games of the same mode and length.
- Rate a quote. After the last quote (and the halfway one in longer games) you're asked to rate a fresh, unattributed quote on five word-pair scales. It takes a few seconds and it's the data that keeps the quiz free of ads.
How the game stays fair
- The correct answer is never sent to your browser. Your picks are scored on the server.
- The other option is chosen at the same rate the real speaker would appear, so no name is a giveaway.
- Speakers are spread evenly across a game; nobody keeps reappearing while others haven't been seen.
- Answers are never revealed, so nobody can memorise the corpus.
- Games that end early are kept out of the public statistics, so the curves only compare full games.
Questions
- Are these real quotes?
- Yes. Every quote in this experiment had its speaker verified by listening to the recording, was checked to make sure the subtitles matched the spoken words, and had context added in square brackets when a line needed it. Nothing was paraphrased.
- What's the difference between the Standard and Ultra quizzes?
- The Standard Quiz draws only on the regulars: hosts and guests who turn up often enough that you could reasonably know their voice. The Ultra Quiz opens the whole corpus, including people heard only a few times. In both, you choose 15, 30 or 50 quotes per game, and both options on every card come from the same pool, so the mode never tips you off.
- How are the two names chosen?
- One is the real speaker. The other is drawn from the same pool with probability proportional to how often that person speaks in the corpus, which is the same rate at which they show up as the real answer. That means a famous face is no more likely to be right than an obscure one, so the only useful signal is the words themselves. Speakers are spread across a game so nobody keeps reappearing while others haven't been seen.
- Why don't you show the answers?
- Two reasons. Quotes would be memorised and the comparison between players would stop meaning anything. And the game is not trying to teach you who said what. It is trying to measure how distinctive people's words are. The aggregate results are published on the stats page.
- What happens when I say I don't know either?
- The quote is swapped for a freshly dealt one and the swap is recorded as its own data point: which quote, which two faces, how long you looked. Neither face is dropped from the corpus. Each mode allows a limited number of swaps; use them all and the game ends there.
- Why am I asked to rate a quote?
- Once or twice a game you're shown a fresh quote, never one you just guessed on and without saying who said it, and asked to place it on five scales such as calm to agitated or populist to elitist, drawn at random from a longer list. The ratings are what the whole project is for: over thousands of games they add up to a picture of how each voice comes across from the words alone. Move only the scales you can judge. An untouched slider is recorded as no opinion, not as a middle score.
- Can I stop a game early?
- Yes. There is an End game button under the cards. A game that ends early is stored for research but kept out of the public score curves and stats, and you only see a score if you answered at least 15 quotes.