SpeechCritic
View code

Diagnostic speech evaluation

Listen to one reference and two candidates, then compare human judgements with audio-grounded model rationales.

Example 1

English → Japanese

Emotion · Timing
ReferenceSource-language speech
Candidate ATarget-language speech
Candidate BTarget-language speech
English transcriptWhat are we gonna do? Hold on to those training bands and fling ourselves onto the padded mat?
Japanese transcriptどう使うの?このトレーニングバンドにつかまって、僕らをマットまで飛ばすの?
Human judgementsOverall A · 3–0Speaker TieEmotion ATiming APronunciation TieArtifacts Tie

Our diagnostic speech judge

✓ All verdicts match human judgements
Reference Anchor
  • Young male voice with a light vocal weight.
  • Highly energetic, frantic, and urgent intent.
  • Fast pacing with a brief pause after “What are we gonna do?”.
  • The phrase “fling ourselves onto the padded mat” is delivered with rapid, urgent momentum.
Speaker Tie

Candidate A: Sounds like a young male, closely matching the light vocal weight and energetic speaker profile of the reference.

Candidate B: Sounds like a young male, but the vocal weight is slightly heavier and the vocal energy is noticeably lower than the reference.

Tie note: Both candidates provide a plausible young male speaker profile, though Candidate A is slightly closer in energy.

Emotion A

Candidate A: Captures the frantic urgency and high energy of the reference, particularly in the urgent delivery of “飛ばす” (tobasu).

Candidate B: Sounds noticeably calmer and more subdued, missing the frantic desperation present in the reference.

Timing A

Candidate A: Matches the rapid, urgent pacing of the reference well, maintaining a natural flow through the entire sentence.

Candidate B: Delivers the line with a slower, more deliberate pace that feels too relaxed relative to the reference.

Pronunciation Tie

Candidate A: Pronunciation is clear and natural throughout the phrase.

Candidate B: Pronunciation is clear and natural throughout the phrase.

Tie note: Both candidates articulate the Japanese text clearly without any noticeable accent leakage.

Audio Artifacts Tie

Candidate A: No clear technical defect.

Candidate B: No clear technical defect.

Tie note: Both are free of obvious technical artifacts or glitches.

Overall A

Candidate A is the clear winner because it successfully captures the frantic urgency and rapid pacing of the reference. Candidate B sounds too calm and deliberate, missing the reference's high energy. Both candidates share a suitable vocal profile, but Candidate A's superior emotional expressiveness and timing make it the much stronger performance.

Final answer: A

Example 2

English → Spanish

Pronunciation · Delivery
ReferenceSource-language speech
Candidate ATarget-language speech
Candidate BTarget-language speech
English transcriptYou invited my friend. Why? Was it just so we could finish the mural before the festival begins?
Spanish transcriptInvitaste a mi amigo. ¿Por qué? ¿Solo para que termináramos el mural antes de que empezara el festival?
Human judgementsOverall B · 3–0Speaker TieEmotion BTiming TiePronunciation BArtifacts Tie

Our diagnostic speech judge

✓ Overall and unanimous verdicts match human judgements
Reference Anchor
  • Adult male with a deep, slightly raspy vocal weight.
  • Intense, confrontational, and accusatory tone.
  • Deliberate pacing with clear pauses after “You invited my friend” and “Why?”.
  • The final question “before the festival begins?” is delivered with a sense of finality.
Speaker Model B · Human Tie

Candidate A: Sounds like a younger male with a lighter vocal weight, missing the mature, raspy texture of the reference.

Candidate B: Closely matches the deep, slightly raspy vocal weight and mature age profile of the reference speaker.

Emotion B

Candidate A: Delivers the lines with a flat, robotic tone that completely misses the intense, confrontational anger of the reference.

Candidate B: Successfully captures the accusatory and intense emotion, particularly in the delivery of “¿Por qué?”.

Timing Model B · Human Tie

Candidate A: Rushes through the delivery, failing to include the necessary pauses after “amigo” and “¿Por qué?”, resulting in a hurried and unnatural rhythm.

Candidate B: Maintains a natural, deliberate pace with appropriate pauses after “amigo” and “¿Por qué?”, aligning well with the reference's rhythm.

Pronunciation B

Candidate A: Exhibits a very heavy English accent, with unnatural vowel sounds and awkward articulation on words like “amigo” and “el mural”.

Candidate B: Features clear, authentic Latin American Spanish pronunciation with natural phoneme realization throughout the entire line.

Audio Artifacts Tie

Candidate A: No clear technical defect.

Candidate B: No clear technical defect.

Tie note: Both candidates are free of obvious technical audio defects, background noise, or synthesis glitches.

Overall B

Candidate B is the clear winner across all performance dimensions. Candidate A suffers from a heavy English accent, rushed pacing, and a flat, robotic delivery that fails to capture the intense emotion of the reference. Candidate B provides a natural, authentic, and emotionally accurate performance that perfectly matches the reference speaker's vocal weight and confrontational tone.

Final answer: B

Example 3

SFT vs. SFT + RL rationales

English → Japanese
ReferenceSource-language speech
Candidate ATarget-language speech
Candidate BTarget-language speech
English transcriptWe don't. By now, they're already outside the park.
Japanese transcriptムリだな。今ごろはもう公園の外だ。
Human judgementsOverall B · 3–0Speaker TieEmotion BTiming BPronunciation BArtifacts B

Matched-verdict rationale comparison

✓ SFT + RL preferred by human listeners

SFT

Reference Anchor
  • Adult male with a low, slightly raspy vocal weight.
  • Weary, resigned, and slightly frustrated tone.
  • Deliberate pause between the two short sentences, with a heavy breath or sigh before “By now, they’re already outside the park.”
  • “By now, they’re already outside the park.” / “今ごろはもう公園の外だ”
Speaker B

Candidate A: Sounds like an adult male, but the vocal weight is slightly lighter and the vocal presence feels less grounded.

Candidate B: Closely matches the reference's low, slightly raspy vocal weight and captures its weary vocal quality.

Emotion B

Candidate A: The emotional delivery is disrupted by awkward hesitations, losing the steady, resigned frustration of the reference.

Candidate B: Successfully captures the weary resignation of the reference, particularly in the delivery of “今ごろはもう公園の外だ”.

Timing B

Candidate A: The pacing is severely disrupted by unnatural pauses and inserted sounds, completely breaking the natural flow of the utterance.

Candidate B: Maintains a natural speaking rate with appropriate, well-placed pauses between the phrases, matching the deliberate rhythm of the reference.

Pronunciation B

Candidate A: The phrasing is disjointed, and the intonation on “ムリだな” sounds unnatural and awkward.

Candidate B: Delivers clear, authentic Japanese pronunciation with natural phrase-level intonation throughout the line.

Audio Artifacts B

Candidate A: Contains severe synthesis artifacts, including hallucinated speech, inserted humming, and a loud, unnatural laugh at the end of the clip.

Candidate B: No clear technical defect.

Overall B

Candidate B is vastly superior across all dimensions, providing a clean, well-paced, and emotionally accurate performance that matches the reference's weary tone. Candidate A suffers from severe synthesis artifacts, including hallucinated speech and inserted noises, which completely ruin the timing, emotion, and technical quality.

Final answer: B

SFT + RL

Reference Anchor
  • Adult male with a low, slightly raspy vocal weight.
  • Weary, resigned, and slightly frustrated tone.
  • Deliberate pause between the two short sentences, with a heavy breath or sigh before “By now, they’re already outside the park.”
  • The final phrase is delivered with a definitive, downward inflection.
Speaker B

Candidate A: Sounds like an adult male, but the vocal weight is heavily obscured by unnatural vocalizations, making the speaker profile feel less grounded.

Candidate B: Closely matches the reference's low, slightly raspy vocal weight and mature vocal presence.

Emotion B

Candidate A: The emotional intent is completely derailed by bizarre, non-speech noises, losing the weary resignation of the reference.

Candidate B: Successfully captures the weary, resigned frustration of the reference, particularly in the delivery of “ムリだな” and the final phrase “今ごろはもう公園の外だ”.

Timing B

Candidate A: The pacing is ruined by unnatural, drawn-out vocalizations inserted between the phrases, completely breaking the natural rhythm.

Candidate B: Maintains a natural speaking rate with appropriate, well-placed pauses between the sentences, aligning well with the reference's deliberate pacing.

Pronunciation B

Candidate A: The articulation is severely compromised by the bizarre noises, and the phrasing is disrupted, making the Japanese sound unnatural.

Candidate B: Delivers clear, authentic Japanese pronunciation with natural phrase-level intonation throughout.

Audio Artifacts B

Candidate A: Suffers from severe synthesis artifacts, including loud, glitchy groans, unnatural syllables, and robotic distortion throughout the entire clip.

Candidate B: No clear technical defect.

Overall B

Candidate B is vastly superior across all dimensions, providing a clean, well-paced, and emotionally accurate performance that matches the reference's weary tone. Candidate A is unusable due to severe synthesis artifacts, bizarre non-speech noises, and disrupted pacing.

Final answer: B

Example 4

A persuasive rationale can still be wrong

English → Japanese
ReferenceSource-language speech
Candidate ATarget-language speech
Candidate BTarget-language speech
English transcriptLook at this mess. The equipment still isn't ready, and the inspection begins at noon.
Japanese transcriptこの有様を見ろ。機材の準備がまだできていない。正午には検査が始まるぞ。
Human judgementsOverall B · 3–0Speaker BEmotion TieTiming TiePronunciation TieArtifacts Tie

OPD with Qwen3-Omni-30B teacher

✓ Preferred by human listenersOverall A · disagrees with human judgement
Reference Anchor
  • Adult male with a deep, resonant, and authoritative voice; perceived as mature and serious.
  • Emotionally charged, delivering a strong, urgent warning with a sense of gravity and frustration.
  • Pacing is deliberate, with clear pauses between phrases: a strong emphasis on “The equipment still isn't ready” and a final, rhetorical “the inspection begins at noon.”
  • “Look at this mess” is delivered with finality, followed by a sharper inflection on “isn't ready” and a downward tone at the end.
Speaker A

Candidate A: The voice is deeper and more resonant than Candidate B, with a mature, authoritative tone that better matches the reference speaker’s perceived age and gravitas.

Candidate B: The voice is noticeably higher-pitched and thinner, sounding younger and less weighty than the reference.

Emotion Tie

Candidate A: Delivers the line with a sense of urgency and conviction, but the emotional arc is somewhat flat. “機材の準備がまだできていない” is delivered with a neutral tone, lacking the rhetorical emphasis of the reference.

Candidate B: The delivery is more emotionally charged, with a rising inflection on the same phrase and a sense of frustration. However, the overall tone is less restrained and more exaggerated than the reference.

Timing Tie

Candidate A: The pacing is rushed and unnatural, with no clear pauses between phrases. The delivery feels clipped and lacks the deliberate rhythm of the reference.

Candidate B: The pacing is slightly slower and more deliberate, with a natural pause after “機材の準備がまだできていない”. However, the delivery still feels rushed and lacks the reference’s deliberate rhythm.

Pronunciation A

Candidate A: The pronunciation is clear and natural, with correct Japanese grammar and intonation.

Candidate B: The pronunciation is heavily affected by non-native accent leakage, particularly in “有様”, which sounds like “ゆうさま” instead of “ありさま”. The final phrase also carries a non-native accent.

Audio Artifacts Tie

Candidate A: No clear technical defect. The audio is clean and free of artifacts.

Candidate B: No clear technical defect. The audio is clean and free of artifacts.

Tie note: Both candidates are free of obvious technical artifacts or glitches.

Overall A

Candidate A is better because it more accurately matches the reference speaker’s vocal profile, and its pronunciation is more authentic. While Candidate B’s delivery is more emotionally charged, it suffers from a non-native accent and less natural pacing. The reference’s authoritative tone is better conveyed by Candidate A’s deeper voice and more deliberate delivery.

Final answer: A