Listen to one reference and two candidates, then compare human judgements with audio-grounded model rationales.
Example 1
English → Japanese
Emotion · Timing
ReferenceSource-language speech
Candidate ATarget-language speech
Candidate BTarget-language speech
English transcriptWhat are we gonna do? Hold on to those training bands and fling ourselves onto the padded mat?
Japanese transcriptどう使うの?このトレーニングバンドにつかまって、僕らをマットまで飛ばすの?
Human judgementsOverall A · 3–0Speaker TieEmotion ATiming APronunciation TieArtifacts Tie
Our diagnostic speech judge
✓ All verdicts match human judgements
Reference Anchor
Young male voice with a light vocal weight.
Highly energetic, frantic, and urgent intent.
Fast pacing with a brief pause after “What are we gonna do?”.
The phrase “fling ourselves onto the padded mat” is delivered with rapid, urgent momentum.
Speaker Tie
Candidate A: Sounds like a young male, closely matching the light vocal weight and energetic speaker profile of the reference.
Candidate B: Sounds like a young male, but the vocal weight is slightly heavier and the vocal energy is noticeably lower than the reference.
Tie note: Both candidates provide a plausible young male speaker profile, though Candidate A is slightly closer in energy.
Emotion A
Candidate A: Captures the frantic urgency and high energy of the reference, particularly in the urgent delivery of “飛ばす” (tobasu).
Candidate B: Sounds noticeably calmer and more subdued, missing the frantic desperation present in the reference.
Timing A
Candidate A: Matches the rapid, urgent pacing of the reference well, maintaining a natural flow through the entire sentence.
Candidate B: Delivers the line with a slower, more deliberate pace that feels too relaxed relative to the reference.
Pronunciation Tie
Candidate A: Pronunciation is clear and natural throughout the phrase.
Candidate B: Pronunciation is clear and natural throughout the phrase.
Tie note: Both candidates articulate the Japanese text clearly without any noticeable accent leakage.
Audio Artifacts Tie
Candidate A: No clear technical defect.
Candidate B: No clear technical defect.
Tie note: Both are free of obvious technical artifacts or glitches.
Overall A
Candidate A is the clear winner because it successfully captures the frantic urgency and rapid pacing of the reference. Candidate B sounds too calm and deliberate, missing the reference's high energy. Both candidates share a suitable vocal profile, but Candidate A's superior emotional expressiveness and timing make it the much stronger performance.
Final answer: A
Example 2
English → Spanish
Pronunciation · Delivery
ReferenceSource-language speech
Candidate ATarget-language speech
Candidate BTarget-language speech
English transcriptYou invited my friend. Why? Was it just so we could finish the mural before the festival begins?
Spanish transcriptInvitaste a mi amigo. ¿Por qué? ¿Solo para que termináramos el mural antes de que empezara el festival?
Human judgementsOverall B · 3–0Speaker TieEmotion BTiming TiePronunciation BArtifacts Tie
Our diagnostic speech judge
✓ Overall and unanimous verdicts match human judgements
Reference Anchor
Adult male with a deep, slightly raspy vocal weight.
Intense, confrontational, and accusatory tone.
Deliberate pacing with clear pauses after “You invited my friend” and “Why?”.
The final question “before the festival begins?” is delivered with a sense of finality.
Speaker Model B · Human Tie
Candidate A: Sounds like a younger male with a lighter vocal weight, missing the mature, raspy texture of the reference.
Candidate B: Closely matches the deep, slightly raspy vocal weight and mature age profile of the reference speaker.
Emotion B
Candidate A: Delivers the lines with a flat, robotic tone that completely misses the intense, confrontational anger of the reference.
Candidate B: Successfully captures the accusatory and intense emotion, particularly in the delivery of “¿Por qué?”.
Timing Model B · Human Tie
Candidate A: Rushes through the delivery, failing to include the necessary pauses after “amigo” and “¿Por qué?”, resulting in a hurried and unnatural rhythm.
Candidate B: Maintains a natural, deliberate pace with appropriate pauses after “amigo” and “¿Por qué?”, aligning well with the reference's rhythm.
Pronunciation B
Candidate A: Exhibits a very heavy English accent, with unnatural vowel sounds and awkward articulation on words like “amigo” and “el mural”.
Candidate B: Features clear, authentic Latin American Spanish pronunciation with natural phoneme realization throughout the entire line.
Audio Artifacts Tie
Candidate A: No clear technical defect.
Candidate B: No clear technical defect.
Tie note: Both candidates are free of obvious technical audio defects, background noise, or synthesis glitches.
Overall B
Candidate B is the clear winner across all performance dimensions. Candidate A suffers from a heavy English accent, rushed pacing, and a flat, robotic delivery that fails to capture the intense emotion of the reference. Candidate B provides a natural, authentic, and emotionally accurate performance that perfectly matches the reference speaker's vocal weight and confrontational tone.
Final answer: B
Example 3
SFT vs. SFT + RL rationales
English → Japanese
ReferenceSource-language speech
Candidate ATarget-language speech
Candidate BTarget-language speech
English transcriptWe don't. By now, they're already outside the park.
Japanese transcriptムリだな。今ごろはもう公園の外だ。
Human judgementsOverall B · 3–0Speaker TieEmotion BTiming BPronunciation BArtifacts B
Matched-verdict rationale comparison
✓ SFT + RL preferred by human listeners
SFT
Reference Anchor
Adult male with a low, slightly raspy vocal weight.
Weary, resigned, and slightly frustrated tone.
Deliberate pause between the two short sentences, with a heavy breath or sigh before “By now, they’re already outside the park.”
“By now, they’re already outside the park.” / “今ごろはもう公園の外だ”
Speaker B
Candidate A: Sounds like an adult male, but the vocal weight is slightly lighter and the vocal presence feels less grounded.
Candidate B: Closely matches the reference's low, slightly raspy vocal weight and captures its weary vocal quality.
Emotion B
Candidate A: The emotional delivery is disrupted by awkward hesitations, losing the steady, resigned frustration of the reference.
Candidate B: Successfully captures the weary resignation of the reference, particularly in the delivery of “今ごろはもう公園の外だ”.
Timing B
Candidate A: The pacing is severely disrupted by unnatural pauses and inserted sounds, completely breaking the natural flow of the utterance.
Candidate B: Maintains a natural speaking rate with appropriate, well-placed pauses between the phrases, matching the deliberate rhythm of the reference.
Pronunciation B
Candidate A: The phrasing is disjointed, and the intonation on “ムリだな” sounds unnatural and awkward.
Candidate B: Delivers clear, authentic Japanese pronunciation with natural phrase-level intonation throughout the line.
Audio Artifacts B
Candidate A: Contains severe synthesis artifacts, including hallucinated speech, inserted humming, and a loud, unnatural laugh at the end of the clip.
Candidate B: No clear technical defect.
Overall B
Candidate B is vastly superior across all dimensions, providing a clean, well-paced, and emotionally accurate performance that matches the reference's weary tone. Candidate A suffers from severe synthesis artifacts, including hallucinated speech and inserted noises, which completely ruin the timing, emotion, and technical quality.
Final answer: B
SFT + RL
Reference Anchor
Adult male with a low, slightly raspy vocal weight.
Weary, resigned, and slightly frustrated tone.
Deliberate pause between the two short sentences, with a heavy breath or sigh before “By now, they’re already outside the park.”
The final phrase is delivered with a definitive, downward inflection.
Speaker B
Candidate A: Sounds like an adult male, but the vocal weight is heavily obscured by unnatural vocalizations, making the speaker profile feel less grounded.
Candidate B: Closely matches the reference's low, slightly raspy vocal weight and mature vocal presence.
Emotion B
Candidate A: The emotional intent is completely derailed by bizarre, non-speech noises, losing the weary resignation of the reference.
Candidate B: Successfully captures the weary, resigned frustration of the reference, particularly in the delivery of “ムリだな” and the final phrase “今ごろはもう公園の外だ”.
Timing B
Candidate A: The pacing is ruined by unnatural, drawn-out vocalizations inserted between the phrases, completely breaking the natural rhythm.
Candidate B: Maintains a natural speaking rate with appropriate, well-placed pauses between the sentences, aligning well with the reference's deliberate pacing.
Pronunciation B
Candidate A: The articulation is severely compromised by the bizarre noises, and the phrasing is disrupted, making the Japanese sound unnatural.
Candidate B: Delivers clear, authentic Japanese pronunciation with natural phrase-level intonation throughout.
Audio Artifacts B
Candidate A: Suffers from severe synthesis artifacts, including loud, glitchy groans, unnatural syllables, and robotic distortion throughout the entire clip.
Candidate B: No clear technical defect.
Overall B
Candidate B is vastly superior across all dimensions, providing a clean, well-paced, and emotionally accurate performance that matches the reference's weary tone. Candidate A is unusable due to severe synthesis artifacts, bizarre non-speech noises, and disrupted pacing.
Final answer: B
Example 4
A persuasive rationale can still be wrong
English → Japanese
ReferenceSource-language speech
Candidate ATarget-language speech
Candidate BTarget-language speech
English transcriptLook at this mess. The equipment still isn't ready, and the inspection begins at noon.
Japanese transcriptこの有様を見ろ。機材の準備がまだできていない。正午には検査が始まるぞ。
Human judgementsOverall B · 3–0Speaker BEmotion TieTiming TiePronunciation TieArtifacts Tie
OPD with Qwen3-Omni-30B teacher
✓ Preferred by human listenersOverall A · disagrees with human judgement
Reference Anchor
Adult male with a deep, resonant, and authoritative voice; perceived as mature and serious.
Emotionally charged, delivering a strong, urgent warning with a sense of gravity and frustration.
Pacing is deliberate, with clear pauses between phrases: a strong emphasis on “The equipment still isn't ready” and a final, rhetorical “the inspection begins at noon.”
“Look at this mess” is delivered with finality, followed by a sharper inflection on “isn't ready” and a downward tone at the end.
Speaker A
Candidate A: The voice is deeper and more resonant than Candidate B, with a mature, authoritative tone that better matches the reference speaker’s perceived age and gravitas.
Candidate B: The voice is noticeably higher-pitched and thinner, sounding younger and less weighty than the reference.
Emotion Tie
Candidate A: Delivers the line with a sense of urgency and conviction, but the emotional arc is somewhat flat. “機材の準備がまだできていない” is delivered with a neutral tone, lacking the rhetorical emphasis of the reference.
Candidate B: The delivery is more emotionally charged, with a rising inflection on the same phrase and a sense of frustration. However, the overall tone is less restrained and more exaggerated than the reference.
Timing Tie
Candidate A: The pacing is rushed and unnatural, with no clear pauses between phrases. The delivery feels clipped and lacks the deliberate rhythm of the reference.
Candidate B: The pacing is slightly slower and more deliberate, with a natural pause after “機材の準備がまだできていない”. However, the delivery still feels rushed and lacks the reference’s deliberate rhythm.
Pronunciation A
Candidate A: The pronunciation is clear and natural, with correct Japanese grammar and intonation.
Candidate B: The pronunciation is heavily affected by non-native accent leakage, particularly in “有様”, which sounds like “ゆうさま” instead of “ありさま”. The final phrase also carries a non-native accent.
Audio Artifacts Tie
Candidate A: No clear technical defect. The audio is clean and free of artifacts.
Candidate B: No clear technical defect. The audio is clean and free of artifacts.
Tie note: Both candidates are free of obvious technical artifacts or glitches.
Overall A
Candidate A is better because it more accurately matches the reference speaker’s vocal profile, and its pronunciation is more authentic. While Candidate B’s delivery is more emotionally charged, it suffers from a non-native accent and less natural pacing. The reference’s authoritative tone is better conveyed by Candidate A’s deeper voice and more deliberate delivery.