Mingyue Huo by the waterfront

I am a PhD candidate in Computational Linguistics at the University of Illinois Urbana-Champaign, advised by Prof. Yan Tang. I expect to graduate in December 2026.

My research focuses on speech and multimodal AI. I work on foundation-model evaluation and post-training, drawing on speech representation learning, multi-speaker understanding, and human speech perception.

This fall, I am a research intern at Mitsubishi Electric Research Laboratories (MERL), working on long-form audio understanding. Previously, I interned at Netflix, Tencent AI Lab, Amazon Prime Video, and ByteDance AI Lab.

I am seeking full-time Research Scientist or Applied Scientist roles and am available in mid-December 2026. Please get in touch.

Research Interests

  • Multimodal foundation models; post-training and alignment: supervised fine-tuning, on-policy distillation, and reinforcement learning
  • End-to-end multi-speaker understanding and speech representation learning
  • Human-aligned evaluation, LLM-as-Judge, and reward modeling

Recent News

Highlighted Projects

SpeechCritic

Diagnostic audio LLM judges from limited human preferences

Developed during my Netflix internship, SpeechCritic combines human-calibrated supervision with SFT, OPD, and RL to train and evaluate diagnostic speech judges across perceptual dimensions.

  • LLM-as-Judge
  • Evaluation
  • SFT
  • OPD
  • RL
  • Human preferences
  • Multimodal AI
Figure 1 · SpeechCritic: diagnostic judging, training, and evaluation. View full figure ↗

TagSpeech

End-to-end meeting transcription: who spoke what, and when

TagSpeech unifies multi-speaker ASR, speaker diarization, and timestamp prediction. Dual speech encoders and interleaved time anchors produce structured, speaker-attributed meeting transcripts.

  • End-to-end ASR
  • Diarization
  • Meeting transcription
  • Temporal grounding
  • Audio-language models
Figure 2 · TagSpeech: dual-stream encoding and time-aware transcription. View full figure ↗

Auden-Voice

General-purpose voice representations

A general-purpose voice encoder for speaker identity, paralinguistic understanding, and LLM-QA, with multi-task and contrastive learning experiments and open-source training recipes in Auden.

  • Representation learning
  • Speaker embedding
  • Speaker verification
  • LLM-QA
  • Contrastive learning
  • Multi-task learning
  • Paralinguistics
Figure 1 · Auden-Voice: voice encoder and speech-language evaluation. View full figure ↗

StyleTSE

Text-guided target speech extraction · ICASSP 2025

Developed during my Amazon Prime Video internship, StyleTSE extracts target speech from mixtures using natural-language descriptions and optional audio clues. TextrolMix supplies paired speech mixtures and speaking-style descriptions.

  • Target speech extraction
  • Speech separation
  • Multimodal learning
  • Text-audio fusion
  • Dataset construction
Figure 1 · StyleTSE: speech separation with text and audio clues. View full figure ↗