Audio-LLM Research: Multimodal Interview Agent

Research on a multi-modal LLM-powered interview agent to assist hiring managers throughout the recruitment process.

  • Fine-tuned Qwen2-Audio with LoRA and Direct Preference Optimization (DPO) for conversational alignment in structured technical interviewing
  • Designed and generated high-quality speech/conversation datasets including preference data, using synthetic data generation and voice cloning to simulate diverse accents and speaking styles
  • Real-time audio processing pipelines with voice activity detection (VAD) and speaker diarization to segment live streaming interviews for downstream evaluation

Stack: PyTorch, Hugging Face Transformers, Qwen2-Audio, LoRA/PEFT, HPC environments