Audio-LLM Research: Multimodal Interview Agent
Research on a multi-modal LLM-powered interview agent to assist hiring managers throughout the recruitment process.
- Fine-tuned Qwen2-Audio with LoRA and Direct Preference Optimization (DPO) for conversational alignment in structured technical interviewing
- Designed and generated high-quality speech/conversation datasets including preference data, using synthetic data generation and voice cloning to simulate diverse accents and speaking styles
- Real-time audio processing pipelines with voice activity detection (VAD) and speaker diarization to segment live streaming interviews for downstream evaluation
Stack: PyTorch, Hugging Face Transformers, Qwen2-Audio, LoRA/PEFT, HPC environments
