Research
My research spans multimodal LLM systems, document intelligence, and applied computer vision. Current interests: Multimodal LLMs and AI Agents, Trustworthy, Privacy-Preserving, and Secure AI Systems, and Optimization and resource efficient ML.
Audio-LLM for Realtime Interview Analysis
- Developing a multimodal interview system using Qwen2-Audio for structured technical interviewing and realtime analysis of interview conversations.
- Improving conversational alignment with preference optimization (LoRA fine-tuning and Direct Preference Optimization) on synthetic speech, conversation, and preference datasets generated with voice cloning for diverse accents and speaking styles.
- Building realtime audio processing pipelines with voice activity detection (VAD) and speaker diarization to segment live streaming interviews for downstream evaluation and analysis.
Handwriting Recognition Model Improvement (Multilingual OCR: English + Nepali)
- Improving a custom handwriting recognition model for multilingual (English + Nepali) document text extraction.
- Leveraging synthetic handwriting data generation to expand training coverage across scripts and writing styles.
- Integrating an n-gram based language model for spelling correction during decoding.
Paddy Disease Detection and Classification (Undergraduate Research)
- Conducted a comparative study of object detection and classification for paddy disease identification using deep learning models such as CNNs, Vision Transformers, and YOLO.
- Developed a mobile application for real-time disease diagnosis and treatment suggestions.
- Published as a preprint: arXiv:2412.05996.
