Research

My research spans multimodal LLM systems, document intelligence, and applied computer vision. Current interests: Multimodal LLMs and AI Agents, Trustworthy, Privacy-Preserving, and Secure AI Systems, and Optimization and resource efficient ML.

Audio-LLM for Realtime Interview Analysis

  • Developing a multimodal interview system using Qwen2-Audio for structured technical interviewing and realtime analysis of interview conversations.
  • Improving conversational alignment with preference optimization (LoRA fine-tuning and Direct Preference Optimization) on synthetic speech, conversation, and preference datasets generated with voice cloning for diverse accents and speaking styles.
  • Building realtime audio processing pipelines with voice activity detection (VAD) and speaker diarization to segment live streaming interviews for downstream evaluation and analysis.

Handwriting Recognition Model Improvement (Multilingual OCR: English + Nepali)

  • Improving a custom handwriting recognition model for multilingual (English + Nepali) document text extraction.
  • Leveraging synthetic handwriting data generation to expand training coverage across scripts and writing styles.
  • Integrating an n-gram based language model for spelling correction during decoding.

Paddy Disease Detection and Classification (Undergraduate Research)

  • Conducted a comparative study of object detection and classification for paddy disease identification using deep learning models such as CNNs, Vision Transformers, and YOLO.
  • Developed a mobile application for real-time disease diagnosis and treatment suggestions.
  • Published as a preprint: arXiv:2412.05996.