Audio-LLM Research: Multimodal Interview Agent
Multimodal interview system with Qwen2-Audio, LoRA + DPO alignment, synthetic speech data, and real-time audio pipelines.
Multimodal interview system with Qwen2-Audio, LoRA + DPO alignment, synthetic speech data, and real-time audio pipelines.
Seq2Seq conversational models in PyTorch: GRU/LSTM with attention and Transformers, incl. counseling dialogue.
End-to-end AI pipelines for KYC, financial, and rental document understanding: classification, alignment, handwriting recognition, RAG agents.
Pix2Pix GAN (PatchGAN) restoration and inpainting of old photos, trained on FFHQ-Small.
Computer vision for agriculture and award-winning hackathon builds: litter detection, news portal, cultural AI.
Medical image segmentation: UNet and UNet++ implemented from scratch with Dice loss and IoU evaluation.
CLI tool to expand YOLO detection datasets with box-aware augmentation: more data, better detectors.
Published in arXiv, 2024
Comparative study of CNNs, Vision Transformers, and YOLO for paddy disease identification, with a mobile app for real-time diagnosis.
Recommended citation: Bimarsha Khanal et al. (2024). "Paddy Disease Detection and Classification Using Computer Vision Techniques." arXiv:2412.05996.
Download Paper
Published:
Session on applied AI for fellowship students and working professionals.
Published:
Session on engineering production AI systems for fellowship students and working professionals.
Mentorship / Workshop, Innovative Computer Engineering Society, 2023
Volunteer mentor for the month-long Call for Enthusiast program (2023, 2024), guiding junior students in Machine Learning.
Fellowship course, Fusemachines AI Fellowship, 2026
Teaching Assistant in a six-month AI fellowship program for students and working professionals, supporting the goal of democratizing AI education.