Audio-LLM Research: Multimodal Interview Agent
Multimodal interview system with Qwen2-Audio, LoRA + DPO alignment, synthetic speech data, and real-time audio pipelines.
Multimodal interview system with Qwen2-Audio, LoRA + DPO alignment, synthetic speech data, and real-time audio pipelines.
Seq2Seq conversational models in PyTorch: GRU/LSTM with attention and Transformers, incl. counseling dialogue.
End-to-end AI pipelines for KYC, financial, and rental document understanding: classification, alignment, handwriting recognition, RAG agents.
Pix2Pix GAN (PatchGAN) restoration and inpainting of old photos, trained on FFHQ-Small.
Computer vision for agriculture and award-winning hackathon builds: litter detection, news portal, cultural AI.
Medical image segmentation: UNet and UNet++ implemented from scratch with Dice loss and IoU evaluation.
CLI tool to expand YOLO detection datasets with box-aware augmentation: more data, better detectors.