Document Extraction Engine
Production document intelligence system built at Fusemachines for financial workflows (real-estate lending), processing thousands of documents daily.
- Document classification, deep homography-based alignment, custom handwriting recognition model, and spelling correction to automate data entry
- Researched and evaluated DiT, Table Transformers, vision-language models (PaliGemma, Qwen-VL), and LLM-based methods for table and document structure recognition and retrieval
- RAG pipelines and AI agents with active learning for document validation, discrepancy detection, and deal-specific Q&A
- Scalable parallel processing pipelines deployed to production
Stack: PyTorch, Hugging Face Transformers, ONNX, OpenCV, vector/graph databases, Docker, AWS
