Document Extraction Engine

Production document intelligence system built at Fusemachines for financial workflows (real-estate lending), processing thousands of documents daily.

  • Document classification and deep homography-based alignment to automate data entry
  • Researched and evaluated DiT, Table Transformers, vision-language models (PaliGemma, Qwen-VL), and LLM-based methods for table and document structure recognition and retrieval
  • RAG pipelines and AI agents with active learning for document validation, discrepancy detection, and deal-specific Q&A
  • Scalable parallel processing pipelines deployed to production

Stack: PyTorch, Hugging Face Transformers, OpenCV, vector/graph databases, FastAPI, Docker, Kubernetes, AWS