Document Extraction Engine

Production document intelligence system built at Fusemachines for financial workflows (real-estate lending), processing thousands of documents daily.

  • Document classification, deep homography-based alignment, custom handwriting recognition model, and spelling correction to automate data entry
  • Researched and evaluated DiT, Table Transformers, vision-language models (PaliGemma, Qwen-VL), and LLM-based methods for table and document structure recognition and retrieval
  • RAG pipelines and AI agents with active learning for document validation, discrepancy detection, and deal-specific Q&A
  • Scalable parallel processing pipelines deployed to production

Stack: PyTorch, Hugging Face Transformers, ONNX, OpenCV, vector/graph databases, Docker, AWS