HistoAI

Ecosystem for digitising historical books with OCR, knowledge graphs, and RAG

HistoAI turns scanned historical books into a searchable, queryable collection. The pipeline runs OCR with image extraction on degraded scans, extracts structured data, builds a knowledge graph from the text, and exposes the collection through a RAG chatbot that answers context-aware questions. I started it at CAIR Lab, DSVV and continue to develop it at Swami Rama Himalayan University.

Stack: Python, OCR, knowledge graphs, RAG, large language models.