An AI-powered document intelligence system designed to transform unstructured PDF documents into searchable, conversational and actionable knowledge.
Research documents often contain valuable information, but extracting specific insights manually can be slow and difficult to scale. This project addresses that problem by providing a conversational interface for querying uploaded PDF documents.
Professionals and researchers may spend significant time manually searching, reading and cross-referencing lengthy documents to answer specific questions. The objective was to create an AI assistant capable of retrieving relevant information and converting it into concise, contextual answers.
Build a production-oriented AI assistant that combines document retrieval with LLM reasoning, while providing measurable answer quality and source traceability.
Important information is distributed across lengthy PDF documents and is difficult to locate efficiently.
Simple keyword search may miss semantically related information, while vector-only retrieval can miss exact terminology.
Users need answers that understand document context rather than simply returning matching text fragments.
The system uses a Retrieval-Augmented Generation architecture to retrieve relevant document chunks before passing them to the LLM. Hybrid retrieval combines semantic FAISS search with BM25 keyword search, while query expansion improves retrieval coverage. Conversation memory enables contextual interactions, and the LLM converts retrieved evidence into natural-language responses. When the uploaded documents do not directly answer a question, the assistant can use reasoning and general knowledge rather than being restricted to document extraction.
The RAG pipeline was evaluated across five questions using Faithfulness and Answer Relevancy metrics.
The evaluation achieved an average faithfulness score of 0.8013 and average answer relevancy of 0.7912 across the five evaluated questions, demonstrating that the system generally produced answers grounded in the retrieved context while maintaining strong relevance to user questions.
The architecture separates document processing, retrieval, reasoning, evaluation and user interaction into modular components.
βββββββββββββββββββββββ
β PDF Upload β
ββββββββββββ¬βββββββββββ
β
βΌ
βββββββββββββββββββββββ
β PDF Text Loader β
ββββββββββββ¬βββββββββββ
β
βΌ
βββββββββββββββββββββββ
β Text Chunking β
ββββββββββββ¬βββββββββββ
β
ββββββββββββββ΄βββββββββββββ
βΌ βΌ
ββββββββββββββββββ ββββββββββββββββββ
β FAISS Vectors β β BM25 Index β
β Semantic Searchβ β Keyword Search β
βββββββββ¬βββββββββ βββββββββ¬βββββββββ
β β
ββββββββββββ¬βββββββββββββββ
βΌ
βββββββββββββββββββββββ
β Hybrid Retrieval β
ββββββββββββ¬βββββββββββ
β
βΌ
βββββββββββββββββββββββ
β Query Expansion β
ββββββββββββ¬βββββββββββ
β
βΌ
βββββββββββββββββββββββ
β Context Constructionβ
ββββββββββββ¬βββββββββββ
β
βΌ
βββββββββββββββββββββββ
β Conversation β
β Memory β
ββββββββββββ¬βββββββββββ
β
βΌ
βββββββββββββββββββββββ
β Llama 3 β
β LLM Generation β
ββββββββββββ¬βββββββββββ
β
βΌ
βββββββββββββββββββββββ
β Answer + Citations β
βββββββββββββββββββββββ
Users upload one or more PDF documents.
Documents are loaded, cleaned and divided into chunks.
Chunks are indexed using vector and keyword retrieval.
Hybrid retrieval finds relevant document information.
Query expansion improves retrieval coverage.
The LLM generates a contextual response.
Relevant document and page information is surfaced.
Answer quality is measured through RAG evaluation.
Combines FAISS semantic retrieval with BM25 keyword search.
Expands user queries to improve retrieval coverage.
Maintains previous interactions for contextual conversations.
Provides document and page information for retrieved content.
Can reason beyond explicit document statements when appropriate.
Caches generated responses to avoid unnecessary repeated processing.
Records questions, responses, retrieved chunks and response time.
Captures helpful and not-helpful feedback for evaluation.
The evaluation report measured both faithfulness and answer relevancy for five representative questions.
| Question | Faithfulness | Answer Relevancy |
|---|---|---|
| Main topic of the document | 75.00% | 100.00% |
| Qualifications required | 66.67% | 61.47% |
| Selection procedure | 66.67% | 94.31% |
| Guidelines for applicants | 92.31% | 72.09% |
| Important dates / deadlines | 100.00% | 67.73% |
The system achieved 100% faithfulness for the question concerning important dates and deadlines, while the main-topic question achieved 100% answer relevancy.
A single retrieval strategy can miss either semantic relationships or exact terminology. Hybrid retrieval was implemented to combine both approaches.
Retrieved information needs to be structured before being passed to the LLM. Document names, pages and chunk boundaries were preserved for traceability.
The system was designed to distinguish between document-supported information and reasoning or general knowledge.
The application incorporates API serving, monitoring, feedback logging, evaluation, containerization and deployment-oriented architecture.
The completed system provides a practical interface for extracting and analyzing information from unstructured documents without requiring users to manually search through every page. With an average 80.13% faithfulness and 79.12% answer relevancy across the evaluated questions, the project demonstrates a measurable foundation for reliable AI-assisted document research.
View the source code, architecture and complete implementation on GitHub.
View GitHub Repository