AI Research Assistant | Dhritikamal Das
GENERATIVE AI β€’ NLP β€’ RAG

AI Research Assistant

An AI-powered document intelligence system designed to transform unstructured PDF documents into searchable, conversational and actionable knowledge.

Turning documents into actionable knowledge

Research documents often contain valuable information, but extracting specific insights manually can be slow and difficult to scale. This project addresses that problem by providing a conversational interface for querying uploaded PDF documents.

Business Problem

Professionals and researchers may spend significant time manually searching, reading and cross-referencing lengthy documents to answer specific questions. The objective was to create an AI assistant capable of retrieving relevant information and converting it into concise, contextual answers.

Project Goal

Build a production-oriented AI assistant that combines document retrieval with LLM reasoning, while providing measurable answer quality and source traceability.

Problems addressed

πŸ“„

Unstructured Documents

Important information is distributed across lengthy PDF documents and is difficult to locate efficiently.

πŸ”Ž

Retrieval Accuracy

Simple keyword search may miss semantically related information, while vector-only retrieval can miss exact terminology.

πŸ€–

Contextual Answers

Users need answers that understand document context rather than simply returning matching text fragments.

Retrieval + reasoning

The system uses a Retrieval-Augmented Generation architecture to retrieve relevant document chunks before passing them to the LLM. Hybrid retrieval combines semantic FAISS search with BM25 keyword search, while query expansion improves retrieval coverage. Conversation memory enables contextual interactions, and the LLM converts retrieved evidence into natural-language responses. When the uploaded documents do not directly answer a question, the assistant can use reasoning and general knowledge rather than being restricted to document extraction.

Evaluation performance

The RAG pipeline was evaluated across five questions using Faithfulness and Answer Relevancy metrics.

80.13% Average Faithfulness
79.12% Average Answer Relevancy
5 Questions Evaluated

What the results indicate

The evaluation achieved an average faithfulness score of 0.8013 and average answer relevancy of 0.7912 across the five evaluated questions, demonstrating that the system generally produced answers grounded in the retrieved context while maintaining strong relevance to user questions.

End-to-end AI pipeline

The architecture separates document processing, retrieval, reasoning, evaluation and user interaction into modular components.

                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚     PDF Upload      β”‚
                         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                    β”‚
                                    β–Ό
                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚   PDF Text Loader   β”‚
                         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                    β”‚
                                    β–Ό
                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚   Text Chunking     β”‚
                         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                    β”‚
                       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                       β–Ό                         β–Ό
              β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
              β”‚ FAISS Vectors  β”‚        β”‚  BM25 Index    β”‚
              β”‚ Semantic Searchβ”‚        β”‚ Keyword Search β”‚
              β””β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜        β””β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                      β”‚                         β”‚
                      β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                 β–Ό
                       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                       β”‚   Hybrid Retrieval  β”‚
                       β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                  β”‚
                                  β–Ό
                       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                       β”‚   Query Expansion   β”‚
                       β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                  β”‚
                                  β–Ό
                       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                       β”‚ Context Constructionβ”‚
                       β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                  β”‚
                                  β–Ό
                       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                       β”‚   Conversation      β”‚
                       β”‚      Memory         β”‚
                       β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                  β”‚
                                  β–Ό
                       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                       β”‚       Llama 3       β”‚
                       β”‚    LLM Generation   β”‚
                       β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                  β”‚
                                  β–Ό
                       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                       β”‚ Answer + Citations  β”‚
                       β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

From document to answer

STEP 01

Upload

Users upload one or more PDF documents.

STEP 02

Process

Documents are loaded, cleaned and divided into chunks.

STEP 03

Index

Chunks are indexed using vector and keyword retrieval.

STEP 04

Retrieve

Hybrid retrieval finds relevant document information.

STEP 05

Expand

Query expansion improves retrieval coverage.

STEP 06

Generate

The LLM generates a contextual response.

STEP 07

Cite

Relevant document and page information is surfaced.

STEP 08

Evaluate

Answer quality is measured through RAG evaluation.

What the system delivers

Hybrid Search

Combines FAISS semantic retrieval with BM25 keyword search.

Query Expansion

Expands user queries to improve retrieval coverage.

Conversational Memory

Maintains previous interactions for contextual conversations.

Source Citations

Provides document and page information for retrieved content.

General AI Reasoning

Can reason beyond explicit document statements when appropriate.

Response Caching

Caches generated responses to avoid unnecessary repeated processing.

Interaction Monitoring

Records questions, responses, retrieved chunks and response time.

User Feedback

Captures helpful and not-helpful feedback for evaluation.

Technology stack

Python LangChain Llama 3 Groq FAISS BM25 Sentence Transformers Streamlit FastAPI Docker Kubernetes GitHub Actions RAGAS

Question-level performance

The evaluation report measured both faithfulness and answer relevancy for five representative questions.

Question Faithfulness Answer Relevancy
Main topic of the document 75.00% 100.00%
Qualifications required 66.67% 61.47%
Selection procedure 66.67% 94.31%
Guidelines for applicants 92.31% 72.09%
Important dates / deadlines 100.00% 67.73%

Strongest result

The system achieved 100% faithfulness for the question concerning important dates and deadlines, while the main-topic question achieved 100% answer relevancy.

Key implementation challenges

Retrieval Quality

A single retrieval strategy can miss either semantic relationships or exact terminology. Hybrid retrieval was implemented to combine both approaches.

Context Management

Retrieved information needs to be structured before being passed to the LLM. Document names, pages and chunk boundaries were preserved for traceability.

LLM Reliability

The system was designed to distinguish between document-supported information and reasoning or general knowledge.

Production Readiness

The application incorporates API serving, monitoring, feedback logging, evaluation, containerization and deployment-oriented architecture.

From manual search to conversational intelligence

The completed system provides a practical interface for extracting and analyzing information from unstructured documents without requiring users to manually search through every page. With an average 80.13% faithfulness and 79.12% answer relevancy across the evaluated questions, the project demonstrates a measurable foundation for reliable AI-assisted document research.

Explore the implementation

View the source code, architecture and complete implementation on GitHub.

View GitHub Repository
← Back to Portfolio