📚 PDF Q&A Chatbot
This is a PDF-based Q&A chatbot built using Streamlit, LangChain, and FAISS. It allows users to upload PDF documents, generate vector embeddings using Hugging Face models, and query the chatbot to get answers based on document content. If the provided documents lack sufficient information, the chatbot falls back on general knowledge.
🚀 Live Demo
Click the link below to upload PDFs and chat with the AI:
🛠Installation
Clone the repository and install dependencies:
git clone https://github.com/VanshajR/ChatPDF.git
cd ChatPDF
pip install -r requirements.txt
🔑 Getting API Keys
This project requires API keys for:
- Groq API (for LLM inference)
- Hugging Face API (for embeddings)
How to Get a Groq API Key
- Visit Groq's API Platform.
- Sign in or create an account.
- Navigate to API Keys under account settings.
- Generate a new API key and copy it.
How to Get a Hugging Face API Key
- Go to Hugging Face.
- Sign in or create an account.
- Click on your profile picture and go to Settings → Access Tokens.
- Generate a new token (choose "Read" access) and copy it.
🎯 Usage
Run the Streamlit app locally:
streamlit run app.py
Steps to Use
- Enter API Keys 🔑 – Provide Groq and Hugging Face API keys in the sidebar.
- Choose Models 🧠– Select an embedding model and an LLM model.
- Upload PDFs 📂 – Upload one or more PDF documents.
- Generate Embeddings 🔄 – The app processes PDFs and creates a vector database.
- Ask Questions 💬 – Type queries in the chat interface and receive context-aware answers.
📜 Technical Overview
- PDF Processing: Uses
PyPDFLoaderto extract text. - Text Splitting: Implements
RecursiveCharacterTextSplitterfor better chunking. - Vector Storage: Uses
FAISSfor storing and retrieving document embeddings. - LLM Responses: Uses Groq's Llama 3 or Mixtral models for answering questions.
- Retrieval Chain: Implements LangChain's RAG pipeline to fetch document-related responses.
📜 License
This project is licensed under the MIT License.
Log in or sign up for Devpost to join the conversation.