All work
RAG
DocuChat AI
Ask questions about any PDF and get grounded answers from a hybrid search index.
The problem
Teams have answers buried in long PDFs — manuals, policies, reports — and keyword search keeps missing them.
How it works
Documents are chunked and embedded with bge-m3, stored in PostgreSQL with pgvector, and retrieved with hybrid search that combines vector similarity and full-text ranking. LangGraph orchestrates retrieval and generation, and answers come from LLaMA 3.1 on Groq for sub-second responses.
Pipeline
- 01Upload
- 02Chunk & embed
- 03Hybrid retrieval
- 04Generate
- 05Answer
Highlights
- Hybrid search: vector and full-text in the same database
- Embedding model loaded once, with GPU → CPU fallback
- FastAPI backend with a Streamlit chat interface
- Free-tier inference keeps running costs close to zero
Stack
LangGraphbge-m3pgvectorGroq · LLaMA 3.1FastAPIStreamlit
Want something like this for your company?
I build production versions of these systems on your data and your infrastructure.
Related service: Knowledge assistants
Book a call