How MedRAG-X works
MedRAG-X is a complete Retrieval-Augmented Generation pipeline. Here is exactly what happens between your question and the cited answer.
1 · Document processing
Uploads (PDF, TXT, Markdown up to 10 MB) are validated, text is extracted (pypdf for PDFs, with page numbers preserved), cleaned of PDF artifacts, and split into overlapping 1,200-character chunks that keep section headings and page numbers as metadata.
2 · Embedding
Each chunk is embedded with BAAI/bge-small-en-v1.5 — a 33M-parameter open-source model (384 dimensions) that runs on CPU via ONNX. No GPU, no API key, no per-token cost.
3 · Vector search
Chunk vectors are indexed in a FAISS flat inner-product index. Your question is embedded with the same model, and cosine similarity retrieves the top-K most relevant chunks (K configurable, default 4). A small lexical-overlap bonus sharpens exact-phrase matches.
4 · Grounded generation
Retrieved chunks and the question go to an LLM (Groq-hosted openai/gpt-oss-120b on the free tier) behind a strict prompt: answer only from the numbered evidence, never invent citations, and say when the documents don't contain the answer.
5 · Sources & transparency
The response carries the exact retrieved passages with document name, page or chunk number, and similarity scores — plus a Technical Details view showing scores and the raw context sent to the model.
Full engineering details — chunking rationale, embedding model selection, evaluation metrics, and deployment architecture — live in the repository: RAG_DOCUMENTATION.md and RAG_EVALUATION.md.