A fully local, cloud-free multimodal retrieval-augmented generation system that indexes and semantically searches across text, images, and audio — powered by FAISS, CLIP, and Whisper.
Semantic search over your private documents, photos, and voice memos. Nothing ever leaves your machine.
- Three modalities, one query — search text, images, and audio with a single natural-language query (e.g. "a dog running on a beach").
- Cross-modal retrieval — image and audio live in the same CLIP embedding space, so you can search images using audio transcripts and vice versa.
- Drag-and-drop dashboard — a Dash web UI for uploading files, viewing index status, and running searches.
- Bulk ingestion script — batch-index text, images, and audio with progress bars.
- Real evaluation harness — Recall@K and MRR per modality, using each indexed item's own caption/transcript as the ground-truth query.
- 100% offline — no cloud API calls, no telemetry, no data leaving the box. All models run locally.
- Service-oriented — split into ingestion, retrieval, embedding, and dashboard services for clean separation of concerns (and optional GPU isolation).
┌─────────────┐ HTTP ┌──────────────────┐
│ Dash UI │ ─────────► │ Ingestion Service│
│ (port 8050)│ │ (text/image/ │
└──────┬──────┘ │ audio) │
│ └────────┬─────────┘
│ HTTP │ embed
▼ ▼
┌─────────────┐ ┌──────────────────┐
│ Retrieval │ ◄───────── │ Embedding Service│
│ Service │ │ (CLIP, Whisper, │
│ (port 8002)│ │ sentence-trans) │
└──────┬──────┘ └──────────────────┘
│
▼
┌─────────────┐
│ FAISS idx │
│ text/image/│
│ audio │
└─────────────┘
| Component | Role |
|---|---|
Dashboard (app.py) |
Dash UI for uploads, search, and index stats. |
Retrieval Service (main.py) |
FastAPI service exposing /search, /upsert, /stats, /export. |
FAISS Store (faiss_store.py) |
Per-modality IndexFlatIP indices + parallel metadata lists. |
Chunker (chunking.py) |
Lightweight word-based chunker with overlap. |
| Scripts | build_index.py, download_sample_data.py, evaluate.py. |
- FAISS — vector similarity search (
IndexFlatIP, optionalIndexIVFFlatupgrade) - CLIP — shared embedding space for images and audio-caption text
- Whisper — local audio transcription
- sentence-transformers — text embeddings
- FastAPI + Dash + httpx
- tqdm for progress bars
python -m venv .venv
\.venv\Scripts\Activate.ps1
pip install -r requirements.txtOn macOS or Linux, activate the environment with source .venv/bin/activate.
.\run_local.ps1This starts the embedding, ingestion, retrieval, and dashboard services in separate windows. Open http://localhost:8050 when they are ready. The first request can take longer while the local models load.
If PowerShell blocks the script, use:
powershell -ExecutionPolicy Bypass -File .\run_local.ps1To start services manually, run these commands in four terminals after activating the environment:
python -m services.embedding_service.main
python -m services.ingestion_service.main
python -m services.retrieval_service.main
python app.pypython scripts/download_sample_data.py --text-n 200 --audio-n 20This fetches:
- Text — arXiv abstracts (no auth)
- Audio — LibriSpeech dev-clean sample (~330 MB) with ground-truth transcripts
- Images — prints Kaggle instructions (Flickr8k requires a free Kaggle account)
Use the dashboard to upload one file at a time, or place files into these folders for bulk ingestion:
data/raw/text/ # .txt, .pdf
data/raw/images/ # .jpg, .jpeg, .png, .webp
data/raw/audio/ # .wav, .mp3, .flac, .m4a
.\.venv\Scripts\python.exe .\scripts\build_index.pyThe script scans each folder recursively and sends supported files to the ingestion service. Use the dashboard's Refresh button to see the updated vector counts.
Re-running bulk ingestion adds new vectors; it does not deduplicate files. Start with an empty
data/indexes/directory when you need a clean rebuild.
Visit http://localhost:8050 to upload content, check index counts, and search across all modalities. Image results are displayed directly in the dashboard.
.\.venv\Scripts\python.exe .\scripts\evaluate.py --top-k 1 5 10 --sample 500Evaluate a single modality, for example images:
.\.venv\Scripts\python.exe .\scripts\evaluate.py --modalities image --top-k 1 5 10For meaningful image evaluation, provide a descriptive caption when uploading an image in the dashboard. Bulk-ingested images without captions use their filename as the evaluation query.
Methodology (the headline accuracy number you can talk about in interviews):
- For each indexed item, we already have a natural-language ground-truth query — an image's caption, an audio clip's Whisper transcript, or a held-out sentence from a text chunk.
- We embed that query and search the corresponding index.
- "Correct" = the source item appears in the top-K results (Recall@K).
- We report Recall@1, Recall@5, Recall@10 per modality and overall, plus MRR (mean reciprocal rank).
This is the standard evaluation protocol for cross-modal retrieval systems (same idea as Recall@K on Flickr30k for CLIP-style models), applied uniformly across all three modalities.
PrivateRAG/
├── app.py # Dash dashboard
├── main.py # Retrieval service (FastAPI)
├── faiss_store.py # FAISS wrapper (per-modality indices)
├── chunking.py # Word-based text chunker w/ overlap
├── build_index.py # Batch indexing script
├── download_sample_data.py# Sample data fetcher
├── evaluate.py # Recall@K / MRR evaluation
├── common/ # Shared config & schemas
│ ├── config.py
│ └── schemas.py
├── services/
│ ├── embedding_service/ # CLIP, Whisper, sentence-transformers
│ ├── ingestion_service/ # /ingest/{text,image,audio}
│ └── retrieval_service/ # this repo's main.py
├── data/
│ ├── raw/{text,images,audio}/
│ └── indexes/
└── scripts/
- One index per modality, one shared space for image+audio. Text uses sentence-transformers; image and audio use CLIP's text tower so they're directly comparable in vector space.
- Inner product == cosine similarity because all vectors are pre-normalized upstream.
IndexFlatIPis the default since the target corpus is ~2-3k vectors per modality. Swap inIndexIVFFlat/IndexHNSWFlatfor larger corpora — a one-line change inModalityIndex.__init__.- Thread-safe writes — each index has its own lock so concurrent uploads don't corrupt the index.
- Dashboard stays GPU-free — it only talks to services over HTTP, so it can run on a tiny container while heavy models live elsewhere.
Copy .env.example to .env and edit its values (or set environment variables) to point at:
INGESTION_SERVICE_URL— defaulthttp://localhost:8001RETRIEVAL_SERVICE_URL— defaulthttp://localhost:8002EMBEDDING_SERVICE_URL— defaulthttp://localhost:8003DATA_DIR,RAW_DATA_DIR,INDEX_DIRTEXT_EMBED_DIM,CLIP_EMBED_DIM
| Modality | Extensions |
|---|---|
| Text | .txt, .pdf |
| Image | .jpg, .jpeg, .png, .webp |
| Audio | .wav, .mp3, .flac, .m4a |
PRs welcome. Keep the offline-first guarantee — no new cloud dependencies.
MIT