Skip to content

Repository files navigation

Offline MultiModal RAG

A fully local, cloud-free multimodal retrieval-augmented generation system that indexes and semantically searches across text, images, and audio — powered by FAISS, CLIP, and Whisper.

Semantic search over your private documents, photos, and voice memos. Nothing ever leaves your machine.


Features

  • Three modalities, one query — search text, images, and audio with a single natural-language query (e.g. "a dog running on a beach").
  • Cross-modal retrieval — image and audio live in the same CLIP embedding space, so you can search images using audio transcripts and vice versa.
  • Drag-and-drop dashboard — a Dash web UI for uploading files, viewing index status, and running searches.
  • Bulk ingestion script — batch-index text, images, and audio with progress bars.
  • Real evaluation harness — Recall@K and MRR per modality, using each indexed item's own caption/transcript as the ground-truth query.
  • 100% offline — no cloud API calls, no telemetry, no data leaving the box. All models run locally.
  • Service-oriented — split into ingestion, retrieval, embedding, and dashboard services for clean separation of concerns (and optional GPU isolation).

Architecture

┌─────────────┐    HTTP    ┌──────────────────┐
│   Dash UI   │ ─────────► │ Ingestion Service│
│  (port 8050)│            │   (text/image/   │
└──────┬──────┘            │     audio)       │
       │                   └────────┬─────────┘
       │ HTTP                       │ embed
       ▼                            ▼
┌─────────────┐            ┌──────────────────┐
│  Retrieval  │ ◄───────── │ Embedding Service│
│   Service   │            │  (CLIP, Whisper, │
│  (port 8002)│            │  sentence-trans) │
└──────┬──────┘            └──────────────────┘
       │
       ▼
┌─────────────┐
│  FAISS idx  │
│  text/image/│
│   audio     │
└─────────────┘
Component Role
Dashboard (app.py) Dash UI for uploads, search, and index stats.
Retrieval Service (main.py) FastAPI service exposing /search, /upsert, /stats, /export.
FAISS Store (faiss_store.py) Per-modality IndexFlatIP indices + parallel metadata lists.
Chunker (chunking.py) Lightweight word-based chunker with overlap.
Scripts build_index.py, download_sample_data.py, evaluate.py.

Tech Stack


Quick Start

1. Install dependencies

python -m venv .venv
\.venv\Scripts\Activate.ps1
pip install -r requirements.txt

On macOS or Linux, activate the environment with source .venv/bin/activate.

2. Start the services (Windows)

.\run_local.ps1

This starts the embedding, ingestion, retrieval, and dashboard services in separate windows. Open http://localhost:8050 when they are ready. The first request can take longer while the local models load.

If PowerShell blocks the script, use:

powershell -ExecutionPolicy Bypass -File .\run_local.ps1

To start services manually, run these commands in four terminals after activating the environment:

python -m services.embedding_service.main
python -m services.ingestion_service.main
python -m services.retrieval_service.main
python app.py

3. Pull sample data

python scripts/download_sample_data.py --text-n 200 --audio-n 20

This fetches:

  • Text — arXiv abstracts (no auth)
  • Audio — LibriSpeech dev-clean sample (~330 MB) with ground-truth transcripts
  • Images — prints Kaggle instructions (Flickr8k requires a free Kaggle account)

4. Add and index your files

Use the dashboard to upload one file at a time, or place files into these folders for bulk ingestion:

data/raw/text/    # .txt, .pdf
data/raw/images/  # .jpg, .jpeg, .png, .webp
data/raw/audio/   # .wav, .mp3, .flac, .m4a
.\.venv\Scripts\python.exe .\scripts\build_index.py

The script scans each folder recursively and sends supported files to the ingestion service. Use the dashboard's Refresh button to see the updated vector counts.

Re-running bulk ingestion adds new vectors; it does not deduplicate files. Start with an empty data/indexes/ directory when you need a clean rebuild.

5. Open the dashboard

Visit http://localhost:8050 to upload content, check index counts, and search across all modalities. Image results are displayed directly in the dashboard.


Evaluation

.\.venv\Scripts\python.exe .\scripts\evaluate.py --top-k 1 5 10 --sample 500

Evaluate a single modality, for example images:

.\.venv\Scripts\python.exe .\scripts\evaluate.py --modalities image --top-k 1 5 10

For meaningful image evaluation, provide a descriptive caption when uploading an image in the dashboard. Bulk-ingested images without captions use their filename as the evaluation query.

Methodology (the headline accuracy number you can talk about in interviews):

  • For each indexed item, we already have a natural-language ground-truth query — an image's caption, an audio clip's Whisper transcript, or a held-out sentence from a text chunk.
  • We embed that query and search the corresponding index.
  • "Correct" = the source item appears in the top-K results (Recall@K).
  • We report Recall@1, Recall@5, Recall@10 per modality and overall, plus MRR (mean reciprocal rank).

This is the standard evaluation protocol for cross-modal retrieval systems (same idea as Recall@K on Flickr30k for CLIP-style models), applied uniformly across all three modalities.


Project Layout

PrivateRAG/
├── app.py                 # Dash dashboard
├── main.py                # Retrieval service (FastAPI)
├── faiss_store.py         # FAISS wrapper (per-modality indices)
├── chunking.py            # Word-based text chunker w/ overlap
├── build_index.py         # Batch indexing script
├── download_sample_data.py# Sample data fetcher
├── evaluate.py            # Recall@K / MRR evaluation
├── common/                # Shared config & schemas
│   ├── config.py
│   └── schemas.py
├── services/
│   ├── embedding_service/ # CLIP, Whisper, sentence-transformers
│   ├── ingestion_service/ # /ingest/{text,image,audio}
│   └── retrieval_service/ # this repo's main.py
├── data/
│   ├── raw/{text,images,audio}/
│   └── indexes/
└── scripts/

Design Notes

  • One index per modality, one shared space for image+audio. Text uses sentence-transformers; image and audio use CLIP's text tower so they're directly comparable in vector space.
  • Inner product == cosine similarity because all vectors are pre-normalized upstream.
  • IndexFlatIP is the default since the target corpus is ~2-3k vectors per modality. Swap in IndexIVFFlat / IndexHNSWFlat for larger corpora — a one-line change in ModalityIndex.__init__.
  • Thread-safe writes — each index has its own lock so concurrent uploads don't corrupt the index.
  • Dashboard stays GPU-free — it only talks to services over HTTP, so it can run on a tiny container while heavy models live elsewhere.

Configuration

Copy .env.example to .env and edit its values (or set environment variables) to point at:

  • INGESTION_SERVICE_URL — default http://localhost:8001
  • RETRIEVAL_SERVICE_URL — default http://localhost:8002
  • EMBEDDING_SERVICE_URL — default http://localhost:8003
  • DATA_DIR, RAW_DATA_DIR, INDEX_DIR
  • TEXT_EMBED_DIM, CLIP_EMBED_DIM

Supported File Types

Modality Extensions
Text .txt, .pdf
Image .jpg, .jpeg, .png, .webp
Audio .wav, .mp3, .flac, .m4a

Contributing

PRs welcome. Keep the offline-first guarantee — no new cloud dependencies.


License

MIT

About

A fully offline multimodal RAG

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages