The BioNeMo Structure Prediction Pipeline is an open-source, GPU-accelerated workflow that goes from protein sequence to predicted 3D structure at scale. It builds multiple sequence alignments (MSAs), predicts structures for single proteins and multi-chain complexes, and scores the confidence of each prediction. Each phase runs as containerized jobs on a Slurm cluster with NVIDIA GPUs.
NVIDIA and its collaborators used this pipeline to generate two datasets that are openly available in the AlphaFold Database:
- Pandemic Preparedness / viral protein complex dataset: predicted structures for the protein complexes of more than 2,800 viruses (NVIDIA blog, EMBL-EBI announcement, Nature article).
- Proteome-scale protein complexes: about 1.8 million high-confidence complexes, selected from about 31 million predictions across 4,777 proteomes (EMBL-EBI announcement, manuscript).
Researchers can run the same workflow on their own protein targets.
Start here: Quickstart · Documentation · Pipeline overview · Architecture · Release status and limitations
| Phase | What it does | Built on |
|---|---|---|
| Preprocessing | Searches sequence databases on the GPU to build an MSA for each target | ColabFold colabfold_search with MMseqs2 GPU search |
| Folding | Predicts 3D structures with per-residue confidence (pLDDT) and predicted aligned error (PAE) | OpenFold2 with AlphaFold2-Multimer weights (optionally OpenFold2 pTM weights for single-chain targets), accelerated by NVIDIA BioNeMo Inference Runtime. The folding phase can also run standard OpenFold2 or ColabFold (AlphaFold2); see release status for what each backend supports |
| Postprocessing | Scores complex interfaces (ipSAE, pDockQ2), checks for steric clashes and exports ModelCIF and BinaryCIF files with Parquet manifests | AFDB Integration Kit |
bsppctl, the pipeline's command-line tool, runs on your workstation (over SSH)
or directly on a Slurm host and submits each phase as containerized Slurm jobs.
Each phase run is pinned to checksum-verified inputs and closed with a receipt
once its job records are verified. Phases hand off results through local
storage or S3-compatible object storage, so they can run on the same cluster or
on different ones. Folding can spread targets across multiple GPUs, balanced by
sequence length.
Phases are run and handed off individually rather than through a single end-to-end command. See release status and limitations for supported configurations.
- A Slurm cluster with Pyxis and Enroot whose nodes have x86-64 CPUs and NVIDIA GPUs with at least 80 GB of memory, such as the A100 80GB (Ampere) or H100 80GB (Hopper)
- Python 3.12, Git and uv on the machine that runs
bsppctl - Docker, to build the container images
- Your input sequences (FASTA), MSA search databases, model weights and postprocessing reference data (UniProt mappings), which you supply
- Fast local storage, such as NVMe, is recommended for the MSA search databases; the UniRef30 search index alone is about 240 GB
- S3-compatible object storage for postprocessing and benchmark validation
From a checkout of this repository:
# Install the bsppctl command-line tool
uv sync --frozen --package bspp-orchestration-control --no-dev
# Confirm the install
uv run --frozen --package bspp-orchestration-control --no-dev bsppctl --help--frozen installs from the shipped lock file without checking that the lock is
current; see installation and lock handling.
This installs only the lightweight control tool; scientific work runs inside the container images below. Rebuilding the public benchmark also needs PyArrow (see the quickstart), and DEVELOPING.md covers the development setup.
The repository includes pdb-temporal-2022-2025-v1, a 1,000-target benchmark of
experimental structures released in the PDB from 2022 through 2025: 750
monomers, 150 dimers, 50 trimers and 50 tetramers. bsppctl prepare-benchmark
rebuilds it from public RCSB downloads with no credentials. The
quickstart covers this step, and the
BioIR benchmark walkthrough continues
through MSA generation and folding.
- Cluster Profile: settings for one cluster, including connection, paths, container images, mounts, account and resource defaults.
- Phase Plan: what one phase should do, and which cluster it runs on.
- Phase RunSpec: the immutable record created from a Plan for a single run attempt.
- Phase Receipt: the sealed record of a successful attempt, issued after its outputs are verified.
The phase lifecycle guide walks through each step: materialize, submit, status, resume, cancel, retry and finalize.
The skills/ directory contains agent skills that help AI coding
assistants configure runs, operate bsppctl and build containers. They use the
same example files as the documentation:
The examples use placeholder paths and identifiers. Replace them with your own values before submitting a run; see the configuration reference.
Each phase has its own container image; folding uses a shared runtime image
plus the image for your chosen backend. Build the images with Docker, push them to your registry (set
BSPP_REGISTRY) and import them on the cluster as SquashFS files for
Pyxis/Enroot. The container guide covers each step.
| Task | Command | Guide |
|---|---|---|
| Run and manage a phase | bsppctl phase … |
Phase lifecycle |
| Publish phase outputs to object storage | bsppctl phase publish-preprocessing / publish-folding |
Phase lifecycle |
| Rebuild the public benchmark | bsppctl prepare-benchmark |
Quickstart |
| Validate benchmark predictions | bsppctl validate-run |
BioIR walkthrough |
| Qualify preprocessing and postprocessing images on a cluster | bsppctl runtime … |
CLI reference |
For problems during a run, see troubleshooting.
This repository is licensed under the Apache License 2.0 — see LICENSE. Third-party component licenses and notices are published in THIRD_PARTY_NOTICES.md.
The benchmark dataset under docs/benchmarks/pdb-temporal-2022-2025-v1 is derived from RCSB Protein Data Bank records and is provided under the CC0 1.0 Universal public-domain dedication; see its DATASET-LICENSE.md.
This project is currently not accepting contributions.