Skip to content

Latest commit

 

History

58 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
��
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Latest news 🔥

  • 2026-08-25 — Relation-level ROOT queries and parallel decoding. read_root() now selects objects and collections as SQL relations, exposes their immediate fields as columns. Both object and serialized readers now parallelize single-file scans, while multi-column serialized reads share basket decoding work across sibling fields. Exact remote ROOT URIs, including s3:// and davix://, use the common ROOT input layer.
  • 2026-08-13 — ROOT histograms in SQL. TH1, TH2, TH3 and profile objects are available as relational views.
  • 2026-08-11 — 10.7 billion values on one node. A direct query scanned 117 remote ROOT files in 10 min 44 s at 16.6 million values/s with approximately 1.3 GiB peak memory. Methodology and results.

Quick start

ROOT4DuckDB Demo

The current binary release requires:

  • Linux x86-64 with glibc 2.34 or newer;
  • DuckDB 1.4.5;
  • a compatible CERN ROOT 6.40 installation.

Download the extension:

curl -fL \
  https://github.com/LordVitiate/root4duckdb/releases/latest/download/root.duckdb_extension \
  -o root.duckdb_extension

The extension is not signed with an official DuckDB key:

duckdb -unsigned

Load it:

LOAD './root.duckdb_extension';

Inspect a ROOT file:

SELECT *
FROM read_root('events.root');

Read a primitive collection:

SELECT value
FROM read_root(
    'events.root',
    path_prefix := '/energy'
);

Read fields of a nested experiment object using its ROOT dictionary:

SELECT momentum
FROM read_root(
    'events.root',
    dictionary := 'libExperiment.so',
    path_prefix := '/Event/tracks'
);

Query a ROOT histogram:

SELECT x_low, x_high, content, error
FROM read_root(
    'analysis.root',
    path_prefix := '/analysis/mass'
);

Why root4duckdb?

ROOT efficiently stores complex scientific data. But an experiment is usually a dataset distributed across thousands of files, branches and baskets.

ROOT4DuckDB provides one logical SQL layer over that storage:

  • one logical path across different ROOT layouts;
  • automatic selection of the safest efficient reader;
  • selective file, basket and entry reads;
  • reusable statistics, Bloom filters and snapshots;
  • direct integration with DuckDB queries, joins and aggregates;
  • no conversion of the original event data.

ROOT stores the data. ROOT4DuckDB makes it a queryable data platform.

Features

  • primitive and deeply nested ROOT fields;
  • fully split, partially split and unsplit layouts;
  • direct, serialized and universal object readers;
  • automatic correctness fallback;
  • local, remote and parallel multi-file scans;
  • projection and predicate pushdown;
  • basket-aware indexes with statistics and Bloom filters;
  • versioned metadata through Apache Iceberg;
  • relational views for ROOT histograms.

Nested collections become ordinary rows:

/Event/tracks/hits

event_id | tracks_idx | hits_idx | energy

ROOT4DuckDB hides the physical traversal while preserving the logical structure through index columns.

Performance

A measured direct scan decoded:

Metric Result
Remote ROOT files 117
Values 10,696,574,044
Wall time 10 min 44 s
Sustained rate 16.6 million values/s
Peak memory approximately 1.3 GiB

An indexed validation query reduced 2,960 baskets to one basket, five entries and 25 decoded values.

These are measured workloads, not universal performance claims. See Direct multi-file scans and performance.

Release files

For normal use, download:

root.duckdb_extension

The release archive additionally contains licenses, checksums and build information for auditing and redistribution.

Apache Iceberg and the compiler runtime are embedded in the extension. Separate Iceberg or GCC runtime libraries are not required.

ROOT remains external because its I/O runtime and experiment dictionaries must match the environment that opens the data.

Building

./build-iceberg.sh --clean --jobs $(nproc)
./build-root4duckdb.sh --clean --package --jobs $(nproc)

Release artifacts are written to dist/.

The build requires a C++23 compiler, CMake 3.28+, Python, Ninja and CERN ROOT.

See Building from source.

Documentation

Generate the C++ API documentation with:

doxygen Doxyfile

Status

ROOT4DuckDB is a research prototype.

Correctness takes priority over forcing an optimization. When a narrow reader cannot prove that it preserves the result, ROOT4DuckDB falls back to the universal ROOT reader or fails explicitly.

License

ROOT4DuckDB is distributed under the MIT License.

Developed by Seraphim S.
sserubin@jinr.ru

About

Query ROOT TTrees directly from DuckDB—without converting the data first. Nested-object decoding, predicate pushdown, dataset indexes, and Apache Iceberg metadata.

Topics

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages