Skip to content

[Data] Fix Dataset.summary() crash on boolean columns - #65541

Open
lonexreb wants to merge 2 commits into
ray-project:masterfrom
lonexreb:fix/62235-summary-bool
Open

lonexreb wants to merge 2 commits into
ray-project:masterfrom
lonexreb:fix/62235-summary-bool

Conversation

@lonexreb

Copy link
Copy Markdown
Contributor

Why are these changes needed?

Dataset.summary() crashes with ArrowNotImplementedError on any boolean column:

import ray
ds = ray.data.from_items([{"b": True}, {"b": False}, {"b": True}])
ds.summary()
# pyarrow.lib.ArrowNotImplementedError:
#   Function 'subtract' has no kernel matching input types (bool, double)

DataType.bool() was mapped to _numerical_aggregators, but Std, ApproximateQuantile, and ZeroPercentage rely on PyArrow kernels with no boolean implementation. I probed each default aggregator individually against a bool column on master: Count, Mean, Min, Max, MissingValuePercentage, and ApproximateTopK all work; Std and ZeroPercentage raise ArrowNotImplementedError (and a boolean quantile is not meaningful).

This PR gives boolean columns a dedicated _boolean_aggregators set — count, mean (= fraction of true values), min, max, missing-value percentage, and approximate top-k with k=2 (true/false counts) — and routes boolean dtypes to it in the heuristic fallback path as well.

Related issue number

Fixes #62235

Checks

  • I've signed off every commit (DCO).
  • Formatting checked with the repo-pinned black==22.10.0.
  • Testing strategy:
    • Added test_summary_boolean_column (end-to-end regression: summary() on a bool column with nulls, asserting count/mean/min/max/missing_pct values).
    • Added test_boolean_aggregators unit test; updated the two existing expectations that encoded "boolean is numerical".
    • pytest python/ray/data/tests/test_dataset_stats.py → 39 passed (run locally against master via Ray nightly wheel + setup-dev.py symlinks).

Notes for reviewers / AI-assistance disclosure

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces specialized aggregators for boolean columns in Ray Data stats, resolving an issue where boolean columns would crash the summary() method due to unsupported PyArrow kernels (like subtraction for standard deviation). It maps boolean data types to _boolean_aggregators instead of _numerical_aggregators and adds comprehensive unit tests to verify this behavior. I have no feedback to provide as there are no review comments to evaluate.

@ray-gardener ray-gardener Bot added data Ray Data-related issues community-contribution Contributed by the community labels Aug 18, 2026
@github-actions

github-actions Bot commented Sep 1, 2026

Copy link
Copy Markdown

This pull request has been automatically marked as stale because it has not had
any activity for 14 days. It will be closed in another 14 days if no further activity occurs.
Thank you for your contributions.

You can always ask for help on our discussion forum or Ray's public slack channel.

If you'd like to keep this open, just leave any comment, and the stale label will be removed.

@github-actions github-actions Bot added the stale The issue is stale. It will be closed within 7 days unless there are further conversation label Sep 1, 2026
@lonexreb
lonexreb force-pushed the fix/62235-summary-bool branch from 2b507f5 to 0913560 Compare September 4, 2026 06:55
@lonexreb

lonexreb commented Sep 4, 2026

Copy link
Copy Markdown
Contributor Author

Not stale — rebased onto current master and re-ran the test suite locally (test_dataset_stats.py: 39 passed). Ready for review whenever a maintainer has bandwidth.

@github-actions github-actions Bot added unstale A PR that has been marked unstale. It will not get marked stale again if this label is on it. and removed stale The issue is stale. It will be closed within 7 days unless there are further conversation labels Sep 4, 2026
@iamjustinhsu

iamjustinhsu commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Hi @lonexreb, how hard would it be to support booleans in those preprocessors? I think I would prefer that over skipping (since all booleans can be converted to numbers)

DataType.bool() was mapped to the numerical aggregators, but Std,
ApproximateQuantile, and ZeroPercentage rely on PyArrow kernels (e.g.
subtract) that have no boolean implementation, so summary() raised
ArrowNotImplementedError on any bool column.

Give boolean columns their own aggregator set: count, mean (fraction of
true values), min, max, missing-value percentage, and approximate top-k
(true/false counts). Also route boolean dtypes to it in the heuristic
fallback path.

Fixes ray-project#62235

Signed-off-by: lonexreb <reach2shubhankar@gmail.com>
…ing stats

Review feedback: rather than omitting Std/ApproximateQuantile/
ZeroPercentage for boolean columns, treat booleans as 0/1 numbers so
they get the full numerical statistics:

- ArrowBlockColumnAccessor.sum_of_squared_diffs_from_mean casts boolean
  columns to float64 (PyArrow has no boolean 'subtract' kernel).
- ZeroPercentage casts boolean columns to int8 before the zero
  comparison, in both the per-block path and the vectorized
  zero_pct_spec path ('equal(bool, int)' has no kernel).
- DataType.bool() maps back to _numerical_aggregators; the dedicated
  boolean aggregator set is removed.

summary() on a bool column now reports count, mean (fraction true),
min, max, std, median, missing_pct, and zero_pct (fraction false).

Signed-off-by: lonexreb <reach2shubhankar@gmail.com>
@lonexreb
lonexreb force-pushed the fix/62235-summary-bool branch from 0913560 to e77b374 Compare October 1, 2026 06:01
@lonexreb

lonexreb commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

Great suggestion @iamjustinhsu — done in e77b374. Booleans now get the full numerical statistics (treated as 0/1) instead of a reduced set:

  • ArrowBlockColumnAccessor.sum_of_squared_diffs_from_mean casts boolean columns to float64 (PyArrow has no boolean subtract kernel) → Std works.
  • ZeroPercentage casts boolean columns to int8 before the zero comparison in both its per-block path and the vectorized zero_pct_spec path (equal(bool, int) has no kernel) → zero_pct = fraction of False.
  • ApproximateQuantile already handled booleans (datasketches treats them as 0/1).
  • DataType.bool() maps back to _numerical_aggregators; the dedicated boolean aggregator set is gone.

summary() on a bool column now reports count, mean (fraction true), min, max, std, median, missing_pct, and zero_pct. The regression test asserts the actual values (e.g. std([1,0,1]), zero_pct == 33.3%), and the casts benefit standalone ds.aggregate(Std(on=bool_col)) / ZeroPercentage(on=bool_col) too.

Rebased onto current master; test_dataset_stats.py passes locally (38 tests).

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.

Fix All in Cursor

Reviewed by Cursor Bugbot for commit e77b374. Configure here.

arrow_compatible = column_accessor._to_arrow_compatible_container()
if pa.types.is_boolean(arrow_compatible.type):
# `equal(bool, int)` has no kernel; treat booleans as 0/1.
arrow_compatible = pc.cast(arrow_compatible, pa.int8())

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Boolean cast assumes Arrow container

Medium Severity

The new boolean handling reads .type from _to_arrow_compatible_container(), but the pandas accessor returns a Python list. ZeroPercentage on a pandas boolean column then fails with AttributeError instead of treating False as zero.

Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit e77b374. Configure here.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

community-contribution Contributed by the community data Ray Data-related issues unstale A PR that has been marked unstale. It will not get marked stale again if this label is on it.

2 participants