Skip to content

[Data] Support appending a subset of columns to a Lance dataset. - #64474

Merged
goutamvenkat-anyscale merged 3 commits into
ray-project:masterfrom
xuxiaoqiang666:data/lance-append-column-subset
Sep 3, 2026
Merged

goutamvenkat-anyscale merged 3 commits into
ray-project:masterfrom
xuxiaoqiang666:data/lance-append-column-subset

Conversation

@xuxiaoqiang666

Copy link
Copy Markdown
Contributor

Description

Dataset.write_lance(..., mode="append") fails when the data being appended
contains only a subset of the target dataset's columns. On append, the sink
sets the write schema to the full existing dataset schema, and the positional
RecordBatchReader.from_batches(schema, ...) path then errors on the columns
the incoming blocks don't provide (reorder_columns_by_schema → Table.select
on a missing field), surfacing as a confusing LanceError / OSError instead
of anything actionable.

This PR makes APPEND align each block to the dataset schema and null-fill
the columns the block omits
— the missing columns become null for the newly
appended rows, analogous to how Iceberg/Delta handle schema evolution on
append. If a block contains a column that is not in the target dataset, a
clear ValueError is raised (adding brand-new columns is out of scope; select
the existing columns first via ds.select_columns(...)).

Scope / compatibility:

  • Only SaveMode.APPEND changes. CREATE / OVERWRITE keep their existing
    strict behavior (still use reorder_columns_by_schema).
  • No public API change — composes with ds.select_columns(...).
  • Extension types (e.g. arrow.json) and nested types are handled when
    building the null-filled columns.

Before (fails):

import ray, lance, pyarrow as pa

uri = "/tmp/lance_subset"
lance.write_dataset(
    pa.table({"a": [1, 2, 3], "b": [4, 5, 6], "c": [7, 8, 9]}), uri, mode="overwrite"
)
# Data provides only (a, b); the table is (a, b, c)
ray.data.from_items([{"a": 10, "b": 40}]).write_lance(uri, mode="append")
# -> OSError: LanceError(Arrow): ... KeyError: 'Field "c" does not exist in schema'
After (this PR): the append succeeds and the new rows have c = null.

Related issues
N/A

Additional information
Tests added in python/ray/data/tests/datasource/test_lance.py:

test_align_block_to_schema_reorders_fills_and_validates — unit test for the
alignment helper: reorders columns, null-fills missing columns, and raises on
columns not in the target schema.
test_null_column_supports_extension_and_nested_types — null-column
construction for extension (arrow.json-like) and nested types.
test_lance_write_append_fills_missing_columns_with_null — end-to-end:
appending (a, b) rows to an (a, b, c) table writes c as null.
@xuxiaoqiang666
xuxiaoqiang666 requested a review from a team as a code owner July 1, 2026 10:44

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces support for appending data with a subset of the target schema's columns in LanceDatasink by filling missing columns with nulls. It implements helper functions _null_column and _align_block_to_schema to align PyArrow tables to the target schema, along with comprehensive unit tests. The review feedback suggests an improvement opportunity in _align_block_to_schema to optimize column lookup complexity from O(M * N) to O(M + N) by using a set, and to pass the field object instead of field.name to tbl.append_column to preserve field metadata.

Comment on lines +127 to +131
for field in schema:
if field.name not in tbl.schema.names:
tbl = tbl.append_column(
field.name, _null_column(field.type, tbl.num_rows)
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Improvement Opportunity: Optimize Lookups and Preserve Field Metadata

  1. Performance: tbl.schema.names is a list. Checking field.name not in tbl.schema.names inside a loop results in an $O(M \times N)$ complexity (where $M$ is the number of fields in schema and $N$ is the number of fields in tbl). Converting tbl.schema.names to a set before the loop reduces this to $O(M + N)$.
  2. Metadata Preservation: Passing field (a pa.Field object) instead of field.name (a string) to tbl.append_column ensures that the exact field metadata, nullability, and other properties from the target schema are preserved rather than using default values.
Suggested change
for field in schema:
if field.name not in tbl.schema.names:
tbl = tbl.append_column(
field.name, _null_column(field.type, tbl.num_rows)
)
existing_names = set(tbl.schema.names)
for field in schema:
if field.name not in existing_names:
tbl = tbl.append_column(
field, _null_column(field.type, tbl.num_rows)
)
@xuxiaoqiang666
xuxiaoqiang666 force-pushed the data/lance-append-column-subset branch from 98f8402 to 3ed8b6e Compare July 1, 2026 10:50
Appending data whose schema is a subset of the target Lance dataset used to
fail: the append path forces the full dataset schema and the positional writer
errored on the missing columns (surfacing as a confusing LanceError/OSError).

This aligns each block to the dataset schema on APPEND and null-fills the
columns the block omits (like Iceberg/Delta schema evolution). CREATE and
OVERWRITE keep their existing strict behavior. No public API change --
composes with ds.select_columns(...).

Signed-off-by: xuxiaoqiang13 <xuxiaoqiang13@jd.com>
@xuxiaoqiang666
xuxiaoqiang666 force-pushed the data/lance-append-column-subset branch from 3ed8b6e to 2f7417f Compare July 1, 2026 12:32
@ray-gardener ray-gardener Bot added data Ray Data-related issues community-contribution Contributed by the community labels Jul 1, 2026
@xuxiaoqiang666

Copy link
Copy Markdown
Contributor Author

@abrarsheikh can you please review it?

@github-actions

Copy link
Copy Markdown

This pull request has been automatically marked as stale because it has not had
any activity for 14 days. It will be closed in another 14 days if no further activity occurs.
Thank you for your contributions.

You can always ask for help on our discussion forum or Ray's public slack channel.

If you'd like to keep this open, just leave any comment, and the stale label will be removed.

@github-actions github-actions Bot added the stale The issue is stale. It will be closed within 7 days unless there are further conversation label Jul 17, 2026
@github-actions

Copy link
Copy Markdown

This pull request has been automatically closed because there has been no more activity in the 14 days
since being marked stale.

Please feel free to reopen or open a new pull request if you'd still like this to be addressed.

Again, you can always ask for help on our discussion forum or Ray's public slack channel.

Thanks again for your contribution!

@github-actions github-actions Bot closed this Jul 31, 2026
@richardliaw richardliaw removed the stale The issue is stale. It will be closed within 7 days unless there are further conversation label Jul 31, 2026
@richardliaw richardliaw reopened this Jul 31, 2026
@github-actions

Copy link
Copy Markdown

This pull request has been automatically marked as stale because it has not had
any activity for 14 days. It will be closed in another 14 days if no further activity occurs.
Thank you for your contributions.

You can always ask for help on our discussion forum or Ray's public slack channel.

If you'd like to keep this open, just leave any comment, and the stale label will be removed.

@github-actions github-actions Bot added the stale The issue is stale. It will be closed within 7 days unless there are further conversation label Aug 14, 2026
@github-actions github-actions Bot added unstale A PR that has been marked unstale. It will not get marked stale again if this label is on it. and removed stale The issue is stale. It will be closed within 7 days unless there are further conversation labels Aug 15, 2026
@bveeramani bveeramani added this to the Data issue and PR backlog milestone Aug 19, 2026
@goutamvenkat-anyscale
goutamvenkat-anyscale enabled auto-merge (squash) September 3, 2026 00:03
@github-actions github-actions Bot added the go add ONLY when ready to merge, run all tests label Sep 3, 2026
@goutamvenkat-anyscale
goutamvenkat-anyscale merged commit 0166706 into ray-project:master Sep 3, 2026
7 of 8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

community-contribution Contributed by the community data Ray Data-related issues go add ONLY when ready to merge, run all tests unstale A PR that has been marked unstale. It will not get marked stale again if this label is on it.

4 participants