[Data] Support appending a subset of columns to a Lance dataset. - #64474
goutamvenkat-anyscale merged 3 commits into
Conversation
There was a problem hiding this comment.
Code Review
This pull request introduces support for appending data with a subset of the target schema's columns in LanceDatasink by filling missing columns with nulls. It implements helper functions _null_column and _align_block_to_schema to align PyArrow tables to the target schema, along with comprehensive unit tests. The review feedback suggests an improvement opportunity in _align_block_to_schema to optimize column lookup complexity from O(M * N) to O(M + N) by using a set, and to pass the field object instead of field.name to tbl.append_column to preserve field metadata.
| for field in schema: | ||
| if field.name not in tbl.schema.names: | ||
| tbl = tbl.append_column( | ||
| field.name, _null_column(field.type, tbl.num_rows) | ||
| ) |
There was a problem hiding this comment.
Improvement Opportunity: Optimize Lookups and Preserve Field Metadata
-
Performance:
tbl.schema.namesis a list. Checkingfield.name not in tbl.schema.namesinside a loop results in an$O(M \times N)$ complexity (where$M$ is the number of fields inschemaand$N$ is the number of fields intbl). Convertingtbl.schema.namesto a set before the loop reduces this to$O(M + N)$ . -
Metadata Preservation: Passing
field(apa.Fieldobject) instead offield.name(a string) totbl.append_columnensures that the exact field metadata, nullability, and other properties from the target schema are preserved rather than using default values.
| for field in schema: | |
| if field.name not in tbl.schema.names: | |
| tbl = tbl.append_column( | |
| field.name, _null_column(field.type, tbl.num_rows) | |
| ) | |
| existing_names = set(tbl.schema.names) | |
| for field in schema: | |
| if field.name not in existing_names: | |
| tbl = tbl.append_column( | |
| field, _null_column(field.type, tbl.num_rows) | |
| ) |
98f8402 to
3ed8b6e
Compare
Appending data whose schema is a subset of the target Lance dataset used to fail: the append path forces the full dataset schema and the positional writer errored on the missing columns (surfacing as a confusing LanceError/OSError). This aligns each block to the dataset schema on APPEND and null-fills the columns the block omits (like Iceberg/Delta schema evolution). CREATE and OVERWRITE keep their existing strict behavior. No public API change -- composes with ds.select_columns(...). Signed-off-by: xuxiaoqiang13 <xuxiaoqiang13@jd.com>
3ed8b6e to
2f7417f
Compare
|
@abrarsheikh can you please review it? |
|
This pull request has been automatically marked as stale because it has not had You can always ask for help on our discussion forum or Ray's public slack channel. If you'd like to keep this open, just leave any comment, and the stale label will be removed. |
|
This pull request has been automatically closed because there has been no more activity in the 14 days Please feel free to reopen or open a new pull request if you'd still like this to be addressed. Again, you can always ask for help on our discussion forum or Ray's public slack channel. Thanks again for your contribution! |
|
This pull request has been automatically marked as stale because it has not had You can always ask for help on our discussion forum or Ray's public slack channel. If you'd like to keep this open, just leave any comment, and the stale label will be removed. |
0166706
into
ray-project:master
Description
Dataset.write_lance(..., mode="append")fails when the data being appendedcontains only a subset of the target dataset's columns. On append, the sink
sets the write schema to the full existing dataset schema, and the positional
RecordBatchReader.from_batches(schema, ...)path then errors on the columnsthe incoming blocks don't provide (
reorder_columns_by_schema→Table.selecton a missing field), surfacing as a confusing
LanceError/OSErrorinsteadof anything actionable.
This PR makes APPEND align each block to the dataset schema and null-fill
the columns the block omits — the missing columns become null for the newly
appended rows, analogous to how Iceberg/Delta handle schema evolution on
append. If a block contains a column that is not in the target dataset, a
clear
ValueErroris raised (adding brand-new columns is out of scope; selectthe existing columns first via
ds.select_columns(...)).Scope / compatibility:
SaveMode.APPENDchanges.CREATE/OVERWRITEkeep their existingstrict behavior (still use
reorder_columns_by_schema).ds.select_columns(...).arrow.json) and nested types are handled whenbuilding the null-filled columns.
Before (fails):