ray icon indicating copy to clipboard operation
ray copied to clipboard

[Datasets] Correct schema unification for Datasets with ragged Arrow arrays

Open scottjlee opened this issue 3 years ago • 0 comments

Signed-off-by: Scott Lee [email protected]

Why are these changes needed?

When creating Datasets with ragged arrays, the resulting Dataset incorrectly uses ArrowTensorArray instead of ArrowVariableShapedTensorArray as the underlying schema type. This PR refactors existing logic for schema unification into a separate function, which is now called during Arrow table concatenation and schema fetching to correct type promotion involving ragged arrays.

Related issue number

Closes #30082

Checks

  • [x] I've signed off every commit(by using the -s flag, i.e., git commit -s) in this PR.
  • [x] I've run scripts/format.sh to lint the changes in this PR.
  • [ ] I've included any doc changes needed for https://docs.ray.io/en/master/.
  • [ ] I've made sure the tests are passing. Note that there might be a few flaky tests, see the recent failures at https://flakey-tests.ray.io/
  • Testing Strategy
    • [x] Unit tests
    • [ ] Release tests
    • [ ] This PR is not tested :(

scottjlee avatar Dec 13 '22 20:12 scottjlee