3 ms·
This actually addresses a huge hole right now in the ecosystem. At the moment, Arrow treats things that are logically the same as different types, for example a
by sakras 2y ago
This actually addresses a huge hole right now in the ecosystem. At the moment, Arrow treats things that are logically the same as different types, for example a dictionary-encoded utf8 is different from a utf8. Really there needs to be a distinction between logical and physical types, and it’s a big source of rough edges in things like Acero. Hacking on my own query engine on the side, I’ve put some thought into how to group physical into logical types in Arrow.
I am however worried that the place to define the logical type system might be inside Arrow instead of Arrow consumers. If everyone has their own logical type system, we’re just going to end up with the same incompatibilities that Arrow was trying to solve.
- nerdponx 2y agoSnowflake and Pandas already have the ability to decode/encode "logical" type information stored in a Parquet file header, which is how you can do things like round-trip categorical-dtype data and data frame indexes between Pandas and Parquet. So I think things are already heading in that direction, but a coherent vision is absolutely necessary to keep it up and drive adoption.