6 ms·
Nimble: A new columnar file format by Meta [video]
- mempko 2y agoI would love to see support in Apache Arrow to read this format. Parquet is already supported.
- nomel 2y agoThis makes me assume it also doesn't have proper support for multidimensional arrays.
- bionhoward 2y agoJust curious, how would you decide if an “it” did have proper support for multidimensional arrays?
- nomel 2y agoIf I put an nd array in, do I get the same nd array out, or do I have to serialize/deserialize myself, with some custom schema/code containing hacks like stuffing data (coordinates) into the column name?
- CharlesW 2y agoI learned that "Nimble" is the new name for "Alpha", discussed in this 2023 report: https://www.cidrdb.org/cidr2023/papers/p77-chattopadhyay.pdf https://www.cidrdb.org/cidr2023/papers/p77-chattopadhyay.pdf Here's an excerpt that may save some folks a click or three… > "While storing analytical and ML tables together in the data lakehouse is beneficial from a management and integration perspective, it also imposes some unique challenges. For example, it is increasingly common for ML tables to outgrow analytical tables by up to an order of magnitude. ML tables are also typically much wider, and tend to have tens of thousands of features usually stored as large maps. > "As we executed on our codec convergence strategy for ORC, it gradually exposed significant weaknesses in the ORC format itself, especially for ML use cases. The most pressing issue with the DWRF format was metadata overhead; our ML use cases needed a very large number of features (typically stored as giant maps), and the DWRF map format, albeit optimized, had too much metadata overhead. Apart from this, DWRF had several other limitations related to encodings and stripe structure, which were very difficult to fix in a backward-compatible way. Therefore, we decided to build a new columnar file format that addresses the needs of the next generation data stack; specifically, one that is targeted from the onset towards ML use cases, but without sacrificing any of the analytical needs. > "The result was a new format we call Alpha. Alpha has several notable characteristics that make it particularly suitable for mixed Analytical nd ML training use cases. It has a custom serialization format for metadata that is significantly faster to decode, especially for very wide tables and deep maps, in addition to more modern compression algorithms. It also provides a richer set of encodings and an adaptive encoding algorithm that can smartly pick the best encoding based on historical data patterns, through an encoding history loopback database. Alpha requires fewer streams per column for many common data types, making read coalescing much easier and saving I/Os, especially for HDDs. Alpha was written in modern C++ from scratch in a way that allows it to be extended easily in the future. > "Alpha is being deployed in production today for several important ML training applications and showing 2-3x better performance than ORC on decoding, with comparable encoding performance and file size."
- 0cf8612b2e1e 2y agoAlpha has got to be one of the worst names I have ever heard for a new product. Did they want to make it impossible to find?
- __MatrixMan__ 2y agoHow could a company called Meta be so shortsighted?
- santoshalper 2y agoWell played.
- keithalewis 2y agoYou Bet!
- d5dhcuvyv 2y ago[dead]
- isodev 2y agoAlpha was also the name of the virtual assistant owned by the bad guy in Extrapolations. https://www.imdb.com/title/tt13821126/ https://www.imdb.com/title/tt13821126/
- headwayoldest 2y agoYes, but before that he was helping Zordon and the Power Rangers. https://www.imdb.com/title/tt0106064/ https://www.imdb.com/title/tt0106064/
- isodev 2y agoHaha, good times!
- metadat 2y ago
- khaledh 2y agoFwiw, the name clashes with Nim's package manager nimble: https://github.com/nim-lang/nimble https://github.com/nim-lang/nimble
- deleted 2y ago[deleted]
- jauntywundrkind 2y agoThere's already been some interesting column format optimization work at Meta, as their Velox execution engine team worked with Apache Arrow to align their columnar formats. This talk is actually happening at VeloxCon, so there's got to be some awareness! https://engineering.fb.com/2024/02/20/developer-tools/velox-apache-arrow-15-composable-data-management/ https://engineering.fb.com/2024/02/20/developer-tools/velox-... https://news.ycombinator.com/item?id=39454763 https://news.ycombinator.com/item?id=39454763 I wonder how much if any overlap there is here, and whether it was intentional or accidentally similar. Ah, "return efficient Velox vectors" is on the list, but still seems likely to be some overlap in encoding strategies etc. The four main points seem to be: a) encoding metadata as part of stream rather than fixed metadata, b) nls are just another encoding, c) no stripe footer/only stream locations is in footer, d) FlatBuffers! Shout out to FlatBuffers, wasn't expecting to see them making a comeback! I do wish there were a lot more diagrams/slides. There's four bullet points, and Yoav Helfman talks to them, but there's not a ton of showing what he's talking about.
- hiyer 2y agoWhat's the future for datafusion (https://arrow.apache.org/datafusion/ https://arrow.apache.org/datafusion/) if Arrow is moving towards Velox?
- theLiminator 2y agoWhat do you mean by Arrow is moving towards Velox? Arrow is a standard for in-memory columnar data? At most arrow might adopt innovations made in Velox into its spec that datafusion will then adopt?
- hiyer 2y agoI meant the Arrow ecosystem. Datafusion is the query processing project therein at present, so I was curious to know what's the future for that. As you said, it's possible Datafition will adopt some stuff from Velox.
- yencabulator 2y ago
- gigatexal 2y agoHmm another conte for in the open table format space. Nice.
- gigatexal 2y ago*contender
- MaximilianEmel 2y agoIs there a quick description of the structure of it anywhere?
- deleted 2y ago[deleted]
- mrtimo 2y agoHow is this compare with parquet format?
- yigitkonur35 2y agoCurious about Clickhouse’s approach to this compression structure.
- horusporus 2y agoI was really hoping to see Cap'N Proto used for the format, since that has fast access without decoding, and reasonable backwards compatibility with old files. Anyone know why Flatbuffers were used?
- Kalanos 2y agoBy the time data has been preprocessed for ML, it is numerically encoded as floats, so .npy/npz is a good fit and `np.memmap` is an incredible way to seek into ndim data.
- RyanHamilton 2y agoParquet + Arrow hopefully seem to be emerging as standards. I would much rather see those standards improved than new formats emerge. Even within those existing formats there has become enough variation than some platforms only support a subset of functionality. That and the performance and size of the libraries is poor. e.g. DuckDB / Clickhouse Parquet nanosecond compatibility. https://github.com/duckdb/duckdb/issues/9852 https://github.com/duckdb/duckdb/issues/9852 e.g. The arrow SQL driver is 70+MB in java.
- HackerThemAll 2y agoThe Meta's own ORC is quite popular, too, in addition to Parquet, Arrow, Iceberg, Delta, Velox, Lance or Avro. So I assume the new one will find its way into lake houses/data warehouses as well. Because we need bigger mess and bloat.