4 ms·
CSV are nice for data interchange or even for storage if compressed (storage is interchange between the past and the future). But the very first thing one shoul
by plafl 5y ago
CSV are nice for data interchange or even for storage if compressed (storage is interchange between the past and the future). But the very first thing one should do when working with CSV data is convert it to a binary representation. It will force you to make a full pass over all the data, understand it, think about the types of the columns, and you will increase iteration speed.
More related to TFA: does someone recommend a good book or other resource about modern code optimization? I have interiorized several lessons but it would be nice to dig deeper.
- shoo 5y ago> recommend a good book or other resource about modern code optimization https://www.agner.org/optimize/ https://www.agner.org/optimize/
- plafl 5y agoThanks, and I must say, very old school web page:)
- errantmind 5y agoAgner Fog's writings are all well worth the read if you care about optimization. Some of the techniques he uses in his optimized versions of common C functions are really interesting. For example, he uses SIMD in several of these, but one of the performance problems associated with SIMD is the last 'n' bytes. If you processing, say, 16 or 32 bytes at a time, what do you do when you get towards the end of your buffer and that process would over-read past the end of the buffer? Most code I've seen just uses a simple linear process for the last few bytes but this is often very slow. Agner's solution was to just read past the end of the buffer, but use some assembly tricks to ensure there is always valid memory available after the end of the buffer, that way reading past the end of the buffer would (almost) never cause a problem.
- jandrewrogers 5y agoAnother useful trick for reading the last few bytes, especially if suitable buffer padding is not possible and the read would cross a memory page boundary, is to use PSHUFB aligned against the memory page boundary to pull the last few bytes into a register. Safe and fast.
- burntsushi 5y agoFor the specific case of substring/byte searching, another useful trick here is to do an unaligned load that lines up with the end of the haystack. Some portion of that load will have already been searched, but you already know nothing is in it.
- jhgb 5y ago> but use some assembly tricks to ensure there is always valid memory available after the end of the buffer Isn't rounding up your buffer's size an allocation technique rather than "as assembly trick"?
- speedgoose 5y agoSome years ago I had huge performance increase by simply converting a csv dataset to parquet before processing it with Apache Spark.
- mgradowski 5y agoAlso, you will sleep better at night knowing that your column dtypes are safe from harm, exactly as you stored them. Moving from CSV (or god forbid, .xlsx) has been such a quality of life improvement. One thing I miss though is how easy it is to inspect .csv and .xlsx. I kinda solved it using [1], but it only works on Windows. More portable recommendations welcome! [1] https://github.com/mukunku/ParquetViewer https://github.com/mukunku/ParquetViewer
- speedgoose 5y agoI used to use Zeppelin, some kind of Jupyter Notebook for Spark (that supports Parquet). But it may be better alternatives. https://zeppelin.apache.org/ https://zeppelin.apache.org/
- cb321 5y agoThe "real" format being binary with debugging tools is absolutely the best way to go. For example, you can use `nio print` (or even just `nio p`) in the Nim multi-program https://github.com/c-blake/nio https://github.com/c-blake/nio to get "text debugging output" of binary files.
- ZeroGravitas 5y agoI really like Visidata for "exploring" csv-type data. Its a vi(m) inspired tool. It also handles xls(x), sqlitedb and a bunch of other random things, and it appears to support parquet via pandas: https://www.visidata.org/docs/loading/ https://www.visidata.org/docs/loading/
- mgradowski 5y agoNice one!
- ComputerGuru 5y agoI have a one-liner alias/script csv_to_sqlite that just starts an sqlite3 shell with the csv converted to a table. It’s a lifesaver.