6 ms·
I shudder to think who needs to process a million lines of csv that fast...
by voidUpdate 1y ago
I shudder to think who needs to process a million lines of csv that fast...
- hermitcrab 1y agoFor all its many weaknesses, I believe CSV is still the most common data interchange format.
- adra 1y agoErm, maybe file based? JSON is the king if you count exchanges worldwide a sec. Maybe no 2 is form-data which is basically email multipart, and if course there's email as a format. Very common =)
- hermitcrab 1y agoI meant file-based.
- devmor 1y agoI honestly wonder if JSON is king. I used to think so until I started working in fintech. XML is unfortunately everywhere.
- hermitcrab 1y agoJSON isn't great for tabular data. And an awful lot of data is tabular.
- devmor 1y agoYeah, I don’t like parsing XML, but I’d rather do that than deal with the Lovecraftian API design that comes with complex JSON representations.
- hajile 1y agoJSON tabular data only adds a couple of brackets per line and at the start/end of the file vs CSV. In exchange for these bits (that basically disappear when compressed), you get a guaranteed standard formatting. Seems like a decent tradeoff to me.
- arcfour 1y agoJSON: because XML is too hard. Developers: hey, let's hack everything XML had back onto JSON except worse and non-standardized. Because it turns out you need those things sometimes!
- segmondy 1y agolots of folks in Finance, you can share csv with any Finance company and they can process it. It's text.
- zzbn00 1y agoHumans generate decisions / text information at rates of ~bytes per second at most. There is barely enough humans around to generate 21GB/s of information even if all they did was make financial decisions! So 21 GB/s would be solely algos talking to algos... Given all the investment in the algos, surely they don't need to be exchanging CSV around?
- internetter 1y ago> Humans generate decisions / text information at rates of ~bytes per second at most Yes, but the consequences of these decisions are worth much more. You attach an ID to the user, and an ID to the transaction. You store the location and time where it was made. Ect.
- zzbn00 1y agoI think these would add only small amount of information (and in a DB would be modelled as joins). Only adds lots of data if done very inefficiently.
- jajko 1y agoWhy are you theoretising? I can tell you from out there its used massively, and its not going away in contrary. Even rather small banks can end up generating various reports etc. which can easily become huge. The speed of human decision has basically 0 role here, as it doesn't with messaging generally, there is way more to companies than just direct keyboard-to-output link.
- adrianN 1y agoYou might have accumulated some decades of data in that format and now want to ingest it into a database.
- sunrunner 1y agoI shudder to think of what it means to be storing the _results_ of processing 21 GB/s of CSV. Hopefully some useful kind of aggregation, but if this was powering some kind of search over structured data then it has to be stored somewhere...
- devmor 1y agoJust because you’re processing 21GB/s of CSV doesn’t mean you need all of it. If your data is coming from a source you don’t own, it’s likely to include data you don’t need. Maybe there’s 30 columns and you only need 3 - or 200 columns and you only need 1. Enterprise ETL is full of such cases.
- trollbridge 1y agoIt's become a very common interchange format, even internally; it's also easy to deflate. I have had to work on codebases where CSV was being pumped out at basically the speed of a NIC card (its origin was Netflow, and then aggregated and otherwise processed, and the results sent via CSV to a master for further aggregation and analysis). I really don't get, though, why people can't just use protocol buffers instead. Is protobuf really that hard?
- nobleach 1y agoExtremely hard to tell an HR person, "Right-click on here in your Workday/Zendesk/Salesforce/etc UI and export a protobuf". Most of these folks in the business world LIVE in Excel/Spreadsheet land so a CSV feels very native. We can agree all day long that for actual data TRANSFER, CSV is riddled with edge cases. But it's what the customers are using.
- heavenlyblue 1y agoIt's extremely unlikely they need to load spreadsheets large enough for 21Gb/s speed to matter
- matja 1y agoKind of, there isn't a 1:1 mapping of protobuf wire types to schema types, so you need to package the protobuf schema with the data and compile it to parse the data, or decide on the schema before-hand. So now you need to decide on a file format to bundle the schema and the data.
- 1y ago
- moregrist 1y agoI have. I think it's a pretty easy situation for certain kinds of startups to find themselves in: - Someone decides on CSV because it's easy to produce and you don't have that much data. Plus it's easier for the <non-software people> to read so they quit asking you to give them Excel sheets. Here <non-software people> is anyone who has a legit need to see your data and knows Excel really well. It can range from business types to lab scientists. - Your internal processes start to consume CSV because it's what you produce. You build out key pipelines where one or more steps consume CSV. - Suddenly your data increases by 10x or 100x or more because something started working: you got some customers, your sensor throughput improved, the science part started working, etc. Then it starts to make sense to optimize ingesting millions or billions of lines of CSV. It buys you time so you can start moving your internal processes (and maybe some other teams' stuff) to a format more suited for this kind of data.
- nly 1y agoWe use parquet extensively at work, and it's really slow to ingest. Slower than a hand rolled binary column oriented format. Sometimes using something standardized is just worth it though.
- deleted 1y ago[deleted]
- ourmandave 1y agoThat cartesian product file accounting sends you at year end?
- constantcrying 1y agoIn basically every situation it is inferior to HDF5. I do not think there is an actual explanation besides ignorance, laziness or "it works".
- pak9rabid 1y agoUgh.....I do unfortunately.