6 ms·
wait… if importing malformed csvs gets automated that’s like half of a data professional’s job gone in a poof of smoke /s. jk -- great use case so often w/ pan
by data_ders 4y ago
wait… if importing malformed csvs gets automated that’s like half of a data professional’s job gone in a poof of smoke /s. jk -- great use case
so often w/ pandas I’d:
1. “yeet” the csv into a dataframe
2. use dataframe methods to massage the data to a “clean” state
3. push as much of the df methods into pd.read_csv() parameter options
it’s be great to iterate more quickly on the above loop. better yet — what if it would could auto-generate a letter to send to the folks from whom you got this data on how they could better output to csv to make ingestion simpler and easier for downstream users…. but maybe that letter would just be "don't use CSV!"
related to flat data formats, it obviously makes sense to start with CSV, but what about the future? If this tool became ubiquitous, how might a SWE or data professional's job change? What opportunities be created? As in:
1. CSV is ubiquitous but has no singularly well-adopted standard.
2. software and data engineers struggle with CSVs as a result of #1.
3. tool is created to reduce pain and friction.
4. profit? a new market? a new standard?
Last, but most personally interestingly, how much do you know about the Apache Arrow ecosystem and how it's mission might overlap with YoBulk's
- sgerenser 4y ago1. CSV is ubiquitous but has no singularly well-adopted standard. 2. software and data engineers struggle with CSVs as a result of #1. 3. tool is created to reduce pain and friction. 4. profit? a new market? a new standard? The "revolutionary new tool" to replace CSVs was XML in the late 1990s.
- nerdponx 4y agoExcept not at all. XML is harder to enter by hand than CSV and not mess it up. CSV optimizes for the easiest cases and performs well on them. XML optimizes for the most complicated cases and therefore performs poorly on the easiest cases. JSON is somewhere in the middle. The main problems with CSV have to do with 1) MS Excel and 2) some kind of delusion among programmers that formatting or parsing arbitrary data is easy and you don't need a library for it, so you get hand-rolled generators and parsers that emit broken files. Otherwise, the problem with CSV has little to do with the CSV format as such and more to do with the fact that the data is stringly-typed. XML has the same problem. JSON interestingly does not. Everything has tradeoffs.
- yosai 4y ago@nerdponx You are spot on..
- aforwardslash 4y agonitpick, I wouldn't place JSON in the middle (the lack of proper integers and precision problems is one of the issues). but other than that, spot on.
- refulgentis 4y agoThat’s really interesting, I wonder if this simplifies down to “you want CSV with column typing and a typesafe CSV editor”, as you note, JSONs win is the lack of issues with typing, and CSV really isn’t complicated at all except for that property. JSON is just a row with keys that are column.
- nerdponx 4y agoI definitely want that! Parquet is great for data interchange, but it's not easily hand-editable. I wonder if there's an open niche in the software world for an Excel-like data entry and manipulation tool, but with stronger/stricter typing of cells and columns, and with direct export to and import from SQLite and Parquet.
- Fnoord 4y agoTo solve the XML issues you described we got schemas and syntax highlighting. I hate non-prettified JSON but its easy to prettify in any editor. So its a meh argument against JSON. But to solve the crap with the comma one needs a variant of JSON, and there's various of these... One other neat feature of CSV is it can be imported in a very popular and powerful IDE, called... Excel.
- yosai 4y ago@sgerenser Yes CSV is everywhere..YoBulk is smartly positioning itself for the data donor/provider or customer..It is the end customer or data provider who bite the bullet and do the time consuming data cleaning.The customer should know about the errors, duplicates,PII data,inconsitency in data.Customer has to be properly guided to clean the data in best possible manner.
- boringg 4y ago"half of a data professional's job" ... you mean like 90%
- textninja 4y agoIt’s useful to make a distinction between what is one’s job and where one spends his time - that is, the theory and practice of job titles. Data cleaning is 0% of a Data Scientist’s job despite occupying 90% of their time. My hope is that AI can help bridge that gap.
- yosai 4y agoyes textninja.YoBulk's vision is to automate the first mile data onboarding and cleaning through AI so that Data scientists are free from doing any mundane task of data cleaning.
- yosai 4y ago@data_ders We realized that more than 70% of business data is shared in CSV and Excel formats, and only a small percentage use API integrations for data exchange..So CSV is here to stay for sure..On the other side,data engine is a sub module inside YoBulk.We are trying to automate complete CSV importing workflow..mostly solving the CSV errors in a collaborative way with the data donor..Yobulk's USP is how we show the errors in human readable way.We have written wrapper on top of some open source data validation engines.Yes..I have used apache Arrow..we not competing with Apache Arrow..We are creating an alternative to flatfile.com
- chaps 4y agoMan, been down this path for a long while. It gets tough! Flattening csvs with hierarchical headers (as in, headers that that apply a category to a second row of headers) are tough. The ways csv can fail is just fucking nuts. Especially when they're half hand written, half automated, or where a failure is 20m rows in. Hard to have speed and strong checks simultaneously.
- yosai 4y agoYes you are right..In YoBulk we flatten the CSV to a JSON schema store it in a document DB and do all the validations.Chunking the CSV and analysing the stream buffers for validation is giving us speed also.
- chaps 4y agoHave you been able to get something that might match a relational database? Auto-generation of a relational schema from a large dataset, or multiple datasets, is a deeply interesting idea.
- anonymouse008 4y agoWould you really pay for this? I made one for a client analyzing some X million line csvs by sharding the records then computing across 500 lambda instances to arrive at a schema
- refulgentis 4y agoyeet is dispose of with haste, not “move, but zoomer” or “sloppily with haste”
- chaps 4y agoWot, no. It's throw hastily we with no care of what it splats into. Pretty sure it came from a video of someone throwing something (food or drink?) in a crowded highschool hallway, from within the hallway. It's chaotic, reckless energy in.. mostly.. harmless form.
- recursive 4y ago> CSV is ubiquitous but has no singularly well-adopted standard RFC4180 exists regardless of adoption level. In a way, the simplicity of the spec causes the proliferation of grammars. No one thinks they can just yolo a PDF by hand in a text editor. Ok, maybe PDF is a bad example. But CSV (as specified in RFC4180) is so dead simple that people take shortcuts.
- yosai 4y agoYes you are absolutely right..We need a solution beyond standard as 80% businesses run on CSV..
- ed_elliott_asc 4y ago99.99%?
- groestl 4y agoAnd that's only the portion the businesses know about.
- thedudeabides5 4y agoworking on it, this is a hard problem actually
- silent_cal 4y agopd.yeet("/data/data_1.csv") pd.yeet(lambda x: yeet(x)) pd.yeet_to_csv("/clean_data/cleaned_data_1.csv")
- anothernewdude 4y agoYou'd have to be insane to trust GPT that much. I wouldn't want anything hallucinated in my data.