19 ms·
Golang – encoding/csv: Reading is slow
- zephyrfalcon 10y agoPython's csv module uses an internal module _csv which is written in C. So I'm not sure it's all that surprising that a Go implementation is a bit slower.
- chrisper 10y agoYou should contribute this knowledge to the GitHub issue.
- jlarocco 10y agoA good idea, but I think it's common knowledge. Lots of modules in the stdlib are written in C.
- chrisper 10y agoWell, it was news to me. Maybe because I only use Python casually.
- baq 10y agoultimately it's of academic interest - the end user doesn't care what precisely is going on under the hood.
- chrisper 10y agoI thought it is rather part of the solution/answer to the issue. It is expected to be slower, so it's not necessarily broken. Kind of it's a feature, not a bug.
- 0xFFC 10y agoHow about Java? It is quite funny Java version is much faster than Python version even when Python version does use C ? Something fishy going on.
- Anderkent 10y agoThe java one doesn't do as much as the python one: `line.split(',')` doesn't handle quoted commas, escape characters, different csv separators, etc.
- andrepd 10y agoYeah, calling from python introduces a layer of inefficiency.
- justin66 10y agoIt's as if there's more to software quality than the choice of tools.
- jerf 10y agoThere's nothing mysterious about the Java version... that's not a CSV parser. The other two things are truly CSV parsers. (Inasmuch as there is such a thing for such an ill-defined format. (No, the RFC is not determinative.)) It's easy to be faster if you do fundamentally less. Not necessarily wrong, depending on your task, but it's not comparable.
- __s 10y agoThe java source is using str.split(',') which is completely different than Python's offering of dialect/delimiter/etc https://github.com/python/cpython/blob/master/Lib/csv.py#L24 https://github.com/python/cpython/blob/master/Lib/csv.py#L24
- nradov 10y agoThe Java code is defective. It's not checking for double quotes. The CSV format allows for commas inside column values by surrounding with double quotes, and then you can also put double quotes within such values by escaping them as double double quotes. Fix those defects and the Java code will be a little slower. With modern JVMs, Java can occasionally actually be faster than native compiled languages due to dynamic optimization at runtime.
- Cyph0n 10y agoI don't think the answer is that simple. Perhaps there's a key optimization that the Go team didn't consider.
- aartur 10y agoI run the benchmark using PyPy (which doesn't have this C extension) and got a result about 20% slower compared to CPython (ie. still faster than Go). EDIT. I also did a funny thing and replaced the CPython C _csv.so extensions with pure Python version _csv.py, from PyPy. It run about 80 (eighty) times slower. It shows what wonders does JIT do (at least to some code).
- jjawssd 10y agoWould be a great experiment to Cythonize PyPy's _csv.so
- masklinn 10y agoThat sounds completely worthless, PyPy doesn't need a Cython version, and its library is a Python version of CPython's native csv module.
- chrisseaton 10y agoPyPy doesn't have a _csv.so - that was the main point of the comment you replied to.
- rdtsc 10y agoBut not sure if it matters. Go is free to use a C module to load csv files as well. It can use assembly or other tricks as well perhaps.
- weberc2 10y agoThere's a performance cost to calling between c and Go, and sharing memory between the two makes for hard to predict GC behavior. I doubt it would be faster than a pure Go implementation.
- djur 10y agoIt seems pretty common for languages to start out with a relatively unoptimized CSV parser (if they have one at all) and then get a faster one contributed by the community once there's enough interest. Ruby had that happen with FasterCSV. The Java comparison here seems inapt, because it doesn't do as much as the other two. It's just a naive "split on commas" implementation that wouldn't handle quoted cells. Really, if Go's CSV reader is only 200% slower than that and 50% slower than Python's optimized C implementation, that's pretty good already.
- justincormack 10y agoIt also seems to support UTF8, which is perhaps common now. (CSV doesnt tell you what encoding it is, in the old days it was 95% ASCII).
- masklinn 10y ago> CSV doesnt tell you what encoding it is, in the old days it was 95% ASCII Or some non-ascii codepage you're not told about and have to guess (e.g. Excel generates CSV in CP1250 by default, with an option to export UTF-16)
- coldtea 10y agoAll 3 support UTF8. The difference with the Go implementation is the mostly useless capability to have a utf-8 multibyte item as the delimiter.
- masklinn 10y ago1. there's nothing useless about it 2. the Python 3 CSV library supports arbitrary codepoints as delimiter, quote character and escape character (if applicable)
- geofft 10y ago> 1. there's nothing useless about it Have you ever seen a "C"SV with a multibyte sequence as a delimiter? I haven't. Even if such a thing exists, the feature is of negative utility if it slows down CSV parsing for everyone else. If you must, write two implementations, and use the slow path if your delimiter is multibyte.
- tmaly 10y agoI wrote my own in Go that is blazing fast using bytes. I know the data is ascii so I was able to use that to my advantage.
- piinbinary 10y agoI did the same [0]. It runs almost as fast as the Java implementation. I'd be interested to see how yours works if you are willing to share it. Edit: Plus one that is ~2x faster than Java by avoiding allocations [1]. [0] https://gist.github.com/jmikkola/6ac96ad6d6f66e772c33ec41ed27f057 https://gist.github.com/jmikkola/6ac96ad6d6f66e772c33ec41ed2... [1] https://gist.github.com/jmikkola/7ded8392226b7659c881f5540be539f5 https://gist.github.com/jmikkola/7ded8392226b7659c881f5540be...
- SeanDav 10y agoEffectively one is comparing library performance here and not language performance. Granted, that line can get very blurry indeed, but in this case this says very little about golang the language and far more about a current implementation of one of the golang libraries.
- jonlawlor 10y agoIt is kind of both; go doesn't allow some approaches in native go code that can make it slower than other languages. (I love go, but that is my experience.) In this case the choice to use utf-8 everywhere, including in the csv delimiters, is making it slower.
- weberc2 10y agoThis isn't what makes CSV parsing slow, there's nothing about Go that requires you to deal in UTF8.
- paulddraper 10y agoThe Java code is not a CSV parser. I added the results of using Apache Commons CSV to the GitHub thread. After using that, Python was actually by far the fastest. Hooray for performance sensitive code in C :)
- paulddraper 10y agoFYI, I later went back and ran PyPy, which implements csv in pure Python. Even including start-up time, it was nearly at CPython speed.
- petters 10y agoAny CSV reader should be limited only be disk access right? I wrapped together a C++ program solving this problem and got 0.124 seconds. But that does not do quotations etc.
- _ph_ 10y agoI would guess, if you directly translate your program to Go it would not be much slower. On low level code, Go 1.7 gets quite close to GCC. The point is, that doing the quotations right, especially if you allow non-ascii quotes, eats a lot of performance. So the benchmark is less about the languages involved, but rather the exact algorithms used and capabilities offered.
- Roboprog 10y agoGood point. But assume if you run a series of tests on the same file, it's in RAM. (discard the time of the first run or two, assume data source - network, DB, file - makes access time moot)
- masklinn 10y ago> Any CSV reader should be limited only be disk access right? Depends on the speed of your storage subsystem. If you're working from RAM or from a fast PCIe SSD, you'll probably bottleneck in the encoding validation and actual parsing.
- endymi0n 10y agoOn a related note, also the Go stdlib regex package is pretty naive and imperformant compared to a full blown and modern backtracking PCRE implementation (at 1/10 the LOC and complexity) - same thing goes for the reflection based JSON package (which is still kinda "fast enough"). The focus wasn't so much on performance but on initial completeness, good interface, versatility, clarity and simplicity - with faster or more specialized implementations left to the community. There might be different opinions about that, but I personally like the approach of having a solid and ordered programming pocket knife - that also doesn't replace a Katana for cutting.
- dsymonds 10y agoThe standard regexp package, unlike PCRE, is actually a proper regular expression parser/matcher. Anything doing backtracking is at risk of exponential blowup and isn't safe. https://swtch.com/~rsc/regexp/regexp1.html https://swtch.com/~rsc/regexp/regexp1.html
- paulddraper 10y agoOn differing opinions: I prefer having a Katana, an automatic shotgun, and a spy drone, and leaving the pocket knife at home. In other words....C/C++ :) I just try not to blow off my leg.
- deno 10y agoBetter than node.js import * as csv from 'csv-parse'; import * as fs from 'fs'; type Line = [string,string,string,string,string,string]; const parser = new csv.Parser({}); parser.on('data', (line: Line) => { if (line[0] === '42') { console.dir(line); } }); fs.createReadStream('mock_data.csv').pipe(parser); $ /usr/bin/time node parse_csv.js 43.61user 0.85system 0:45.61elapsed 97%CPU (0avgtext+0avgdata 60076maxresident)k $ node --version v6.4.0 Edit: Using fast-csv 24.28user 0.20system 0:24.58elapsed 99%CPU (0avgtext+0avgdata 91780maxresident)k
- masklinn 10y agocsv-parse is hardly the only CSV parser for node, and it is by far the slowest: https://github.com/phihag/csv-speedtest https://github.com/phihag/csv-speedtest (csv2json depends on csv-parse, so it's unsurprising that it's even slower)
- deno 10y agoI chose the most popular one on npm because Go and Python are using stdlib.
- elmigranto 10y agoBut you still wrote "faster than node.js" and not "faster that most popular npm module" (which aren't always of a great quality or performance-oriented).
- deno 10y agoYup[1]. Sorry! [1] https://meta.wikimedia.org/wiki/Cunningham%27s_Law https://meta.wikimedia.org/wiki/Cunningham%27s_Law
- deno 10y agoThe fastest streaming example on that list (csv-parser) is still 5× slower than Python.
- twotwotwo 10y agoAs someone notes on the bug, if you were rolling your own, there are some other things you could do--return a [][]byte that's a pointer to its internal buffer, only usable until the next row is read. Making a version of encoding/csv that retains most of its features (custom delimiters, handling backslashes and quoting and \r) but streams like that would be a fun open source project for someone who likes Making Things Go Fast.
- peterwaller 10y agoI did this a couple of months ago and got a >5x speedup. It's at the expense of dropping quoting though, so no commas or newlines can be in the input data. https://github.com/pwaller/usv https://github.com/pwaller/usv
- weberc2 10y agoThat guy here. I'm also interested in rolling a version that only supports standard delimiters so I can forego rune parsing. Rune parsing accounts for about 30% of the processing; not sure how much a bytes implementation could save, but I'm hopeful.
- lcarlson 10y agoThis may be a bit off topic but I've found sqlite to be quite a powerful csv parser. Once posted you can manipulate the data in lots of ways. When you're working with reports that need to get back into some sort of table format, it's very intuitive and easy for SQL people.
- WestCoastJustin 10y agoRelated to this, was a Reddit thread from a few days ago in /r/golang about improving a csv/reader. See: https://www.reddit.com/r/golang/comments/50ncer/implementing_a_streaming_csv_reader_reduced/ https://www.reddit.com/r/golang/comments/50ncer/implementing...
- Roboprog 10y agoSo, where's the Perl (+ CPAN...) version for comparison? I guess "Nobody cares about your dead religion" :-) Still, it would probably be faster, at the expense of being unreadable.
- burntsushi 10y agoOn the CSV Game benchmark, the Go csv reader is around 10x slower than the fastest: https://bitbucket.org/ewanhiggs/csv-game https://bitbucket.org/ewanhiggs/csv-game