4 ms·
This is reasonably idiomatic Python and 10x faster than the implementation in the original post: with open("orthocoronavirinae.fasta") as f: text = ''.
by soundmasterj 5y ago
This is reasonably idiomatic Python and 10x faster than the implementation in the original post:
with open("orthocoronavirinae.fasta") as f:
text = ''.join((line.rstrip() for line in f.readlines() if not line.startswith('>')))
gc = text.count('G') + text.count('C')
total = len(text)
Or if you want to be explicit, this is just as fast (and might scale better for particularly long genomes):
gc = 0
total = 0
with open("orthocoronavirinae.fasta") as f:
for line in f.readlines():
if not line.startswith('>'):
line = line.rstrip()
gc += line.count('C') + line.count('G')
total += len(line)
I didn't test Nim but the author reports Nim is 30x faster than his Python implementation, so mine would be about 3x slower than his Nim.
- epidemian 5y agoI think this is missing the point of the article. Yes, you can implement a faster Python version, but notice also: * This faster version is reading all the file into memory (except comment lines). The article mentions the data being 150MB, which should fit in memory, but for larger datasets, this approach would be unfeasible * The faster version is actually delegating a lot of work to Python's C internals by using text.count('G'). All the internal looping and comparisons is done in C, while on the original version, goes through Python So yes, you can definitely write faster Python by delegating most of the work to C. The point of the article is not about how to optimize Python, but about how given almost identical implementations in Python and Nim, Nim can outperform Python by 1 or 2 orders of magnitude without resorting to use C internals for basic things like looping or comparing characters.
- soundmasterj 5y agoI didn't try to write optimized code, but idiomatic Python. Which also happens to be 10x faster. To make it streaming, take the second version and remove the readlines (directly iterate over f). Delegating work to Python's C internals is fine IMO because "batteries included" is a key feature of Python. "Nim outperforms unidiomatic Python that deliberately ignores key language features" is perhaps true, but less flashy of a headline. And to be honest, I mainly wrote this because the other top level Python implementations for this one were terrible at the time of the post.
- user5994461 5y agoOne liner to count gc, without buffering. import io f = io.StringIO( """ AB CD EF GH """ ) total = sum(map(lambda s: 0 if s[0]==">" else s.count('G') + s.count('C'), f.readlines())) print(total)
- user5994461 5y agoAnd reading the file as binary. There's a lesson about the overhead of unicode strings here ;) Your first example takes 3.1 seconds, my previous comment takes 2.3 seconds, this one takes 1.4 seconds. start = time.perf_counter() with open("orthocoronavirinae.fasta", "rb") as f: total = sum(map(lambda s: 0 if s[0]==65 else s.count(b"G") + s.count(b"C"), f.readlines())) end = time.perf_counter() print(total, " total") print(end-start, " seconds")
- soundmasterj 5y agoThis is good too. I would use a generator expression instead of the map though probably.
- user5994461 5y agomap is a generator