4 ms·
50ms doesn't sound especially fast. If you had a million of them to process thats 4 hours, maybe an hour with some parallelism? A 100x speedup from going native
by Ultimatt 4y ago
50ms doesn't sound especially fast. If you had a million of them to process thats 4 hours, maybe an hour with some parallelism? A 100x speedup from going native then becomes quite welcome. Python has its place, but the fact you even think the tens of ms domain is "fast" or even fast enough on modern CPUs shows the real strength of Python, which is most people dont actually care at all about performance. Thats not to say its performant though. Just that no one cares anymore. Some rando boss is happy to wait 4 hours for your script to chug through a million statements, because no one actually told them it can take 2.5 minutes in a native language. If they did maybe the boss would suddenly speak about their dream to not only process the data but have a near real-time BI panel instead of a batch report so they can react within one business day. The issue with Python is missed value, not the value it can deliver.
- squeaky-clean 4y agoWhy limit parallelism to 4x? Spin up a ton of lambdas and get it done in 10 seconds. You're also forgetting that 50ms time to load each file from cloud storage. It's still 4 hours and 2.5 minutes with your native code compared to 8 hours with Python. Suddenly not such a massive improvement. Even then I don't see the issue with it taking 4 hours to do a million of them. Do you need to do a million per hour? Is it even likely they'll need to do a million of them total? Do bank statements even come in frequently enough to do a realtime dashboard with? I get my bank statement every 30 days. How much longer does it take to develop the native version? How much longer does it take to modify when a bug is found or a bank changes their statement layout? How much more do you have to pay a native code engineer compared to a Python dev and how easy are they going to be to replace eventually?
- blindhippo 4y agoOne dimension to consider here is cost of compute. Going from 8 hours to 4 hours is a 50% reduction in the time to compute and we're assuming this is occurring on the same relative hardware/instance size. At scale, that could translate into hundreds of thousands of dollars in savings. But your points are relevant - as with anything related to development, "it depends" rules the day. There isn't a clear cut "x is objectively better than y" in general.
- stevesimmons 4y ago> The issue with Python is missed value, not the value it can deliver. I'm very aware of the business value here! The limiting factor in these types of "messy real-world data" problems is the developer time to get 100+ templates right on all the different variations encountered in the wild. I can iterate on each template extremely effectively in a Jupyter notebook REPL, and immediately rerun a sample of 100 statements for that bank in a few seconds. While the total corpus of statements I have access to is actually around a million, no one cares how quick processing them all are if the extraction isn't reliable enough!
- aidos 4y agoExactly. The time spent developing a solution for your problem in pretty much any other language is going to cost magnitudes more than processing millions of documents in python. And that’s ignoring the ongoing maintenance where the hackability of teasing data out of PDFs in ipython is going to top any other system. I have a soft spot for the work you’re doing since I’ve spent a good portion of my life now extracting data from PDFs and can appreciate the joys of the process more than most.
- justinsaccount 4y ago> If you had a million of them to process thats 4 hours If you had 20 bank accounts sending you monthly statements for 100 years you'd have 24,000 of them. 50ms each would take you 20 minutes to process the 100 year backlog. If you have a million of them to process then you're a bank or similar type of institution that can devote resources to this, either computational or developer time to optimize things. A 128 core c6a.32xlarge would turn the runtime from 4 hours into 1.9 minutes and cost $5 to run for an hour.
- tmtvl 4y agoA laptop generates about 27 grams of CO2 per hour. So going from 4 hours to 2.5 minutes would save about 0.1 kg of CO2 each time the app is ran. So for a company of 1,000 people working 48 weeks per year that's almost 5 tons of CO2 that could be saved. Ah well, 5 tons of CO2 is a fair trade off for being able to use Python, right?
- polygamous_bat 4y agoI understand where you're coming from, and I do care about the climate as much as the next person, but this comparison seems downright silly. Basically, for your example to work every single person have to run the same script every week and wait for 4 hours. Are there even that many bank statements? And then, if you, your spouse, and a child take a single flight coast to coast, round trip, that is 4.5 tons right there.
- mmcnl 4y agoI agree with your reasoning and your insight about missed value is valuable, but this is not a Python problem. It's a scaling problem. You can solve scaling problems in many different ways with different costs. Any reasonably sized company has a Kubernetes cluster, just launch 10 instances and you're down to 24 minutes instead of 4 hours. You could also buy faster hardware. You might increase single-node performance by rewriting in a performant language (this won't help much if I/O is the bottleneck though, which often is the case). Talking to your data supplier to check if they can send the data in a machine-readable format is probably also something to consider. Or you might conclude that any effort is not worthwhile at all. Who knows? Scaling problems are not be avoided as a negative thing, they're beautiful and ought to be celebrated. What's not beautiful is letting the thought of potential scaling problems lead you to make decisions that optimize for long-term results instead of short-terms results.