47 ms·
The computers are fast, but you don't know it
- FpUser 4y ago"...but you do not know it" Believe me I do. This is why my backends are single file native C++ with no Docker/VM/etc. The performance on decent hardware (dedicated servers rented from OVH/Hetzner/Selfhost) is nothing short of amazing.
- user_7832 4y agoI wonder how much power (and resulting CO2 emissions) could be saved if all code had to go through such optimization. And on a slightly ranty note, Apple's A12z and A14 are still apparently "too weak" to run multiple windows simultaneously :)
- MR4D 4y agoThat’s a ram issue not a processor issue. At least, that’s according to Apple https://appleinsider.com/articles/22/06/11/stage-manager-for-ipados-16-limited-to-m1-over-memory-storage-speed-requirements https://appleinsider.com/articles/22/06/11/stage-manager-for...
- aaaaaaaaaaab 4y agoThat's the biggest bullshit ever.
- lynguist 4y agoThe funny thing is that actual Macs with 4GB RAM are also supported and they can run the entire desktop environment.
- dijit 4y agoWindowServer runs like a pig though, consistently the most expensive process (in terms of memory) on my mac systems.
- david38 4y agoWorse CO2 emissions. You think optimizations are energy free?
- TheDong 4y agoI think it depends and could be either worse or better depending. Some code is compiled more often than it is run, and some code is run more often than it's compiled. If you can spend 100k operations per compilation to save 50k operations at runtime on average... That'll probably be a net positive for chromium or glibc functions or linux syscalls, all of which end up being run by users more often than they are built by developers. If it's 100k operations at build-time to remove 50k operations from a test function only hit by CI, then yeah, you'll be in the hole 50k operations per CI run. All of this ignores the human cost; I don't really want to try (and fail) to approximate the CO2 emissions of converting coffee to performance optimizations.
- tomrod 4y agoMany can be pareto improvements from current state.
- hashingroll 4y agoNot all optimizations are more energy consuming. For an analogy, does a using a car consume more energy than a bicycle? Yes. But a using a bicycle does not consume more energy than a man running on feet.
- Maclennan 4y ago[dead]
- wodenokoto 4y agoDid the author beat pandas group an aggregate by using standard Python lists?
- morelisp 4y agoThe main optimization at that stage seems to be preallocating the weights. I don't know pandas but such a thing would have been possible without dropping any of the linalg libraries I do know how to use. I doubt the author's C++ implementations beat BLAS/LAPACK, but since they're not shown I can only guess. I've done stuff like this before but the tooling is really no fun, somewhere between 2 and 3 I'd just write it all in C++. Changing the interface just to get parallelism out seems not great - give it to the user for free if the array is long enough - but maybe it was more reasonable for the non-trivial real problem.
- BiteCode_dev 4y agoMost likely a missued of Pandas. DF are heavy to create, but calculations on them are fast if you stay in the numpy world and stay vectorized.
- jjoonathan 4y ago"It's fast so long as you don't use any of the many parts that aren't fast!" This isn't great.
- deleted 4y ago[deleted]
- BiteCode_dev 4y agoThat's true for everything in computing. Don't use a hammer as a screwdriver. I'm not even implying they shouldn't have used pandas for this, I'm suggesting they probably wrote the wrong pandas code for this. Pandas is typically 3 times faster than raw Python, not 10 times slower.
- 4y ago
- muziq 4y agoThis afternoon, discussing with my boss, why issuing two x 64 byte loads per cycle is pushing it; to the point where l1 says no.. 400GB of l1 bandwidth is all we have.. Is all we have.. I remember when we could move maybe 50KB/s.. Ans that was more than enough..
- jerf 4y agoI've been lightly banging the drum the last few years that a lot of programmers don't seem to understand how fast computers are, and often ship code that is just miserably slower than it needs to be, like the code in this article, because they simply don't realize that their code ought to be much, much faster. There's still a lot of very early-2000s ideas of how fast computers are floating around. I've wondered how much of it is the still-extensive use of dynamic scripting languages and programmers not understanding just how much performance you can throw away how quickly with those things. It isn't even just the slowdown you get just from using one at all; it's really easy to pile on several layers of indirection without really noticing it. And in the end, the code seems to run "fast enough" and nobody involved really notices that what is running in 750ms really ought to run in something more like 200us. I have a hard time using (pure) Python anymore for any task that speed is even remotely a consideration for anymore. Not only is it slow even at the best of times, but so many of its features beg you to slow down even more without thinking about it.
- agumonkey 4y agoRemember when people were counting cpu cycles and instruction size to ensure performance ?
- SV_BubbleTime 4y agoI program for embedded… still do that.
- agumonkey 4y agofind people like you, form a club, write articles and enjoy .. 1 reader :)
- Yajirobe 4y agoPython allows one to save development time in exchange for execution time
- morelisp 4y ago
- vjerancrnjak 4y agoHmm, interesting that single threaded C++ is 25% of Python exec time. It feels like C++ implementation might have area for improvement. My usual 1-to-1 translations result in C++ being 1-5% of Python exec time, even on combinatorial stuff.
- bbojan 4y agoI recently ported some very simple combinatorial code from Python to Rust. I was expecting around 100x speed up. I was surprised when the code ended running only 14 times faster.
- SekstiNi 4y agoJust to be sure, you did compile the Rust program using the --release flag?
- bbojan 4y agoYup!
- remuskaos 4y agoDid you use python specific functions like list comprehensions, or "classic" for/while loops? Because I've found the former to be surprisingly fast, while naive for loops are incredibly slow in python.
- bbojan 4y agoI've used list comprehensions where they make sense. But it's a non-trivial (although a simple) program, so there are for loops too. So I'm using a mix of chain and permutations from itertools, list comprehensions and for loops.
- momojo 4y agoAren't most of the python primitives implemented in C?
- 4y ago
- eterm 4y agoIt's hard to evaluate this article without seeing the detail of the "algorithm_wizardry", there's no detail here just where it would be interesting.
- geph2021 4y agoThe author says: "The function looks something like this:" And then shows some grouping and sorting functions using pandas. Then he says: "I replaced Pandas with simple python lists and implemented the algorithm manually to do the group-by and sort." I think the point of the first optimization is you can do the relatively expenseive group/sort operations without pandas, and improve performance. For the rest of the article it's just "algorithm_wizardry", which no longer deals with that portion of the code.
- eterm 4y agoWe never get a good sense of how much time was actually saved with that change not least because the original function calls "initialise weights" inside every loop, the new function does not. It would have been interesting to see what difference that alone made. The takeaway of the article, that computers are blindingly fast and we make them do unecessary work (and often sit around waiting on I/O) with most their time is true of course. I'm currently writing a utility to do a basic benchmark of data structures and I/O and it's been a real learning experience for me in just how fast computers can be, but also just how slow a little bit of overhead or contention can cause things, but that's better left for a full write up another day.
- geph2021 4y agoWe never get a good sense of how much time was actually saved with that change not least because the original function calls "initialise weights" inside every loop, the new function does not. Good point. Furthermore to your point, I would assume a library like pandas has fairly well optimized group and sort operations. It would not occur to me that pandas is the bottleneck, but the author does clarify in his footnote that pandas operations, by virtue of creating more complex pandas objects, can indeed be a bottleneck. [1] Please don't get me wrong. Pandas is pretty fast for a typical dataset but it's not the processing that slows down pandas in my case. It's the creation of Pandas objects itself which can be slow. If your service needs to respond in less than 500ms, then you will feel the effect of each line of Pandas code.
- porcoda 4y agoYup. We have gotten into the habit of leaving a lot of potential performance on the floor in the interest of productivity/accessibility. What always amazes me is when I have to work with a person who only speaks Python or only speaks JS and is completely unaware of the actual performance potential of a system. I think a lot of people just accept the performance they get as normal even if they are doing things that take 1000x (or worse) the time and/or space than it could (even without heroic work).
- LAC-Tech 4y agoI don't think we can really blame slow languages. Implementations of languages like javascript, ruby - and I would presume python and php - are a lot faster than they used to be. I think most slowness is architectural.
- vladvasiliu 4y ago> I think a lot of people just accept the performance they get as normal even if they are doing things that take 1000x (or worse) the time and/or space than it could (even without heroic work). Habit is a very powerful force. Performance is somewhat abstract, as in "just throw more CPUs at it" / it works for me (on my top of the line PC). But people will happily keep on using unergonomic tools just because they've always done so. I work for a shop that's mainly Windows (but I'm a Linux guy). I won't even get into how annoying the OS is and how unnecessary, since we're mostly using web apps through Chrome. But pretty much all my colleagues have no issue with using VNC for remote administration of computers. It's so painful, it hurts to see them do it. And for some reason, they absolutely refuse to use RDP (I'm talking about local connections, over a controlled network). And they don't particularly need to see what the user in front of the computer is seeing, they just need to see that some random app starts or something. I won't even get into Windows Remote Management and controlling those systems from the comfort of their local terminal with 0 lag. But for some reason, "we've always done it this way" is stronger than the inconvenience through which they have to suffer every day.
- etaioinshrdlu 4y agoMy entire career, we never optimize code as well as we can, we optimize as well as we need to. Obviously the result is that computer performance is only "just okay" despite the hardware being capable of much more. This pattern repeats itself across the industry over decades without changing much.
- pjvsvsrtrxc 4y agoThe problem is that performance for most common tasks that people do (f.e. browsing the web, opening a word processor, hell even opening an IM app) has gone from "just okay" to "bad" over the past couple of decades despite our computers getting many times more powerful across every possible dimension (from instructions-per-clock to clock-rate to cache-size to memory-speed to memory-size to ...) For all this decreased performance, what new features do we have to show for it? Oh great, I can search my Start menu and my taskbar had a shiny gradient for a decade.
- etaioinshrdlu 4y agoI think a lot of this is actually somewhat misremembering how slow computers used to be. We used to use spinning hard disks, and we were so often waiting for them to open programs. Thinking about it some more, the iPhone and iPad actually comes to mind as devices that perform well and are practically always snappy.
- pjvsvsrtrxc 4y ago> I think a lot of this is actually somewhat misremembering how slow computers used to be Suffice to say: I wish. I have a decently powerful computer now, but that only happened a few years ago. > We used to use spinning hard disks, and we were so often waiting for them to open programs. Indeed, SSDs are much faster than HDDs. That is part (but not all) of how computers have gotten faster. And yet we still wait just as long or longer for common applications to start up. > the iPhone and iPad actually comes to mind as devices that perform well and are practically always snappy Terribly written programs are perfectly common on iP* and can certainly be slow. But you're right, having a high-end device does make the bloat much less noticeable.
- forinti 4y agoOn a 3GHz CPU, one clock cycle is enough time for light to travel only 10cm. If you hold up a sign with, say, a multiplication, a CPU will produce the result before light reaches a person a few metres away.
- dragontamer 4y ago> If you hold up a sign with, say, a multiplication, a CPU will produce the result before light reaches a person a few metres away. The latency on multiplication (register input to register output) is 5-clock ticks, and many computers are 4GHz or 5GHz these days. 5-clock cycles at 5GHz is 1ns, which is 30-centimeters of light travel. If we include L1 cache read and L1 cache write, IIRC its 4 clock cycles for read + 4 more for the write. So 13 clock ticks, which is almost 70 centimeters. ------------ DDR4 read and L1 cache write will add 50 nanoseconds (~250 cycles) of delay, and we're up to 13 meters. And now you know why cache exists, otherwise computers will be waiting on DDR4 RAM all day, rather than doing work.
- moonchild 4y ago> The latency on multiplication (register input to register output) is 5-clock ticks 3 https://www.agner.org/optimize/instruction_tables.pdf https://www.agner.org/optimize/instruction_tables.pdf
- jnordwick 4y agoThose insrtruction latencies are in addition to the pipeline created latency. (They are actually the number of cycles added to the dependency chain specifically). The mult port has a small pipeline itself of 3 stages (that why 3 cycles latency). Intel has a 5 stage pipeline so the minimum latency is going to be 8 for just those two things.
- gpderetta 4y agoI don't understand what you are trying to say. The dependency chain length is what is normally intended as instruction latency. Also the pipeline length is certainly not 5 stages but more like 20-30.
- tiffanyh 4y agoNIM NIM should be part of the conversation. Typically, people trade slower compute time for faster development time. With NIM, you don’t need to make that trade-off. It allows you to develop in a high-level but get C like performance. I’m surprise its not more widely used.
- goodpoint 4y agoIt's written Nim, not NIM.
- dgan 4y agoat that point, almost anything compiled will be at least an order of magnitude faster than python
- Spivak 4y agoThe dichotomy between "compiled/interpreted" languages is completely meaningless at this point. You can argue that Python is compiled and Java is interpreted. I mean one of our deployment stages is to compile all our Python files. The thing that makes the difference isn't the compilation steps, it's how dynamic the language is and how much behind the scenes work has to be done per line and what tools the language gives you to express stronger guarantees that can be optimized (like __slots__ in Python).
- Too 4y agoNo, compiling JIT like Java/.NET/V8 becomes real cpu machine instructions on the fly. Compiling to .pyc files is still byte code that later the python interpreter will read in what is more or less a gigantic while(true)switch/case (operation_type). What you save with pyc compilation is just the text parsing of the source code. How dynamic the language is, does however affect how feasible it is to do JIT compilation and here python has done itself a big disfavor of simply being too dynamic. Some attempts at JITing it have been done before (pypy, ironpy), usually by nerfing the language a bit to get rid of the most dynamic parts - something that is probably for the better anyway.
- 4y ago
- newaccount2021 4y agowhere I work, every frontend dev has a 64gb ram/2tb ssd/multicore laptop to develop web pages...everything is lightning fast apparently!...so they never do performance engineering of any kind
- bcatanzaro 4y ago1. There's no real limit to how slow you can make code. So that means there can be surprising large speedups if you start from very slow code. 2. But, there is a real limit to the speed of a particular piece of code. You can try finding it with a roofline model, for example. This post didn't do that. So we don't know if 201ms is good for this benchmark. It could still be very slow.
- varispeed 4y agoI wonder if eventually there is going to be consideration for environment required when building software. For instance running unoptimised code can eat a lot of energy unnecessarily, which has an impact on carbon footprint. Do you think we are going to see regulation in this area akin to car emission bands? Even to an extent that some algorithms would be illegal to use when there are more optimal ways to perform a task? Like using BubbleSort when QuickSort would perform much better.
- jcelerier 4y ago> Do you think we are going to see regulation in this area akin to car emission bands? it has thankfully started: https://www.blauer-engel.de/en/productworld/resources-and-energy-efficient-software-products https://www.blauer-engel.de/en/productworld/resources-and-en... I think KDE's Okular has been one of the first certified software :-)
- deleted 4y ago[deleted]
- kzrdude 4y agoWell, there is some rumbling about making proof of work cryptocurrencies illegal, and that falls under this topic. To some extent they can claim to deliver a unique feature where there is no replacement for the algorithm they are using.
- modeless 4y agoNow do it on the GPU. There's at least a factor of 10 more there. And a lot of things people think aren't possible with GPUs are actually possible.
- hamstergene 4y agoOn mobile devices it is more serious than just bad craftsmanship & hurt pride, bad code is short battery life. Think mobile game that could last 8 hours instead of 2 of it wasn’t doing unnecessary linear searches on timer in JavaScript.
- ben_w 4y agoThere was one place where a coworker had written a function that converted data from a proprietary format into a SQL database. On some data, this took 20 minutes on the test iPhone. The coworker swore blind it was as optimised as possible and could not possibly go faster, even though it didn't take that long to load from either the original file format or the database in normal use. By the next morning, I'd found it was doing an O(n^2) operation that, while probably sensible when the app had first been released, was now totally unnecessary and which I could safely remove. That alone reduced the 20 minutes to 200 milliseconds. (And this is despite that coworker repeatedly emphasising the importance of making the phone battery last as long as possible).
- xigoi 4y agoThis is why all programmers should know their asymptotic complexity.
- BeFlatXIII 4y agoAs the poster child of poor performance even on new iPhones, there is always Pokémon Go (produced with Niantic levels of competence)
- javajosh 4y agoThat's really cool but I somewhat resent the use of percentages here. Just use a straight factor or even better just the order of magnitude. In this case it's four orders of magnitude of an improvement.
- charlie0 4y agoI've always been tempted to make things fast, but for what I personally do on a day to day basis, it all lands under the category of premature optimization. I suspect this is the case for 90% of development out there. I will optimize, but only after the problem presents itself. Unfortunately, as devs, we need to provide "value to the business". This means cranking out features quickly rather than as performant as possible and leaving those optimization itches for later. I don't like it, but it is what it is.
- wildrhythms 4y agoI agree... for a business "fast" means shipping a feature quickly. I have personally seen the convos from upper management where they handwave away or even justify making the application slower or unusable for certain users (usually people in developing countries with crappy devices). Oh it will cost +500KB per page load, but we can ship it in 2 weeks? Sounds good!
- blub 4y agoLots of businesses have nearly zero engineering in them and cobble together libraries and frameworks that they sell or rent as software. On the other end of the spectrum you have companies hiring specialists at all points of the stack to squeeze out the last drops of performance, dedicated perf teams, etc. The latter also typically produce the tools that enable the former to function.
- MyNameIsFred 4y agoYou are completely correct and I really wish you weren't.
- momojo 4y ago> for what I personally do on a day to day basis, it all lands under the category of premature optimization Another perspective on premature opt: When my software tool is used for an hour in the middle of a 20-day data pipeline, most optimization becomes negligible unless it's saving time on the scale of hours. And even then, some of my coworkers just shrug and run the job over the weekend.
- dragontamer 4y agoAs a hobby, I still write Win32 programs (WTL framework). Its hilarious how quickly things work these days if you just used the 90s-era APIs. Its also fun to play with ControlSpy++ and see the dozens, maybe hundreds, of messages that your Win32 windows receive, and imagine all the function calls that occur in a short period of time (ie: moving your mouse cursor over a button and moving it around a bit).
- jansommer 4y agoWin32 is so really, really fast. And with Tiny C Compiler the program compiles and boots faster than the Win10 calculator app takes to start.
- pjvsvsrtrxc 4y agoLinux windows get just as many (run xev from a terminal and do the same thing). Our modern processors, even the crappiest Atoms and ARMs, are actually really, really fast.
- dragontamer 4y agoGPUs even faster. Vega64 can explore the entire 32-bit space roughly 1-thousand times per second. (4096 shaders, each handling a 32-bit number per clock tick, albeit requiring 16384 threads to actually utilize all those shaders due to how the hardware works, at 1200 MHz) One of my toy programs was brute forcing all 32-bit constants looking for the maximum amount of "bit-avalanche" in my home-brew random number generators. It only takes a couple of seconds to run on the GPU, despite exhaustive searching and calculating the RNG across every possible 32-bit number and running statistics on the results.
- xupybd 4y agoYou also have to optimize for the constraints you have. If you're like me then development time is expensive. Is optimizing a function really the best use of that time? Sometimes yes, often no. Using Pandas in production might make sense if your production system only has a few users. Who cares if 3 people have to wait 20 minutes 4 times a year? But if you're public facing and speed equals user retention then no way can you be that slow.
- pjvsvsrtrxc 4y ago> If you're like me then development time is expensive. Is optimizing a function really the best use of that time? Sometimes yes, often no. Almost always yes, because software is almost always used many more times than it is written. Even if you doubled your dev time to only get a 5% increase of speed at runtime, that's usually worth it! (Of course, capitalism is really bad at dealing with externalities and it makes our society that much worse. But that's an argument against capitalism, not an argument against optimization.)
- wbsss4412 4y agoNitpick: software optimization isn’t an example of an externality. Externalities are costs/benefits that accrue to parties not involved in a transaction.
- pjvsvsrtrxc 4y agoYes, it is. > ex·ter·nal·i·ty: a side effect or consequence of an industrial or commercial activity that affects other parties without this being reflected in the cost of the goods or services involved The buyer is an "other part[y]" from the seller's (edit: or better yet, developer, who might just be contracted by the ultimate seller...) perspective, and performance is basically impossible to quantify, therefore price. Moreover, even if you want to limit externalities to being completely third-party... sure: Pollution. More electrical generation capacity needed.
- 4y ago
- liprais 4y agomost likely misused pandas / numpy,as long as you stay in numpy land,it is quite fast.
- aaaaaaaaaaab 4y agoDevelopers should be mandated to use artificially slow machines.
- april_22 4y agoToday, in many cases, it truly is about optimising algorithms instead of building faster machines. I overheard this quote recently: 'I'd rather have today's algorithms on an old computer, than a new computer with old algorithms'
- mg 4y agoGood example is this high performance Fizz Buzz challenge: https://codegolf.stackexchange.com/questions/215216/high-throughput-fizz-buzz https://codegolf.stackexchange.com/questions/215216/high-thr... An optimized assembler implementation is 500 times faster than a naive Python implementation. By the way, it is still missing a Javascript entry!
- abraxas 4y agoYep, many (especially younger) programmers don't get the "feel" for how fast things should run and as a result often "optimize" things horribly by either "scaling out" i.e. running things on clusters way larger than the problem justifies or putting queuing in front and dealing with the wait.
- Taywee 4y ago> It's crazy how fast pure C++ can be. We have reduced the time for the computation by ~119%! The pure C++ version is so fast, it finishes before you even start it!
- FirstLvR 4y agoThis is exactly what I was dealing last year, some particular costumer came to meeting with the idea developers has to be aware of making the code Inclusive and sustainable... We told them that we must set priorities on the performance and the literal result from the operation (a transaction development from an integration) Nothing really happened at the end but it's a funny history in the office
- physicsguy 4y agoThe return value on the function in C++ is of the wrong type :) I agree though. I used these tricks a lot in scientific computing. Go to the world outside and people are just unaware. With that said - there is a cost to introducing those tricks. Either in needing your team to learn new tools and techniques, maintaining the build process across different operating systems, etc. - Python extension modules on Windows for e.g. are still a PITA if you’re not able to use Conda.
- reedjosh 4y agoPython and Pandas are absolutely excellent until you notice you need performance. I say write everything in Python with Pandas until you notice something take 20 seconds. Then rewrite it with a more performant language or cython hooks. Developing features quickly is greatly aided by nice tools like Python and Pandas. And these tools make it easy to drop into something better when needed. Eat your cake and have it too!
- ineedasername 4y agoYes, there have been times that I have called linux command line utilities from python to process something rather do it in python.
- vlovich123 4y ago> extra_compile_args = ["-O3", "-ffast-math", "-march=native", "-fopenmp" ], > Some say -O3 flag is dangerous but that's how we roll No. O3 is fine. -ffast-math is dangerous.
- tomrod 4y agoWhy?
- xcdzvyn 4y agoIt reorders instructions in ways that are mathematically but not computationally equivalent (as is the nature of FP). This also breaks IEEE compliance.
- TheRealPomax 4y agoSome really good reasons: https://stackoverflow.com/a/22135559/740553 https://stackoverflow.com/a/22135559/740553 It basically assumes all maths is finite and defined, then ignores how floating point arithmetic actually works, optimizing based purely on "what the operations suggest should work if we wrote them on paper" (alongside using approximations of certain functions that are super fast, while also being guaranteed inaccurate)
- pointernil 4y agoMaybe it's been stated already by someone else here but I really hope that CO2 pricing on the major Cloud platforms will help with this. It boils down to resources used (like energy) and waste/CO2 generated. Software/System Developers using 'good enough' stacks/solutions are externalising costs for their own benefit. Making those externalities transparent will drive alot of the transformation needed.
- ineedasername 4y agoSlow Code Conjecture: inefficient code slows down computers incrementally such that any increase in computer power is offset by slower code. This is for normal computer tasks-- browser, desktop applications, UI. The exception to this seem to be tasks that were previously bottlenecked by HDD speeds which have been much improved by solid state disks. It amazes me, for example, that keeping a dozen miscellaneous tabs open in Chrome will eat roughly the same amount of idling CPU time as a dozen tabs did a decade ago, while RAM usage is 5-10x higher.
- andrewclunn 4y agoHow are we supposed to optimize coding languages, when the underlying hardware architecture keeps changing? I mean you don't write assembly anymore, you would right in the LLVM. Optimization was done because it was required. It will come back when complete commoditization of cpus occur. Enforcement of standards and consistent targets allow for high optimizations. Just see what people are able to do with outdated hardware in the demo and homebrew scene for old game consoles! We don't need better computers, but so long as we keep getting them, we will get unoptimized software, which will necessitate better computers. The vicious cycle of consumerism continues.
- pelorat 4y agoMaybe, stop using Python for anything but better shell scripts? Pretty sure it was invented to be a bash replacement.
- tintor 4y agoTLDR: How to optimize Python function? Use C++.
- julius_deane 4y agoComing up next: Want to ship a C++ project this quarter? Use Python.
- mulmboy 4y agoDoubtful that moving from vectorised pandas & numpy to vanilla python is faster unless the dataset is small (sub 1k values) or you haven't been mindful of access patterns (that is, you're bad at pandas & numpy)
- julius_deane 4y agobut how else do you get to the front page of hn?
- thrwyoilarticle 4y ago>Optimization 3: Writing your function in pure C++ >double score_array[]
- jiggawatts 4y agoSomething all architecture astronauts deploying microservices on Kubernetes should try is benchmarking the latency of function calls. E.g.: call a "ping" function that does no computation using different styles. In-process function call. In-process virtual ("abstract") function. Cross-process RPC call in the same operating system. Cross-VM call on the same box (2 VMs on the same host). Remote call across a network switch. Remote call across a firewall and a load balancer. Remote call across the above, but with HTTPS and JSON encoding. Same as above, but across Availability Zones. In my tests these scenarios have a performance range of about 1 million from the fastest to slowest. Languages like C++ and Rust will inline most local calls, but even when that's not possible overhead is typically less than 10 CPU clocks, or about 3 nanoseconds. Remote calls in the typical case start at around 1.5 milliseconds and HTTPS+JSON and intermediate hops like firewalls or layer-7 load balancers can blow this out to 3+ milliseconds surprisingly easily. To put it another way, a synchronous/sequential stream of remote RPC calls in the typical case can only provide about 300-600 calls per second to a function that does nothing. Performance only goes downhill from here if the function does more work, or calls other remote functions. Yet, every enterprise architecture you will ever see, without exception has layers and layers, hop upon hop, and everything is HTTPS and JSON as far as the eye can see. I see K8s architectures growing side-cars, envoys, and proxies like mushrooms, and then having all of that go across external L7 proxies ("ingress"), multiple firewall hops, web application firewalls, etc...
- taeric 4y agoBeing fair, for many of the things that it is worth using a microservice for, you should already have some sort of dominant factor to the call that would more than justify the added latency of the remote call. Be it a database read/write or some other heavy calculation. Granted, this is exacerbated when architectures don't make a good division between control/compute/data planes. Control plane, which is exposed to users, should almost certainly be limited to a single (or handful, at most) microservice calls. Preferably to the fastest storage mechanism that you have, such that what latency it does add is minimized entirely.
- nickjj 4y agoI think folks often make trade offs with their working requirements. If you provide an end result response from your web app to a user's browser in 50ms-100ms (before external latency) then things like 200 microseconds vs 4 milliseconds have less of a meaningful difference. If your app makes a couple of internal service calls (over HTTP inside of the same Kubernetes cluster) it's not breaking the bank in terms of performance even if you're using "slow" frameworks like Rails and get a few million requests a month. I'm not defending microservices and using Kubernetes for everything but I could see how people don't end up choosing raw performance over everything. Personally my preference is to keep things as a monolith until you can't and in a lot of cases the time never comes to break it up for a large class of web apps. I also really like the idea of getting performance wins when I can (creating good indexes, caching as needed, going the extra mile to ensure a hot code path is efficient, generally avoiding slow things when I have a hunch it'll be slow, etc.) but I wouldn't choose a different language based only on execution speed for most of the web apps I build.
- quickthrower2 4y agoIt is fitting that it is hosted on bearblog.dev, which produces fast minimal sites. It is my favorite blogging platform so far.
- fullstackchris 4y agoAnd if you wrote your instructions in assembly, it would be even faster! /s Sorry for the rude sarcasm, but isn't this a post truly just about the efficiency pitfalls of Python? (or any language / framework choice for that matter) Of course modern computers are lightning fast. The overhead of every language, framework, and tool will add significant additional compute however, reducing this lightning speed more and more with each complex abstraction level. I don't know, I guess I'm just surprised this post is so popular, this stuff seems quite obvious.
- lucidguppy 4y agoIf your fast language is talking to a database how fast will your language be?
- dqpb 4y agoNetsuite takes upwards of 10 seconds to transition from one blank page to another.
- w0mbat 4y agoArticle says at one point, "We have reduced the time for the computation by ~119%!", which is impossible. If you reduce it by 100% it is taking zero time already.
- nerdbaggy 4y agoI always get confused by stuff with this. 100% would actually be 50% in the context you are thinking. https://math.stackexchange.com/a/1404242 https://math.stackexchange.com/a/1404242
- Panzer04 4y ago"faster" vs "reduced time". Many people confuse rate of work with reduction in time, and it's exceptionally annoying :(
- pindab0ter 4y agoBonus points if the 'speed is faster'.
- necovek 4y agoThat one might not be strictly correct (speed is greater), but it's at least non-ambiguous and understandable. I love me some of those "discount -50%" signs though.
- hinkley 4y agoOh, but we let people say "acceleration is faster". It's like we've reserved 'faster' for a single derivative and banned it for all the others.
- pindab0ter 4y agoI don't think it's about banning words at all. It's about words making sense. "What's cheaper? The price is." Now that just doesn't make any sense, since a price isn't cheap or expensive, it's high or low. The thing that is priced can be cheap or expensive, but that's not what's being said. "What's faster? The speed is." Doesn't make sense either. Speed isn't fast, the speedy thing is. However, "What's faster? The acceleration is." is fine, because you can have slow or fast acceleration (I think?). I'm an ESL speaker, so please do tell me if I'm wrong and how.
- streamlining 4y agoFor years I stuck with MATE, Xfce4, LXQT, etc. to get optimal performance on old hardware but nothing can top a tiling window manager. With Nixos I switch between Gnome 40 (I do like the Gnome workflow) and i3 w/ some Xfce4 packages, but lately on my older machine the performance of Gnome (especially while running Firefox) is so sluggish in comparison that I may have switched back permanently now.
- skohan 4y agoI remember the moment I realized how fast computers are at uni. I was in an algorithms course, and one of our projects was to make a program which would read in the entire dataset from IMDB of films and actors, and calculate the shortest path between any actor and Kevin Bacon using actors and movies as nodes and roles as edges. I was working in C, and looking back I came up with a quite performant solution mostly by accident: all the memory allocated up front in a very cache-friendly way. The first time I ran the program, it finished in a couple seconds. I was sure something must have failed, so I looked at the output to try to find the error, but to my surprise it was totally correct. I added some debug statements to check that all the data was indeed being read, and it was working totally as expected. I think before then I had a mental model of a little person inside the CPU looking over each line of code and dutifully executing it, and that was a real eye-opener about how computers actually work.
- deleted 4y ago[deleted]
- ano88888 4y agoWhy can't we have a language easy to read and maintain but also have the speed of C?
- killingtime74 4y agoTranspile to C? Zig or Rust?
- skohan 4y agoYeah I think Zig in particular is trying to be just exactly a "better C". I don't know if transpiling will get you there, because for instance if you're transpiling a dynamic language, you're going to have to output C that is essentially emulating all those dynamic language features, so it might be faster than say, the original Python, but it's not going to be as fast as a pure C implementation.
- bruce343434 4y ago
- hermitcrab 4y agoI have written some data wrangling software in pure C++. I would like to benchmark it again Pandas to see how the speed compares. Does anyone know if there is a good set of Pandas benchmarks that I can create a comparison to? Even better if it has an R comparison.
- hnhn 4y agoThe data.table package in R often produces benchmarks that include pandas, e.g. https://h2oai.github.io/db-benchmark/ https://h2oai.github.io/db-benchmark/
- hermitcrab 4y agoThanks.
- tonto 4y agoAlso fun: test your intuition on the speed of basic operations https://computers-are-fast.github.io/ https://computers-are-fast.github.io/
- deleted 4y ago[deleted]
- justsomeuser 4y agoI imagine for most web dev’s using a fast memory unsafe language is like taking a bullet train to the local shop to get milk.
- Gigachad 4y agoIt also doesn’t stop when you reach your destination so you have to jump and roll out. Get it wrong and you die. Questioning this method is widely frowned on.
- hoseja 4y agoThe alternative is crawling around with your tongue and circling the shop hundred times before coming in. So intuitive!
- floucky 4y agoI would say to the local farm, then you have to wait for the cow to be milked (like an external api call...). At the end you just reduced you journey time by 0.1%, and incread you code complexity by 100%.
- inkblotuniverse 4y agoOn the other hand, to non-webdevs, webdevs are like an obese american woman trundling down the road on her mobility scooter, giving the evil eye to people overtaking her on foot as she takes a bite out of a hunk of RAMcheese.
- jwozn 4y agoAs an overweight-midwestern-american, RAMcheese sounds great!
- bawolff 4y agoIt feels like the gist of this article is just, don't use python.
- journey_16162 4y agoAs a front-end developer, I can't help but notice how much useless computation is going on in a fairly popular library - Redux. It's a store of items, if just one tiny items change in the whole store, every subscriber of every item gets notified and a compare function is ran to check if it changes. Perhaps I'm misunderstanding something and not to bash on Redux - I'm sure there are well-deserved reasons it got popular, but to me that just sounds insane and the fact that it got so much widespread adoption perfectly reflects how little care about performance is given nowadays. I don't use a high-end laptop and I'm not eager to upgrade is because I can relate to the average user of the software I develop. I saw plenty of popular web apps feeling really sluggish.
- korla 4y agoI think developer speed is more important than optimising clock cycles unnecessarily. Generally writing to dom is much much slower than evaluting a few thousand expressions. For the cases when it's not, use memo.
- xaedes 4y ago> I think developer speed is more important than optimising clock cycles unnecessarily. Developer time is spent once. Users will always have to pay the price of additional run time. For. Each. Single. User. Always. It scales! Due to the scale of, e.g. slow front-ends, with millions of users, this takes a HUGE amount of time. Only to save a few hours or days to develop it better. Having 1 million users each wait a single second is already 11 days. If they have to wait that single second for each interaction, it quickly adds up. It is also bad for the environment due to scaled up inefficiency and resulting increase of power usage.
- doctor_eval 4y agoAlthough I 100% agree with you, the problem is that these costs don't affect the original developer; it is an externality; a lot like carbon pollution. It's cheaper for the organisation to optimise for developer speed, even if the cost of that is borne by all the users.
- Someone 4y agoFTA: Note that the output of this function needs to be computed in less than 500ms for it to even make it to production. I was asked to optimize it. […] Took ~8 seconds to do 1000 calls. Not good at all :( Isn’t that 8ms per call, way faster than the target performance? Or should that “500ms” be “*500 μs”?
- julius_deane 4y agopercentages in the post are wrong too no surprise pandas was "slow"
- Havoc 4y agoYes in general for me the limitation is now my ability / Knowledge. Every cloud / SaaS is throwing free tier compute capacity at people and it’s just overwhelming (in a good way I suppose)
- thanzex 4y agoIf anything this is a testament to how slow python can be, and most importantly how easily it pushes you to write miserably unoptimized code. It could be a bit overkill, but whenever I'm writing code on top of optimizing data structures and memory allocations I always try to minimize the use of if statements to reduce the possibility of branch prediction errors. Seeing woefully unoptimized python code being used in a production environment just breaks my heart.
- baobob 4y agothe CPU branch predictor is so many levels down it will have almost no discernible effect on anything you might call a branch in Python code. Even a statement like "a = 1" likely executes a few tens if not a few hundred branches That is not to say aiming for generally unbranchy code is not a good thing - that often implies well designed code and well chosen data structures anyway
- Ultimatt 4y agoThe missing comparison is to Numpy and Numba as the first optimisation post pandas... I suspect nothing else there would beat it.
- tpoacher 4y agoThe point about pandas resonates with me. Don't get me wrong, pandas is a nice library ... but the odd thing is, numpy already has, like, 99% of that functionality built in in the form of structured arrays and records, is super-optimised under the hood, and it's just that nobody uses it or knows anything about it. Most people will have never heard of it. To me pandas seems to be the sort of library that because popular because it mimics the interface of a popular library from another language that people wanted to migrate to (namely dataframes from R), but that's about it. Compounding this, is that, it is now becoming an effective library to do things, even if backward, because the network effect means that people are building stuff to work on top of pandas, rather than on top of numpy. The only times I've had to use pandas in my personal projects was either: a) when I needed a library that 'used pandas rather than numpy' to hijack a function I couldn't care writing by myself (most recently seaborn heatmaps, and exponentially weighted averages - both relatively trivial things to do with pure numpy, and probably faster, but, eh. Leftpad mentality etc ...) b) when I knew I'd have to share the code with people who would then be looking for the pandas stuff. I'm probably wrong, but ...
- sgillen 4y agoIronically as a Phd in an ECE department, almost everyone has heard of and uses numpy, but many people have never even heard of pandas!
- kortex 4y ago> numpy already has, like, 99% of that functionality built in in the form of structured arrays and records Respectfully, this is pretty wrong. Pandas does vastly more out of the box than numpy. Off the top of my head: I/O from over a dozen of data formats, joins/merges, sql queries directly to dataframes, sql-like queries on dataframes, index slicing by time, multi-indexes, much more ergonomic grouping/aggregation functions, ergonomic wrappers around common graphing use-cases, rolling windows. I'm not even really a power user of it, so there's probably a zillion more things it does that numpy can't out of the box, and I don't wanna spend time writing time and validating if an implementation exists.
- 4y ago
- avianes 4y agoWhen I have to explain the speed of a processor to a neophyte I always begin by avoiding using GHz unit which has the weakness of hiding the magnitude of the number, so I explain things in terms of billions of cycles each second. As an example, with an ILP ~4 instruction/cycle at 5GHz we get 20 billion instructions executed each second in a single core. This number is not really tangible but it shocks
- wdroz 4y agoIf you are unhappy with pandas, give a try to polars[0] it's so fast! [0] -- https://www.pola.rs/ https://www.pola.rs/
- pdimitar 4y agoWhile I find this comment section fascinating and will read it top to bottom, I can't help but make an observation that such articles often comply with: +-------------------------------------------------+ | People really do love Python to death, do they? | +-------------------------------------------------+ I find that extremely weird. As a bystander who never relied on Python for anything important, and as a person who regularly had to wrestle with it and tried to use it several times, the language is non-intuitive in terms of syntax, ecosystem, package management, different language version management, probably 10+ ways to install dependencies by now, subpar standard library and an absolute cosmic-wide Wild West state of things in general. Not to mention people keep making command-line tools with it, ignoring the fact that it often takes 0.3 seconds to even boot. Why would a programmer that wants semi-predictable productivity choose Python today (or even 10 years ago) remains a mystery to me. (Example: I don't like Go that much but it seems to do everything that Python does, and better.) Can somebody chime in and give me something better than "I got taught Python in university and never moved on since" or "it pays the bills and I don't want to learn more"? And please don't give me the fabled "Python is good, you are just biased" crap. Python is, technically and factually and objectively, not that good at all. There are languages out there that do everything that it does much better, and some are pretty popular too (Go, Nim). I suppose it's the well-trodden path on integrating with pandas and numpy? Or is it a collective delusion and a self-feeding cycle of "we only ever hired for Python" from companies and "professors teach Python because it's all they know" from universities? Perhaps this is the most plausible explanation -- inertia. Maybe people just want to believe because they are scared they have to learn something else. I am interested in what people think about why is Python popular regardless of a lot of objective evidence that as a tech it's not impressive at all.
- jazzyjackson 4y agoI've started using micropython to interact with embedded arm chips, it's a revelation to interact with hardware through a REPL instead of compiling, transferring, resetting, and writing print statements to serial... This talk by the creator of micropython [0] gives his reasoning for why to implement python on microcontrollers despite it being hundreds of times slower than C. Starts @ 3:00 - it has nice features like list comprehension, generators, and good exception handling - it has a big, friendly, helpful community with lots of online learning resources - it has a shallow but long learning curve. It's easy to get started as a beginner, but you never get bored of the language, there's always more advanced features to learn. - it has native bitwise operations - has good distinction between ints and floats, and floats are arbitrary precision, you're not restricted to doubles or even long longs. (I'll add that built in complex numbers is a plus) - compiled language, so it can be optimized to improve performance [0] https://www.youtube.com/watch?v=EvGhPmPPzko https://www.youtube.com/watch?v=EvGhPmPPzko
- xvilka 4y agoAnd all they are used for to run slow and heavy Electron "apps" with three buttons.
- deleted 4y ago[deleted]
- deleted 4y ago[deleted]
- Zetaphor 4y agoFor some reason my employer is blocking this site as malware using Cisco's OpenDNS service
- Shorel 4y agoThe fact that now AWS CPU cost is a constant consideration in software development is making developers use better algorithms and languages, a trend that seems the opposite of the 2010s.
- djmips 4y agoBingo. It's not that software engineers are stupid, it's that they don't 'see' when they do something stupid and don't have a good mental model because of that lack of sight. Everyone figures out quickly to efficiently clean out their garage or other repetitive chores because it's personally painful to do it poorly and it's right in front of your nose. If only computers were more transparent and/or people learned and used profilers daily...
- illys 4y agoI am amazed by the discussions below on computer performance vs. software inefficiency: I remember the same discussions and arguments about software running on 8088 vs 80286 vs 80386 vs i486 vs Pentium... and so on. You could have had those discussion at anytime since the upgraded computers and microprocessors have become compatible with the previous generation (i.e. the x86 and PC lines). The point is that software efficiency measurement has never changed: it is human patience. The developers and their bosses decide the user can wait a reasonable time for the provided service. It is one-to-five seconds for non-real-time applications, it is often about a target framerate or refresh in 3D or real-time applications... The optimization stops when the target is met with current hardware, no matter how powerful it is. This measure drives the use of programming languages, libraries, data load... all getting heavier and heavier when more processing power gets available. And that will probably never change. Not sure about it? Just open your browser debugger on the Network tab and load the Google homepage (a field, a logo and 2 buttons). I just did: 2.2 MB, loaded in 2 seconds. It is sized for current hardware and 100 Mbps fiber, not for the actually provided service!