5 ms·
Colin's focus has been on speeding up EC2 boot time. You pay per second from time on EC2. A few milliseconds at scale probably ads up to a decent amount of sa
by count 3y ago
Colin's focus has been on speeding up EC2 boot time. You pay per second from time on EC2. A few milliseconds at scale probably ads up to a decent amount of savings - easily imaginable it's enough to cover the work it took to find the time savings.
- taeric 3y agoYes, you pay per second. But, per other notes, this is slated to save at most 2ms. Realistically, it won't save that much, as they are still doing a sort, not completely skipping it, but lets assume qsort does get this down to 0 and they somehow saved it all. That is trivially 500 boots before you see a second. Which... is a lot of reboots. If you have processes that are spawning off 500 instances, you would almost certainly see better savings by condensing those down to fewer instances. So, realistically, assuming that Colin is a decently paid software engineer, I find it doubtable the savings from this particular change will ever add up to mean anything even close to what it costs to have just one person assigned to it. Now, every other change they found up to this point may help sway that needle, but at this point, though 7% is a lot of percent, they are well into the area where savings they are finding are very unlikely to ever be worth finding. Edit: I saw that qsort does indeed get it into basically zero at this range of discussion. I'd love to see more of the concrete numbers.
- vidarh 3y agoMissing the point, which is to shave off cold start times for things like Lambdas, where shorter start times means you can more aggressively shut them down when idle, which means you can pack more onto the same machine.
- taeric 3y agoSee, the binpacking that this implies for lambda instances is insane to me. Certainly sounds impressive and great, and they may even pull it off. And, certainly, for the Lambda team, this at least sounds like it makes sense. I'll even root for them to hope it works out. It just reminds me of several failed attempts to move to JVM based technologies to replace perl, because "just think of the benefits if we get JVM startup so that each request is its own JVM?" They even had papers showing resource advantages that you would gain. Turned out, many times, that competing against how cheap perl processes were was not something you wanted to do. At least, not if you wanted to succeed.
- vidarh 3y agoA realization here is that while reaching cold start times that allow for individual requests would be awesome, you don't need to reach that to benefit from attacking the cold start time: Every millisecond you shave off means you can afford to run closer to max capacity before scaling up while maintaining whatever you see as the acceptable level of risk that users will sometimes wait. Of course if the customer stack is slow to start, the lower bound you can get to might still be awful, but you can only address the part of the stack you control.
- insanitybit 3y agoI think your argument is effectively "Colin's time could be better spent". That might be true, but consider that: 1. He was already profiling the system, which is almost certainly where the bulk of his eng costs are 2. This is one piece of a larger optimization story, he has made many many changes to significantly improve performance 3. This is a trivial change 4. 500 Firecracker executions is nothing - FaaS targets 10s of thousands of executions per second Given (1) and (3) in particular the only sensible thing Colin could do other than fix this is ignore it, which seems insane.
- taeric 3y agoI think my argument is easily seen as this, but that is not my intent of "argumemt." As you say, and if I'm not mistaken, in this case they only "cared" about this because it was easy and "in the way," as it were. I don't think it was a waste of time to pick up the penny in the road as you are already there and walking. What I am talking towards, almost certainly needlessly, is that it would be a waste of most people's time to profile anything that is already down to ms timing in the hopes of finding a speedup. In this case, it sounds like they were basically doing final touches on some speedups they already did and sanity testing/questioning the results. In doing so, they saw a few last changes they can make. So, to summarize, I do not at all intend this as a criticism of this particular change. Kudos to the team for all the work they did to get so that this was 7% of the remaining time, and might as well pick up the extra gains while there.
- greggyb 3y agoI've spent plenty of time optimizing for milliseconds (for very high hourly rates, not on staff, so my work has a very clear cost and ROI) and I operate probably at least a dozen layers above the OS. Modern CPUs execute a few million instructions per millisecond. I think people who operate at a level where microsecond and nanoseconds matter may see it as dismissive to question gains of milliseconds.
- taeric 3y agoI'm sure that happens. And I'm sure they do. I'm also equally sure in thinking that the vast majority of the interactions I have, are with people that are not doing this. You can think of my "argument" here, as reminding people that shaving 200grams off your bicycle is almost certainly not worth it for anyone that is casually reading this forum. Yes, there are people for whom that is not the case. At large, that number of people are countable. And I can't remember if it was this thread or another, but I didn't really intend an argument. Just a conversation. I thought I lead off with a kudos on the improvement. Criticism would be to the headline, if there is any real criticism to care about.
- esprehn 3y agoI think the focus here was for firecracker VMs used by lambda? If you're paying 2ms on every invocation of every function that'll add up. OTOH it seems like SnapStart is a more general fix.
- count 3y agoJust to be clear, I mean 'enough saved across all pay-per-second ec2 users', not a given specific account or user, which, yeah, it's probably minimal. The scale of lambda and ec2 is...enough to make any small change a very large number in aggregate.
- taeric 3y agoRight, but this is akin to summing up all of the time saved by every typist learning to type an extra 5 words a minute. Certainly you can frame this in such a way that it is impressive, but will anyone notice?
- count 3y agoHaha, there's a thread here on hacker news with tons of comments, so, yeah, somebody noticed! :)