4 ms·
Heya, great to see you pop up here! I gotta be honest -- I think EC2 (and, in general, doing computation in units of VMware/Xen-style virtual PCs) is the actual
by keithwinstein 5y ago
Heya, great to see you pop up here! I gotta be honest -- I think EC2 (and, in general, doing computation in units of VMware/Xen-style virtual PCs) is the actual hipster compute substrate. AWS Lambda feels closer to cgi-bin from 1995, i.e. back when things still made sense. (Have you ever joined a tech company and been handed a 10 gigabyte VM image that uses a Vagrant pipeline to provision itself so you can get a working dev environment, except the pipeline only works if 100% of its 10,000 downloads succeed, so the whole thing is super-flaky, but nobody at the company knows because they only ran it once when they first joined and have just kept the same local dev VM ever since? That's what hipster compute means to me.)
All that aside, Cloudflare workers/fly.io/Fastly Compute@edge/Lucet/Google Cloud Run seem really cool, and the resulting work on Wasm and its ecosystem is fantastic, but they're also not exactly what excites me. Deploying code close to the edge (or "anywhere" in particular) isn't very important if the application only makes one round-trip. Even if my code is pure, it's not like fly.io is willing to sign a certificate saying, "We evaluated function <y> on input <z> and the correct answer is <x>, signed Fly Inc., and if you can prove us wrong in the next 10 years, our insurance company will pay you $1 million from our E&O policy." Which would really be cool. And, I don't know of people spinning up 4,000 nodes on those systems in 100 ms to do a 1-second-long computation. I haven't seen any of the providers or outsiders benchmarking the "burst-to-N,000-nodes latency" numbers averaged over many trials at various times of day. (We measured GKE a small number of times in the gg paper [fig. 7] and found it to be... really slow at that particular metric.)
I don't think we want to ignore locality! But I do want the OS to be able to secure access to thousands of cores in <1 second for <10 second duration workloads, and I think many applications would be willing to compromise on locality, or accept heterogeneous/irregular locality, in exchange for that. I'd still love visibility into the locality I end up with, I'd love not to have to do flaky NAT-traversal hacks to get direct communication among nodes, and I could imagine the application bidding more to persuade the infrastructure owner to provide computation in larger units (i.e. more cores on fewer machines, machines in a placement group with full bisection bandwidth, etc.), which is sort of where Lambda seems to be heading already.
(Long term, I don't really think applications should be renting cores and RAM per unit time and thinking about locality; I'd love to be dealing with the infrastructure provider in terms of some higher-level abstraction, because then you could imagine the provider might be genuinely incentivized to discover better ways of computing the same answer, to our mutual benefit.)
- r3trohack3r 5y agoI’m loving this train of thought Keith. What are your thoughts on program correctness and runaway cost. I’m a little uncomfortable running a workload that could scale unexpectedly to a denial of wallet. For this research, how did you enforce bounds on your workload to prevent exceeding your funding budget? Is the whole compute graph calculated locally? The recursive workloads seem particularly anxiety inducing.
- dflock 5y agoYou can set daily (and weekly, monthly etc...) budgets in AWS now, which helps. I think other providers have something similar.
- alfiedotwtf 5y ago> but nobody at the company knows because they only ran it once when they first joined and have just kept the same local dev VM ever since? Lol. I've got a blog post for exactly this, and is the reason why I try to push the opposite - Dockerise everything and treat machines like git branches i.e disposable. Life's great when you don't have to tread on eggshells around a pristine environment and you're only a single `apt-get install` away from disaster!