4 ms·
Original paper author here. I think we are in vehement agreement! 1a. we don't run big systems. Well, I don't, at least -- my coauthors wrote spark and similar
by stochastician 9y ago
Original paper author here. I think we are in vehement agreement!
1a. we don't run big systems. Well, I don't, at least -- my coauthors wrote spark and similar systems, so they're certainly aware of some of the challenges, but for us the major appeal here is someone else (cloud provider) is running a tremendous amount of the system!
1b. I think you're totally right that these sorts of systems tend to be generalized beyond their sweet spot. We have had quite passionate debates internally as to the amount of complexity we should add to the underlying system to support richer BSP efficiently. I just want to make sure the "obvious" stuff works as easily and quickly as possible (that is, basic map)
2. Segregating compute from storage _is_ untenable for a lot of workloads! But for a lot of the compute-heavy work we do, it's not. And we were really, truly shocked at how big the bandwidth to S3 really is -- for our compute applications, this degree of data disaggregation
3. "Model what the dev is using for dev" is _exactly_ what we're trying to accomplish. We even try and replicate the dev environment (thanks to anaconda, cloudpickle, etc.) as transparently as possible. That's exactly what we're trying to achieve, at the end of the day -- get back to where condor (another bird-named job system that we love) had us at 20 years ago.
- BatFastard 9y agoNo one runs big systems on such a platform until is has been used and tests by many parties. But it has to start somewhere! I remember having many discussions about exact this kind of platform 20 years ago. Still waiting for it.... My question for you is how do you bill for usage? This would seems like just the kind of platform that could use an ICO.
- oconnore 9y agoWhy didn't you write a nice python wrapper for AWS Batch? You'd get all the benefits of fully managed code, but none of the limits of Lambda functions. I don't see how any of your workloads are super latency sensitive -- a trivial AWS Batch job completes in ~2 minutes on a cold start.
- peterwwillis 9y agoI re-read the article and now it makes more sense. The takeaway seems to be "serialize your functions in bulk and de-serialize later in bulk": in essence, treat your compute like your storage. But I don't see how this is distributed computing for the 99%? PyWren appears to be what some of the serverless platform providers were missing, and also repeats a bit of history from the bad old days of parallel processing. While it totally works for a lot of scientific and a few business cases, it doesn't work for the general case. It's like saying MPI was the best solution for the 99%. But I might be overlooking another detail. The biggest problem with depending on one aspect of a system for performance is that it becomes the bottleneck. We'll rely on low-latency, high-bandwidth storage? Now if it becomes low latency or low bandwidth, our app is boned. We'll rely on loads of cheap compute cycles? If compute nodes start choking, our app is boned. We'll rely on RAM? Our data becomes bigger, and our app is boned. We'll just scale horizontally? Our network gets partitioned, and our app is boned. It can work fine for systems that don't require performance guarantees. If that's the 99%, then this seems like a good solution, though it also seems like you're trading off learning a complex stack for learning HPC development models.