4 ms·
(Co-author here) It really depends (see Figure 2). On the question of Lambda vs. EC2, EC2 instances take much longer to start. So depending on the job, to get
by keithwinstein 7y ago
(Co-author here) It really depends (see Figure 2).
On the question of Lambda vs. EC2, EC2 instances take much longer to start. So depending on the job, to get the performance you can get with a "burst parallel" flock of Lambda workers, you would need to keep a warm cluster of EC2 instances ready to take your job. At which point, the cost comparison depends on how often you have work to execute. EC2 is cheaper if you have a 100% duty cycle (but you probably don't).
The gg tool, though, is mostly agnostic to the backend -- you can take a job that's expressed in gg IR (e.g., "compile this program") and then execute it with any of the gg back-ends. We have one for Lambda and one for a cluster of warm VMs. The performance of gg-to-EC2 is generally better than outsourcing methods that leave your laptop in the driver's seat (e.g. bazel-to-icecc) and give less semantic information about data- and control-flow of the job to the remote execution engine. (E.g. in Figure 9, you can see that gg-on-EC2 is much faster than icecc-on-EC2 for compiling GIMP and Inkscape.)
- s_Hogg 7y agoThat makes sense, thanks for taking the time to expand on this point.
- crb002 7y agoWhere is the DAG scheduler in the source? About to take a look. Interested in labeling thunks with file size and CPU time then feeding into a solver I am writing to optimize cost within a runtime budget.
- sadjad 7y agoCo-author here. You may wanna take a look at https://github.com/StanfordSNR/gg/blob/master/src/execution/reductor.hh https://github.com/StanfordSNR/gg/blob/master/src/execution/... and https://github.com/StanfordSNR/gg/blob/master/src/execution/reductor.cc https://github.com/StanfordSNR/gg/blob/master/src/execution/....
- dajohnson89 7y agothe startup time isn't always that important. whether a task takes 20 seconds or 60 seconds isnt always something the user cares that much about. and as always we have to weigh the developer headache versus the performance improvement. the customer always wins, but the less headaches the devs have, the more responsive they are to issues that are more pressing than marginal performance improvements.
- deleted 7y ago[deleted]
- polskibus 7y agoAny plans to evolve gg to use open source serverless platforms like knative, openwhisk or openfaas?
- keithwinstein 7y agoSure -- we actually already have an OpenWhisk backend. The IR is sufficiently stupid that it's pretty easy to write a new backend. At least a low-performing one -- it gets harder if you want to use (a) persistent workers [instead of invoking a new platform worker for every "thunk"] and (b) direct inter-worker networking [instead of putting every intermediate result in a storage medium].