5 ms·
With Google Cloud, Serverless has been a reality from 2015! And welcome to Serverless to those coming from AWS. Here is how Google Cloud achieves serverless. I
by obulpathi 9y ago
With Google Cloud, Serverless has been a reality from 2015! And welcome to Serverless to those coming from AWS.
Here is how Google Cloud achieves serverless.
Ingesting Data: PubSub. Scales to millions of messages instantaneously, no need to spin up and spin down the capacity/shards (like as in Kinesis), exposes RESTful interface for ingesting messages from anywhere (web/mobile ... ). If you want a high-performance interface, you get gRPC as well. With RESTful interface to consuming messages, you can connect PubSub to anything you want or trigger Cloud Functions / write to storage / ingest to Stackdriver (Monitoring system) using managed services.
Querying Data: Big Query. No need to spin up a single server. Just write your query in SQL and watch the magic of a thousand servers being spun up to serve the query in a fraction of second and process a petabyte of data in a couple of minutes. btw ... you can ingest data into Big Query in real-time and analyze results within in seconds.
Process Data using Dataflow: For workflows that are more complex than SQL queries to ones that need to be running continuously on streaming data. Dataflow is the serverless version of Spark. No more running our of memory errors, much fewer hassles with hotkeys, no more manual performance tuning of buffers, no more cleaning up log folders. If you are using Spark, but have not tried Dataflow, you are missing some serious magic.
Google Cloud ML for Machine Learning: Cloud ML is a hosted solution for running TensorFlow jobs. No more manual hyperparameter optimization, no more spinning up GPUs, no more scaling up the cluster size. It's all taken care for you.
Container Engine: Hosted version of Kubernetes. I am sure, everyone is aware of what K8 is and its capabilities.
I am from a Data / Analytics / ML background. The above was my reality of Serverless since 2015. Google Cloud has serverless options for Web (App Engine) / Mobile (Firebase), and other purposes as well.
Having done PhD in Cloud Computing and Big Data and I don't find Infrastructure/Big Data a sexy problem anymore. To a large extent, it's a solved problem. Building a data platform with a team of 3 people that can handle 10+ petabytes of data is easy. What is on the horizon and unsolved yet, is AI!
Would love to see if there are any other better/compelling Serverless options.
- stonogo 9y ago> Having done PhD in Cloud Computing and Big Data and I don't find Infrastructure/Big Data a sexy problem anymore. To a large extent, it's a solved problem. This kind of sounds like you've never managed a production system, sorry. There's also the whole "how do you process and manage lifecycle for an exabyte of classified/sensitive data" thing that nobody's nailed down yet. "Serverless" is not the panacea people make it out to be; as usual, it's marketing-speak for outsourcing ops, which isn't always feasible. People who talk about "data platforms" in terms of size are like people who judge coder productivity in lines of code. I guarantee your 10pb "data platform" will choke and die the minute someone shows up with 10gb of historical data and you have to get the timezones right.
- obulpathi 9y ago> This kind of sounds like you've never managed a production system, sorry. Don't be sorry. Instead, challange me with specific problems you have and I will show you how to build solutions. > how do you process and manage lifecycle for an exabyte of classified/sensitive data The infrastructure for storing/processing data is in place. If you doubt me, please go ahead and give the above-mentioned tech stack a try and let me know if you still have any problems. Or give me a detailed description of what you have to do and I will tell you how to do it. > People who talk about "data platforms" in terms of size are like people who judge coder productivity in lines of code Let me be more clear. I know how to build the infrastructure for that scales of data. I am not saying that I will write the code for processing data. That depends on what kind of data you got and what you want to do with it. > I guarantee your 10pb "data platform" will choke and die the minute someone shows up with 10gb of historical data and you have to get the timezones right. See my above comment for the answer.
- mschuster91 9y agoThat's not serverless by definition, you're describing Software-as-a-Service.
- obulpathi 9y agoTo me, the definition of Serverless is this: Developer writes code and the management of resources for running that code is taken care of by a Serverless platform. As a developer, I don't have to deal with servers, disks, networks, logs, metrics, ...
- doubleplusgood 9y agoSo your code assumes that the network/storage/DB are fault-free? You don't emit logs?
- obulpathi 9y ago> You don't emit logs? I emit logs. I mentioned that in my above comment as well. > So your code assumes that the network/storage/DB are fault-free? No, it does not assume that the network/storage/DB are fault-free, completely. It assumes they are fault free with the SLA limits provided by your Cloud provider. The software is written in a way that enables the platform to know about errors when then happen and remediate them. Like, when you build a website, you make the app serving layer stateless and choose a backend datastore that is replicated across regions and is highly available (like Cloud Datastore or Spanner). If your platform detects that a disk failed and your app is returning errors, then that instance is killed and an another instance is brought up. Very similar mechanisms exist for the above-mentioned data services as well for auto scaling, sharding, ... If your Cloud provider can not guarantee the SLA's or shows a green sign even when the service is down, IMO they are not competent enough.
- takeda 9y ago> As a developer, I don't have to deal with servers, disks, networks, logs, metrics, ... Yes, until things don't work as they supposed to. Disks and networks also don't matter if your application is simple. Disks are pointless if you don't keep a state. Networks will suddenly matter more when you start hiring its capacity. If your application is simple, you can also run it on a single machine and don't care about any of the things you listed. But just because those things were abstracted from you doesn't mean they don't exist and their limitations don't apply.
- falcolas 9y ago> Building a data platform with a team of 3 people that can handle 10+ petabytes of data is easy. I'd hate to see that server bill. I'm sure Google or Amazon would be happy to host that for you, at the cost of a full round of funding a month. Point of reference: I worked for a company doing some basic ML analysis on the Twitter firehose; only a few tens of gigs a day. Their hosting costs were significantly over half of their overall operating costs (despite huge discounts, and including downtown New York offices), and they still hadn't gotten profitable when I left (they had actually gotten to the point where they were laying people off). None of it was serverless either - it was measurably incapable of operating at the scale they required.