4 ms·
As a data guy, I've have an entirely different perspective on serverless. While it may not work well for running entire parts of an application, it works great
by blakeburch 6y ago
As a data guy, I've have an entirely different perspective on serverless. While it may not work well for running entire parts of an application, it works great for building out and managing internal data pipelines. You don't want your pipelines crashing because your jobs overlapped and you ran out of memory/cpu on the servers you set up. You want each stage of the pipeline to run independently and scale as needed.
In my eyes, the bigger problem with many FaaS products (which most equate with serverless) is the barrier to entry. The setup isn't very user friendly. The code doesn't "just work" like it does locally. Limitations on data size, runtime, etc. cause you to building workarounds in your scripts just to get them running. Not to mention that once you have it all set up, visibility into everything running is a nightmare and only available to the most technical users.
Based on my experience, I'm currently building a [platform](https://www.shipyardapp.com https://www.shipyardapp.com) to try and make serverless data pipelines easier for teams to setup and manage. Would love to hear someone else's perspectives on serverless setups for data management. I know I'm not alone on these existing frustrations.
- alexpetralia 6y agoInteresting, I have found AWS Lambda quite poor for data pipelines. Batch jobs can't exceed 15 minutes. Memory limit is 3GB. Payload sizes are 256kb max. If you are used to batch processing large amounts of data, Lambda seems to be the exact opposite of this. I think it is good for "small data, highly concurrent event processing" but this is a very different use case from batch processing data pipelines.