8 ms·
Uses and abuses of cloud data warehouses
- andrenotgiant 3y agoIt seems like Snowflake is going all-in on building features and doing marketing that encourage their customers to build applications, serving operational workloads, etc... on them. Things like in-product analytics, usage-based billing, personalization, etc... Anyone here taking them up on it? I'm genuinely curious how it's going.
- datadrivenangel 3y agoI assume they're angling for a salesforce acquisition as they move towards being a micro-hosting service like salesforce.
- rubiquity 3y agoSnowflake is worth at least 25% of Salesforce so such an acquisition is very unlikely unless Salesforce has $60 billion or more burning a hole in their pocket.
- politelemon 3y agoI've noticed that too. I think the marketing is definitely working, I'm seeing a few organisations starting to shift more and more workloads onto them, and some are also publishing datasets on their marketplace. One of their most interesting offerings coming up is Snowpark which lets you run a Python function as a UDF, within Snowflake. This way you don't have to transfer data around everywhere, just run it as part of your normal SQL statements. It's also possible to pickle a function and send it over... so conceivably one could train a data science model and run that as part of a SQL statement. This could get very interesting.
- jamesblonde 3y agoIn theory, fine. Then you look at the walled garden that is Snowpark - only "approved" python libraries are allowed there. It will be a very constrictive set of models you can train, and very constrictive feature engineering in Python. And, wait, aren't Python UDFs super-slow (GIL) - what about Pandas UDFs (wait that's PySpark.....)
- noazdad 3y agoDisclaimer: Snowflake employee here. You can add any Python library you want - as long as its dependencies are also 100% Python. Takes about a minute: pip install the package, zip it up, upload it to an internal Snowflake stage, then reference it in the IMPORTS=() directive in your Python. I did this with pydicom just the other day - worked a treat. So yes, not the depth and breadth of the entire Python ecosystem, but 1500+ native packages/versions on the Anaconda repo, plus this technique? Hardly a "walled garden".
- jamesblonde 3y agoGood luck with trying to install any non-trivial python library this way. And with AI moving so fast, do you think people will accept that they can't use the libraries they need, because you haven't approved them yet?!?
- Pils 3y agoHaving worked with a team using Snowpark, there are a couple things that bother me about it as a platform. For example, it only supported Python 3.8 until 3.9/10 recently entered preview mode. It feels a bit like a rushed project designed to compete with Databricks/Spark at the bullet point level, but not quite at the same quality level. But that's fine! It has only existed for around a year in public preview, and appears to be improving quickly. My issue was with how aggressively Snowflake sales tried to push it as a production-ready ML platform. Whenever I asked questions about version control/CI, model versioning/ops, package managers, etc. the sales engineers and data scientists consistently oversold the product.
- 3y ago
- lokar 3y agoAlso containers: https://www.snowflake.com/blog/snowpark-container-services-deploy-genai-full-stack-apps/ https://www.snowflake.com/blog/snowpark-container-services-d...
- atwebb 3y ago> run a Python function as a UDF Is that a differentiator? I'm unfamiliar with Snowpark's actual implementation but know SQL Server introduced Python/R in engine in 2016? something like that.
- munchor 3y agoDisclaimer: I work at SingleStoreDB. Building a database that can handle both analytics and operations is what we've been working on for the past 10+ years. Our customers use us to build applications with a strong analytical component to them (all of the use cases you mentioned and many more). How's it going? It's going really well! And we're working on some really cool things that will expand our offering from being a pure data storage solution to much more of a platform[1]. If you want to learn more about our architecture, we published this paper at SIGMOD in late 2022 about it[2]. [1]: https://davidgomes.com/databases-cant-be-just-databases-anymore/ https://davidgomes.com/databases-cant-be-just-databases-anym... [2]: https://dl.acm.org/doi/pdf/10.1145/3514221.3526055 https://dl.acm.org/doi/pdf/10.1145/3514221.3526055
- GeneralAntilles 3y agoYeah, they're providing a path-of-least-resistance for getting stuff done in your existing data environment. A common challenge in a lot of organizations is IT as a roadblock to deployment of internal tools coming from data teams. Snowflake is answering this with Streamlit. You get an easy platform for data people to use and deploy on and it can all be done within the business firewall under data governance within Snowflake.
- weego 3y agoAfter a series of calls, examples and explanations with them we never managed to get close to a reasonable projection of what our monthly costs would be like on Snowflake. I understand why companies in this field use abstract notions of 'processing' /'compute' units but it's a no go finance wise. Without some close to real world projections we don't have time to consider implementation to find out for ourselves.
- benjaminwootton 3y agoSnowflake is one of the easier tools to measure because it’s a simple function of region, instance size, uptime. If you can simulate some real loads and understand the usage then you do have a shot at forecasting. Of course the number is going to be high, but you have to remember it rolls up compute and requires less manpower. This is also a win for finance if they are comfortable with usage based billing.
- code_biologist 3y agoWho's finance team likes usage based billing? It makes sense for elastic use cases and is definitely "fair", but there are a lot of issues: Forecasting is hard. "dev team had an oops" situations. I had frog getting boiled situation at one job that was exactly the process described in the posted article: usage of the cloud data warehouse grew as people trusted the infrastructure and used more fresh data for more and more use cases. They were all good, sane use cases. I repeatedly under-forecast our cost growth until we made large changes and it really frustrated the finance people, rightly so.
- ramraj07 3y agoSnowflake is capturing a large market share in analytics industries thanks to its “just works” feature. I’m a massive fan. But in the end, snowflake stores the data in S3 as partitions. If you want to update a single value you have to replace the entire s3 partition. Similarly you need to read a reasonable amount of s3 data to retrieve even a single record. Thus you’re never going to get responses shorter than half a second (at best). As long as you don’t try and game around that limitation it works great. Materialize up here also follows the same model in the end FWIW.
- kwillets 3y agoIt's really bad for in-product analytics. Slow and expensive to keep running on a 24/7 SLA. Even with Vertica doing that, we're seeing 10x costs just doing back-office DWH. My job now is keeping Vertica running so we can pay our Snowflake bill.
- kitanata 3y ago[flagged]
- mritchie712 3y agoI caught myself wondering how Google, Microsoft and Amazon let Snowflake win. You can argue they haven't won, but lets assume they have. Two things: 1. SNOW's market cap is $50B. GOOGL, MSFT, AMZN are all over $1T. Owning Snowflake would be a drop in the bucket for any of them (let alone if they were splitting the revenue). 2. Snowflake runs on AWS, GCP or Azure (customers choice), so a good chunk of their revenue goes back to these services. Looking at these two points as the CEO of GOOGL, MSFT, or AMZN, I'd shrug away Snowflake "beating us". It's crazy that you can build a $50B company that your largest competitors barely care about.
- hnthrowaway0328 3y agoI agree. The cloud providers are basically the guys who sell shovels in gold rush. Snowflake still needs to build on top the clouds so MAG never lose. I heard that SNOW is offering its own cloud services but I could be wrong -- and even if I'm correct they have a super long way to catch up.
- pm90 3y agoWhat I heard is that AWS got there first with Redshift but then didn’t really invest as much as was required by users so Snowflake found an opening and pounced on it. BigQuery in GCP is a pretty great alternative and I know that GCP invests/promotes it heavily, but they were slightly late to the market.
- datadrivenangel 3y agoBigQuery is pretty great. The serverless by default setup works very well for most BI use cases. There are some weird issues when you're a heavy user and start hitting the normally hidden quotas.
- dsaavy 3y agoThere are some ways around the heavy user issues that aren't ideal but will work for BI-oriented heavy users.
- 3y ago
- albert_e 3y agoArent a lot of businesses being sold on "real time analytics" these days? That mixes the uses cases of analytics and operations because everyone is led to believe that things that happened in last 10 minutes must go through the analytics lens and yield actionable insights in real time so their operational systems can react/adapt instantly. Most business processes probably don't need anywhere near such real time analytics capability but it is very easy to think (or be convinced that) we do. Especially if I am a owner of a given business process (with an IT budget) why wouldn't I want the ability to understand trends in real-time and react to it if not get ahead of them and predict/be prepared. Anything less than that is seen as being shamefully behind on the tech curve. In this context-- the section in article where it says present data is of virtually zero importance to analytics is no longer true. We need a real solution even if we apply those (presumably complex and costly) solutions to only the most deserving use cases (and not abuse them). What is the current thinking in this space? I am sure there are technical solutions here but what is the framework to evaluate which use case actually deserves pursuing such a setup. Curious to hear.
- higeorge13 3y agoI work in a real time subscription analytics company (chartmogul.com). We fetch, normalize and aggregate various billing systems data and eventually visualize them into graphs and tables. I had this discussion with key people and i would say it depends on multiple factors. Small companies really like and require real-time analytics: they want to see how a couple invoices translate into updated saas metrics or why they didn’t get a slack/email notification as soon asit happened. Larger ones will check their data less frequently per day or week, but again it depends on the people and their role. Most of them are happy with getting their data once per day into their mailboxes or warehouses. But we try to make everyone happy so we aim for real time analytics.
- mrbungie 3y agoI think GP's point is that is not about the perceived value of real time data/analytics, but rather, its actual value. Decision makers may ask for RT or NRT, but most of the time won't make a decision or action in a timeframe that actually justifies RT/NRT data/analytics. For most operations RT/NRT data stuff normally is about novelty/vanity rather than a real existing business need.
- dontupvoteme 3y agoThe random bolding of words reeks of adtech. is the usage of such an old html tag itself now a trigger to send something to /dev/null?
- xiasongh 3y agoWhy does bolding words imply adtech?
- dontupvoteme 3y agoThey are bad. They indicate poorly written text.
- spullara 3y agoThese reasons are why Snowflake is building hybrid tables (under the Unistore umbrella). Those tables keep recent data in an operational store and historical data in their typical data warehouse storage systems. Best of both worlds. Still in private preview but definitely the answer to how you build applications that need both without using multiple databases and syncing. https://www.snowflake.com/guides/htap-hybrid-transactional-and-analytical-processing https://www.snowflake.com/guides/htap-hybrid-transactional-a...
- datavirtue 3y agoConveniently leave out the issue of cost. Snowflake is piling on features that encourage more compute. Customers abuse the system and they (Snowflake) respond by helping cement them into continuing the abuse (spending more) by developing features to make bad habits and horrible engineering decisions look like something they should be doing. Typical.
- disgruntledphd2 3y agoSnowflake are the Oracle of the cloud.
- ed_elliott_asc 3y agoOh come on snowflake isn’t cheap but there are none of the license auditing nonsense. (Also aren’t oracle the oracle of the cloud?)
- tcoff91 3y agoOracle is so far beyond anyone other major player in crookedness that it's not even funny. Oh, you happened to run your oracle database in a VM and got audited? They'll try to shake you down to pay for however many oracle licenses for every other box you are running hypervisors on, because they claim that you could have transferred the database to any of those other boxes. So if you have a datacenter with 1000 boxes running VMWare, and you ran Oracle on one of them, they try to shake you down for paying for not buying 1000 licenses. Then they say but if you just buy a bunch of cloud credits, we can make your 1000x license violation go away.
- atwong 3y agoThere are other databases today that do real time analytics (ClickHouse, Apache Druid, StarRocks along with Apache Pinot). I'd look at the ClickHouse Benchmark to see who are the competitors in that space and their relative performance.
- slotrans 3y agoYeah ClickHouse is definitely the way to go here. Its ability to serve queries with low latency and high concurrency is in an entirely different league from Snowflake, Redshift, BigQuery, etc.
- biggestdummy 3y agoStarRocks handles latency and concurrency as well as Clickhouse but also does joins. Less denormalization, and you can use the same platform for traditional BI/ad-hoc queries.
- riku_iki 3y agoClickhouse also does joins. Somehow StarRocks dudes appear in every relevant post with this false claim.
- biggestdummy 3y agoThere's a difference between "supports the syntax for joins" and "does joins efficiently enough that they are useful." My experience with Clickhouse is that its joins are not performant enough to be useful. So the best practice in most cases is to denormalize. I should have been more specific in my earlier comment.
- riku_iki 3y agoack that anonymous user in internet said he couldn't make CLickhouse joins perform well in his case which he didn't describe
- mrbungie 3y agoI remeber one time I was working as a Data & Analytics Lead (almost a Chief Data Officer but without the title) in a company were I don't work anymore and I was "challenged" by our parent company CDO about our data tech stack and operations. Just for context, my team at the time was me working as the lead and main Data Engineer plus 3 Data Analysts that I was coaching/teaching to convert into DEngs/DScientists. At the time we were mostly a batch data shop, based on Apache Airflow + K8S + BigQuery + GCS in Google Cloud Platform, with BigQuery + GCS as the central datalake techs for analytics and processing. We still had RT capabilities due to having also some Flink processes running in the K8S cluster, and also having time-critical (time, not latency) processes running in microbatches of minutes for NRT. It was pretty cheap and sufficiently reliable, with both Airflow and Flink having self-healing capabilities at least at the node/process level (and even cluster/region level should we need it and be willing to increase the costs), while also allowing for some changes down the road like moving out of BQ if the costs scaled up too much. What they wanted us to implement what according to them was the industry "best practices" circa 2021: a Kafka-based Datalake (KSQL and co.), at least other 4 engines (Trino, Pinot, Postgres and Flink) and an external object storage with most of the stuff running inside Docker containers orchestrated by Ansible in N compute instances manually controlled from a bastion instance. For some reason, they insisted on having a real time datalake based on Kafka. It was an insane mix of cargo cult, FOMO, high operational complexity and low reliability in one package. I resisted the idea until the last second I was in that place. I reunited with some of my team members for drinks months later after my departure and they told me the new CDO was already convinced that said "RT-based" datalake was the way to go forward. I still shudder every time I remember the architectural diagram and I hope they didn't finally follow that terrible advice. tl;dr: I will never understand the cargo cult around real time data and analytics but it is a thing that appeals to both decision makers and "data workers". Most businesses and operations (especially those whose main focus is not IT by itself) won't act or decide in hours, but rather in days. Build around your main use case and then make exceptions, not the other way around.
- chuckhend 3y agoI agree that is a great approach - build around the main use cases and then make exceptions. I think a lot of companies have legitimate use cases for real-time analytics (outside of their internal decision making), but as you mention, preemptively optimize for the aspiration and leads them towards unnecessary tool and tech sprawl. For example, a marketplace application that shows you the quantity of an item currently available -- you as a consumer use that information to make a decision in seconds, so its a great use-case. Internally, the org probably uses that data for weekly or quarterly forecasting. I've seen use cases like that lead to the "let's make everything real-time", but not every use case benefits the same from real-time.
- hodgesrm 3y agoThis article uses an either or definition that leaves out a big set of use cases that combine operational and analytic usage: > First, a working definition. An operational tool facilitates the day-to-day operation of your business. Think of it in contrast to analytical tools that facilitate historical analysis of your business to inform longer term resource allocation or strategy. Security event and incident management (SEIM) is a typical example. You want fast notification on events combined with the ability to sift through history extremely quickly to assesss problems. This is precisely the niche occupied by real-time analytic databases like ClickHouse, Druid, and Pinot.
- debarshri 3y ago15 years ago when I joined workforce business intelligence was all the rage. Data world was pretty much straight forward. You had transactional data in OLTP databases which would be shipped to Operational data stores, then rolled into the data warehouse. Datawarehouses were actual specialised hardware appliances (netezza et al) reporting tools were robust too. Everytime I moved from one org to another, these concepts of data warehouse somehow got muddled.
- bob1029 3y ago> Operational workloads have fundamental requirements that are diametrically opposite from the requirements for analytical systems, and we’re finding that a tool designed for the latter doesn’t always solve for the former. We aren't even going to consider the other direction? Running your analytics on top of a basic-ass SQL database? In our shop, we aren't going for a separation between operational and analytical. The scale of our business and the technology available means we can use one big database for everything [0]. The only remaining challenge is to construct the schema such that consumers of the data are made aware of the rates of change and freshness of the rows (load interval, load date, etc). If someone wants to join operational with analytical, I think they shouldn't have to reach for a weird abstraction. Just write SQL like you always would and be aware that certain data sources might change faster than others. Sticking everything onto one target might sound like a risky thing, but I find many of these other "best practices" DW architectures to be far less palatable (aka sketchier) than one big box. Disaster recovery of 100% of our data is handled with replication of a single transaction log and is easy to validate. [0]: https://learn.microsoft.com/en-us/azure/azure-sql/database/hyperscale-architecture?view=azuresql#hyperscale-architecture-overview https://learn.microsoft.com/en-us/azure/azure-sql/database/h...