21 ms·
Materialize Raises a $32M Series B
- mrits 6y agoThe headline refers to "incrementally updated materialize views". How does a company get funding for a feature that has already existed in other DBs for at least a decade? E.g, Vertica refers to this as Live Aggregate Projections. It's a cool concept but comes with huge caveats. Keeping tracking of non-estimated cardinality for COUNT DISTINCT type queries, as an example.
- frankmcsherry 6y agoHi, I work at Materialize. You can read about Vertica's "Live Aggregate Projections" here: https://www.vertica.com/docs/9.2.x/HTML/Content/Authoring/AnalyzingData/AggregatedData/LiveAggregateProjectionCreate.htm https://www.vertica.com/docs/9.2.x/HTML/Content/Authoring/An... In particular, there are important constraints like (among others) > The projections can reference only one table. In Materialize you can spin up just about any SQL92 query, join eight relations together, have correlated subqueries, count distinct if you want. It is then all maintained incrementally. The lack of caveats is the main difference from the existing systems.
- jacques_chester 6y ago> The headline refers to "incrementally updated materialize views". How does a company get funding for a feature that has already existed in other DBs for at least a decade? They're getting funding for doing it much more efficiently. I read into the background papers when it first popped up. This is legitimate, deep computer science that other DBs don't yet have.
- hnmullany 6y agoMaterialize is the real deal - completely different architecture under the hood. Origin project is Timely Dataflow & Naiad. https://docs.rs/timely/0.11.1/timely/ https://docs.rs/timely/0.11.1/timely/
- benesch 6y ago(Disclaimer: I'm one of the engineers at Materialize.) > How does a company get funding for a feature that has already existed in other DBs for at least a decade? ... It's a cool concept but comes with huge caveats. I think you answered your own question here. Incrementally-maintained views in existing database systems typically come with huge caveats. In Materialize, they largely don't. Most other systems place severe restrictions on the kind of queries that can be incrementally maintained, limiting the queries to certain functions only, or aggregations only, or only queries without joins—or if they do support maintaining joins, often the joins must occur only on the involved tables' keys. In Materialize, by contrast, there are approximately no such restrictions. Want to incrementally-maintain a five-way join where some of the join keys are expressions, not key columns? No problem. That's not to say there aren't some caveats. We don't yet have a good story for incrementally-maintaining queries that observe the current wall-clock time [0]. And our query optimizer is still young (optimization of streaming queries is a rather open research problem), so for some more complicated queries you may not get the resource utilization you want out of the box. But, for many queries of impressive complexity, Materialize can incrementally-maintain results far faster than competing products—if those products can incrementally maintain those queries at all. The technology that makes Materialize special, in our opinion, is a novel incremental-compute framework called differential dataflow. There was an extensive HN discussion on the subject a while back that you might be interested in [1]. [0]: https://github.com/MaterializeInc/materialize/issues/2439 https://github.com/MaterializeInc/materialize/issues/2439 [1]: https://news.ycombinator.com/item?id=22359769 https://news.ycombinator.com/item?id=22359769
- mrits 6y agoThanks for the explanation. I'm going to look more into this as I'm working on a new service on top of Vertica. There is a lot I don't like about Vertica and don't see alternatives such as Snowflake to be much of an improvement.
- jamesblonde 6y agoWhat about the other big problem ignored here: does your streaming platform separate compute and storage? Because GCP DataFlow does. Flink doesn't. DataFlow allows you to elastically scale the compute you need (Snowflake, Databricks). If you can't do that, materialized views will be a more niche feature for bigger 24x7 deployments with predictable workflows.
- mwcampbell 6y ago> All of this comes in a single binary that is easy to install, easy to use, and easy to deploy. And it looks like they chose a sensible license for that binary [1], so they're not giving too much away. I wonder though if they could have made this work as a bootstrapped business, so they would answer only to customers, not to investors chasing growth at all costs. [1]: https://materialize.com/download/ https://materialize.com/download/
- offtop5 6y agoBootstrapping is fun until you can't make payroll. If your goal is an exit, and you can raise this much, why not.
- georgewfraser 6y agoMaterialize has tackled the hardest problem in data warehousing, materialized views, which has never really been solved, and built a solution on a completely new architecture. This solution is useful by itself, but I'm also watching eagerly how their road map [1] plays out, as they go back and build out features like persistence and start to look more like a full-fledged data warehouse, but one with the first correct implementation of materialized views. [1] https://materialize.com/blog-roadmap/ https://materialize.com/blog-roadmap/
- adamnemecek 6y agoHow were previous implementations of materialized views deficient?
- georgewfraser 6y agoJoins were unavailable or subject to extreme limitations. Or just plain wrong!
- ahupp 6y agoHere's a nice writeup of Materialize: https://lucperkins.dev/blog/new-db-tech-1/#materialize https://lucperkins.dev/blog/new-db-tech-1/#materialize Not really mentioned here, but in standard postgres it might be quite expensive to update the view so you can only do it periodically. Materialize keeps that up-to-date continuously.
- ako 6y agoRDBMSes enable you to create materialized views only for data in the database. Materialize enables you to do this for any streaming data source in your organization, with the ease of writing SQL. This enables you to simply write a SQL statement joining data from Salesforce + SAP + Siebel as soon as the data changes, and store the results as a near real-time up to date database table. It does depend on a lot of underlying plumbing: streaming platform (e.g. kafka), and streaming data sources (e.g., kafka connect + debezium).
- dataplayer 6y ago
- adamnemecek 6y agoThis is a big win for Rust.
- thuongleit 6y ago+1
- npiit 6y agoI wonder if BSL becomes the new standard for open source commercial products. It's a good trade-off between freedom and real world business pressure.
- nickstinemates 6y agoDoubt it. Lots of aversion to the license given its limited use and some ambiguous terms/education around the various windows.
- npiit 6y agoI think it can be, I know a few other potentially successful examples like CockroachDB and ZeroTier. The BSL license makes the entire project basically FOSS for you and me, but not for the big sharks. Which I guess is much better for the world compared to open-core and of course proprietary SaaS.
- deleted 6y ago[deleted]
- deleted 6y ago[deleted]
- haggy 6y agoCan you point me at documentation for the fault tolerance of the system? A huge issue for streaming systems (and largely unsolved AFAIK) is being able to guarantee that counts aren't duplicated when things fail. How does Materialize handle the relevant failure scenarios in order to prevent inaccurate counts/sums/etc?
- frankmcsherry 6y agoHi! I work at Materialize. I think the right starter take is that Materialize is a deterministic compute engine, one that relies on other infrastructure to act as the source of truth for your data. It can pull data out of your RDBMS's binlog, out of Debezium events you've put in to Kafka, out of local files, etc. On failure and restart, Materialize leans on the ability to return to the assumed source of truth, again a RDBMS + CDC or perhaps Kafka. I don't recommend thinking about Materialize as a place to sink your streaming events at the moment (there is movement in that direction, because the operational overhead of things like Kafka is real). The main difference is that unlike an OLTP system, Materialize doesn't have to make and persist non-deterministic choices about e.g. which transactions commit and which do not. That makes fault-tolerance a performance feature rather than a correctness feature, at which point there are a few other options as well (e.g. active-active). Hope this helps!
- deleted 6y ago[deleted]
- jgraettinger1 6y agoThis is a solved problem, for a few years now. The basic trick is to publish "pending" messages to the broker which are ACK'd by a later written message, only after the transaction and all it's effects have been committed to stable storage (somewhere). Meanwhile, you also capture consumption state (e.x. offsets) into the same database and transaction within which you're updating the materialization results of a streaming computation. Here's [1] a nice blog post from the Kafka folks on how they approached it. Gazette [2] (I'm the primary architect) also solves in with some different trade-offs: a "thicker" client, but with no head-of-line blocking and reduced end-to-end latency. Estuary Flow [3], built on Gazette, leverages this to provide exactly-once, incremental map/reduce and materializations into arbitrary databases. [1]: https://www.confluent.io/blog/exactly-once-semantics-are-possible-heres-how-apache-kafka-does-it/ https://www.confluent.io/blog/exactly-once-semantics-are-pos... [2]: https://gazette.readthedocs.io/en/latest/architecture-exactly-once.html https://gazette.readthedocs.io/en/latest/architecture-exactl... [3]: https://estuary.readthedocs.io/en/latest/README.html https://estuary.readthedocs.io/en/latest/README.html
- temuze 6y agoI'm glad more people are tackling this problem. There still isn't a good solution to real-time aggregation data at large scale. At a previous company, we dealt with huge data streams (~1TB data / minute) and our customers expected real-time aggregations. Making an in-house solution for this was incredibly difficult because each customer's data differed wildly. For example: - Customer A's shards might have so much cardinality where memory becomes an issue. - Customer B's shards might have so much throughput where CPU becomes a constraint. Sometimes a single aggregation may have so much throughput where you need to artificially increase the cardinality and aggregate the aggregations! This makes the optimal sharding strategy very complex. Ideally, you want to bin-pack memory-constrained aggregations with CPU-constrained aggregations. In my opinion, the ideal approach involves detecting the cardinality of each shard and bin-packing them.
- jstrong 6y agoI've always found that when you are solving a concrete problem, like you were, it's vastly easier than the case of a general-purpose database because you can make all the tradeoffs that benefit your exact use case. but it sounds like that's not what you experienced. was it just how heterogeneous the clients' needs were? I guess what I'm saying is, if you are capable of handling 1TB/minute, seems like you're plenty able to and would want to be designing the system yourself - but interested what I'm missing about this.
- jstrong 6y agocongrats to Frank McSherry and the rest of the materialized team! very impressed by your project.
- pgt 6y agoMaterialize can help us manifest The Web After Tomorrow [^1]. My previous comments persuading you why DDF is so crucial to the future of the Web: > "There is a big upset coming in the UX world as we converge toward a generalized implementation of the "diff & patch" pattern which underlies Git, React, compiler optimization, scene rendering, and query optimization." — https://news.ycombinator.com/item?id=21683385 https://news.ycombinator.com/item?id=21683385 also with links to prior art like Adapton and Incremental. > "DD (Differential Dataflow) is commercialized in Materialize" — https://news.ycombinator.com/item?id=24846119 https://news.ycombinator.com/item?id=24846119 > "Materialize exists to efficiently solve the view maintenance problem" https://news.ycombinator.com/item?id=22888396 https://news.ycombinator.com/item?id=22888396 [^1]: https://tonsky.me/blog/the-web-after-tomorrow/
- cocoflunchy 6y agoThanks for this, I'm glad to see I'm not the only one tired of writing everything twice (once in the frontend and once in the backend). I'll revisit the links later.
- coinwitcher 6y agoThis is interesting given what AWS just announced (AWS Glue Elastic Views): https://news.ycombinator.com/item?id=25267734 https://news.ycombinator.com/item?id=25267734
- deleted 6y ago[deleted]
- acjohnson55 6y agoI'm so psyched about Materialize. An old coworker explained to me about how his previous company used DBT to create many different projections of messy data to serve many applications, rather than trying to come up with the One Canonical Representation. It truly blew my mind in terms of thinking about how to model data within a business. The huge limitation with this vision is that it only works in places where you can tolerate some pretty significant staleness. So the promise of this approach excludes most OLTP applications. I simply assumed it wouldn't be reasonable to create something that allows for unconstrained SQL-based transformations in real time, and that no one was working on this. Oh well. But several months back, I discovered Materialize and it was an "oh shit" moment. Someone was actually doing this, and in a very first principles-driven approach. I'm really excited for how this project evolves.
- thejosh 6y agodbt is fantastic. It depends on your usecase, but for ours having an hourly sync of data is fine for reporting.
- drewbanin 6y agoWe're super excited about materialize at Fishtown Analytics (the company that makes dbt). I think that the "reporting" use-case is well served by Snowflake/BigQuery/etc, but I do think that operational use-cases are left behind by batch-based transformation models. The thing I'm most excited about in materialize is the ability to create Sinks (https://materialize.com/docs/sql/create-sink/ https://materialize.com/docs/sql/create-sink/). The combination of 1) streaming transformations over realtime source data and 2) streaming outputs into external systems feels like the way of the future IMO. I saw a dbt-materialize plugin (https://github.com/jwills/dbt-materialize https://github.com/jwills/dbt-materialize) out there in the wild. My guess is it's not ready for primetime yet, but would love to bake support for materialize into dbt when the time is right :)
- beoberha 6y agoLate to the post, but if anyone wants a good primer on Materialize (beyond what their actual engineers and a cofounder are saying in the comments), check out the Materialize Quarantine Database Lecture: https://db.cs.cmu.edu/events/db-seminar-spring-2020-db-group-building-materialize-a-streaming-sql-database-powered-by-timely-dataflow/ https://db.cs.cmu.edu/events/db-seminar-spring-2020-db-group...
- mavelikara 6y agoThe actual talk seems to be here: https://www.youtube.com/watch?v=9XTg09W5USM https://www.youtube.com/watch?v=9XTg09W5USM
- beoberha 6y agoThanks! Accidentally copied the wrong link in haste.
- gfody 6y agois there a single comprehensive list of restrictions on what can and can't be materialized? for example, if SQL Server can't efficiently maintain your materialized view then it doesn't let you create it - the whole list of restrictions is here: https://docs.microsoft.com/en-us/sql/relational-databases/views/create-indexed-views?view=sql-server-ver15 https://docs.microsoft.com/en-us/sql/relational-databases/vi... I'd love to be able to directly compare this with that Materialize is capable of - does a similar document exist?
- frankmcsherry 6y agoIt's easier to describe the things that cannot be materialized. The only rule at the moment is that you cannot currently maintain queries that use the functions `current_time()`, `now()`, and `mz_logical_timestamp()`. These are quantities that change automatically without data changing, and shaking out what maintaining them should mean is still open. Other than that, any SELECT query you can write can be materialized and incrementally maintained. https://materialize.com/docs/sql/select/ https://materialize.com/docs/sql/select/
- gfody 6y agothere are messages like this in the docs: > "WARNING! LATERAL subqueries can be very expensive to compute. For best results, do not materialize a view containing a LATERAL subquery without first inspecting the plan via the EXPLAIN statement. In many common patterns involving LATERAL joins, Materialize can optimize away the join entirely. " I take this to mean that Materialize cannot always efficiently maintain a view with lateral joins - that's fine neither can SQL Server, but it would be nice if I could find all these exceptions in one place like I can for SQL Server. ..fwiw I prefer the behavior of failing early rather than letting potential severe performance problems into prod. [1] https://materialize.com/docs/sql/join/#lateral-subqueries https://materialize.com/docs/sql/join/#lateral-subqueries
- frankmcsherry 6y ago> I take this to mean that Materialize cannot efficiently maintain a view with lateral joins [...] Well, no this isn't a correct take. Lateral joins introduce what is essentially a correlated subquery, and that can be surprisingly expensive, or it can be fine. If you aren't sure that it will be fine, check out the plan with the EXPLAIN statement. Here's some more to read about lateral joins in Materialize: https://materialize.com/lateral-joins-and-demand-driven-queries/ https://materialize.com/lateral-joins-and-demand-driven-quer...
- animeshjain 6y agoI was wondering if Materialize is meant to be used in analytical workloads only, or would it be equally up to the task for consumer app kind of workloads as well?
- frankmcsherry 6y agoHere's my take on this, from a few months back: https://materialize.com/lateral-joins-and-demand-driven-queries/ https://materialize.com/lateral-joins-and-demand-driven-quer...
- shay_ker 6y agoThis is maybe a silly question, but what's the difference between timely dataflow and Spark's execution engine? From my understanding they're doing very similar things - break down a sequence of functions on a stream of data, parallelize them on several machines, and then gather the results. I understand that the feature set of timely dataflow is more flexible than Spark - I just don't understand why (I couldn't figure it out from the paper, academic papers really go over my head).
- sixdimensional 6y agoDon’t forget to keep your eyes on the architectural concept of Command Query Record Separation (CQRS). When combined with event sourcing [1], there is a new unified architecture possible that solves the problem that microservices create by fragmenting data [2], and performant querying on data updating in real time. This architecture represents more complexity but increased flexibility. I recently saw this article about federated GraphQL [3], and while a cool idea and probably the ultimate solution (API composition), I expect that with network and physical boundaries between services still adding latency, we need materialized views as part of the architecture to compensate for the overhead of bringing together aggregate root objects from multiple systems. [1] https://www.confluent.io/blog/event-sourcing-cqrs-stream-processing-apache-kafka-whats-connection/ https://www.confluent.io/blog/event-sourcing-cqrs-stream-pro... [2] https://microservices.io/patterns/data/cqrs.html https://microservices.io/patterns/data/cqrs.html [3] https://netflixtechblog.com/how-netflix-scales-its-api-with-graphql-federation-part-1-ae3557c187e2 https://netflixtechblog.com/how-netflix-scales-its-api-with-...
- leeuw01 6y agoGreat to hear they got more funding!
- leeuw01 6y agoDoes anyone know how Materialize stacks up against VIATRA in terms of performance? VIATRA seems very similar to Materialize. They have multiple algorithms implemented to incrementalize queries, including Differential Dataflow. The main difference seems to be that it's based on Graph Patterns instead of SQL.
- frankmcsherry 6y agoIt's a good question, but you'd have to ask them I think. Tamas (from Itemis) and I were in touch for a while, mostly shaking out why DD was out-performing their previous approach, but I haven't heard from him since. My context at the time was that they were focused on doing single rounds of incremental updates, as in a PL UX, whereas DD aims at high throughput changes across multiple concurrent timestamps. That's old information though, so it could be very different now!
- leeuw01 6y agoThanks for the reply! A while ago (2018), the people behind VIATRA performed a cross-technology benchmark where they compared their performance to 9 other incremental and non-incremental solutions (Neo4j, Drools, OCL, SQLite, MySQL, among others) [1]. Perhaps it could be interesting to rerun that benchmark while including Materialize? This would give us a direct comparison between Materialize and other existing solutions. Their benchmark is however based on a kind of UX case, so the tests might be a bit biased towards that use case. [1] The Train Benchmark: cross-technology performance evaluation of continuous model queries