5 ms·
Hadoop can do realtime. One example is Cloudera Impala, which can do small SQL queries in seconds or less. Another, non-SQL example is using the Lambda Archit
by cmccabe 13y ago
Hadoop can do realtime. One example is Cloudera Impala, which can do small SQL queries in seconds or less. Another, non-SQL example is using the Lambda Architecture (http://jameskinley.tumblr.com/post/37398560534/the-lambda-architecture-principles-for-architecting http://jameskinley.tumblr.com/post/37398560534/the-lambda-ar...) with something like Storm or S4.
- jandrewrogers 13y agoReal-time is round-trip time: the latency between when new data is available for ingest and when that data shows up in queries. Database engines that are designed for these types of workloads have a round-trip latency measured in milliseconds, and some of them can do this while ingesting millions of records per second at petabyte scales. As in, very fast SQL queries concurrent with extremely high continuous ingest workloads. The term does not mean "fast queries". Otherwise, many parallel SQL databases would be "real-time" because some of those are even faster than Impala in this regard. A system that uses offline data loading is not real-time. Stream processing systems like Storm are real-time in this sense but are not databases. I don't think you can run ad hoc SQL queries against the data being processed by Storm nor can it store those streams to disk for future queries. They also aren't "big" data due to the limitations of the architecture.
- cmccabe 13y agoYou're missing the point, which is that the current limitations of the system may not be the limitations in the future.
- jandrewrogers 13y agoCurrent limitations do imply future limitations. Database engines (and similar software) necessarily embed a large number of tradeoffs and assumptions in their design at the most fundamental levels. Every line of code is written to support the target workload to the exclusion of others because of the tradeoffs required. Design decisions are extremely sticky over the long-term because they are tacitly embedded in every piece of code. The point you are missing is that (1) significantly altering the basic architectural characteristics is tantamount to a complete rewrite from scratch and (2) existing users design their applications around the design assumptions of the platform so fundamentally changing the architecture abandons the user base as well. This is why in practice almost everyone starts from a blank slate if they need new architectural capabilities. And in the few cases where they did manage to re-architect an existing system, not only did it cost more than doing it from scratch but they lost their user base anyway. I've designed these types of systems for a long time. When a database engine is first designed, its capabilities, limitations, and future performance ceiling are essentially set in stone. All potential modifications have to fit within those constraints so you need to be cognizant of the choices you are making but not really thinking about. An experienced database engine designer can look at a system design and tell you what types of data models and workloads the system will always do poorly on. And every database engine design has weaknesses designed into them. Hadoop just happens to be a particularly weak engine with a low ceiling in terms of its expressiveness.