Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
georgewfraser
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
61.
▲
by
georgewfraser
6y ago
I bought a model Y with FSD a few months ago. It’s fantastic. “Navigate on autopilot” is game changing for highway driving. It’s 80% a useful as a completely self driving car.
62.
▲
by
georgewfraser
6y ago
Typo :)
63.
▲
by
georgewfraser
6y ago
Replicating a Postgres instance to a data warehouse takes ~1 hour to set up and $500/month, depending on the database size. The cost of the data warehouse is similar. If you can’t afford that cost, you’re not going to be able to afford
64.
▲
by
georgewfraser
6y ago
This is a needlessly complex solution. You will get better performance, with simpler maintenance, by replicating everything into an appropriate analytical database (Snowflake and BugQuery are both good choices). Setting up multiple Postgres
65.
▲
by
georgewfraser
6y ago
The entire category of "is X associated with COVID" research is a recipe for p-hacking, because you're looking at relationships between slowly changing variables, with all sorts of confounds, and there are so many ways to do
66.
▲
by
georgewfraser
6y ago
Beware that simply adding a column-oriented storage engine to a row store like Postgres is not going to get you anywhere near the performance of a ground-up columnar system like Redshift or Snowflake. This paper explains why [1]. Short vers
67.
▲
by
georgewfraser
6y ago
Do you have a source for this, or a code sample that can demonstrate it? This would be an extremely naive implementation of columnar storage. There are some truly hard cases around long variable-length strings, but any halfway decent column
68.
▲
by
georgewfraser
6y ago
Great article, one quibble: there isn’t really a clear dividing line between batch and streaming. If you process data one row at a time, that is clearly a streaming pipeline, but most systems that call themselves streaming actually process
69.
▲
by
georgewfraser
6y ago
Much of the value of Arrow is in the things that will get built after Arrow is widely supported by data warehouses. Much of the data ecosystem we have today was designed to avoid the cost of moving data between systems. The whole Hadoop e
70.
▲
by
georgewfraser
6y ago
The main benefit of a columnar representation in memory is it's more cache friendly for a typical analytical workload. For example, if I have a dataframe: (A int, B int, C int, D int) And I write: A + B In a columnar re
71.
▲
by
georgewfraser
6y ago
Arrow is the most important thing happening in the data ecosystem right now. It's going to allow you to run your choice of execution engine, on top of your choice of data store, as though they are designed to work together. It will m
72.
▲
by
georgewfraser
6y ago
Because it actually costs the same, and if you process them in Snowflake using SQL or UDFs, you will get your results in seconds and you won't have to manage any of the underlying infrastructure.
73.
▲
by
georgewfraser
6y ago
EXACTLY. You absolutely can store unstructured and semi structured data in Snowflake. I find it baffling and at this point a bit irritating that there is this community of people insisting that is not allowed for...some unspecified reason.
74.
▲
by
georgewfraser
6y ago
If Teradata is faster for workload X or Y, it isn’t because of a shared nothing architecture. Databricks and Snowflake both cache the working set on local disks of the worker nodes, so in practice they are shared nothing systems from a perf
75.
▲
by
georgewfraser
6y ago
Interestingly, BigQuery is also missing multi-table transactions. There are ways to live without this feature, but I agree it's a gap.
76.
▲
by
georgewfraser
6y ago
It's a lot simpler to use a single system as both your data lake, and your data warehouse. As Databricks gets better and better at the core data warehouse features, it becomes feasible to use it for both. Meanwhile, Snowflake and BQ ar
77.
▲
Databricks is an RDBMS
(fivetran.com)
148 points
by
georgewfraser
6y ago
|
89 comments
78.
▲
by
georgewfraser
6y ago
The fundamental challenge of open-source ETL is that high-quality connectors require understanding and working around all kinds of corner cases in the API of each data source. It’s very hard to get open source contributors to do this kind o
79.
▲
by
georgewfraser
6y ago
My point is that it doesn’t fail in edge cases. It has very clear limitations, there are no surprises. Even if you were literally asleep at the wheel, the worst thing that would happen is it’d end up stopped in front of a traffic light, wai
80.
▲
by
georgewfraser
6y ago
I have a model Y and what it has today is 80% as useful as a completely self driving car. It’s mostly self driving on the highway, and the situations it can’t handle are very predictable: toll booths, stopped cars in the middle of the road,
81.
▲
by
georgewfraser
6y ago
This kinda tells you everything about the culture of Google: > Sadly, despite the team’s groundbreaking technical achievements over the last 9 years — ... — the road to commercial viability has proven much longer and riskier than hoped.
82.
▲
by
georgewfraser
6y ago
These new GCs are amazing technology, but they primarily target pause time, whereas in data processing the primary concern is the “headroom” of extra space in your heap to allow the GC to work efficiently.
83.
▲
by
georgewfraser
6y ago
This kind of data infrastructure is a great use case for Rust. A lot of data infrastructure is memory-bound, so saving the memory overhead of GC is a huge win. The use of Arrow to support multiple programming languages is also a great conc
84.
▲
by
georgewfraser
6y ago
What is not said in this article is that you can use modern data warehouses, like Snowflake and BigQuery, in the exact same way: a single system that serves as both your data lake and your data warehouse. Databricks and the cloud data wareh
85.
▲
by
georgewfraser
6y ago
This is really bizarre. MicroStrategy is a solid BI tool, it's been solid for a long time, and the company basically operates as a cash machine. This is when you start issuing a dividend. If the founder-CEO is bored, he should promote
86.
▲
by
georgewfraser
6y ago
"Read-committed isolation" is not a meaningful implementation of transactions. If you can't do read, then a write, while guaranteeing the database didn't change in between, then you don't really have transactions.
87.
▲
by
georgewfraser
6y ago
We're trying to address a real problem that is happening in our industry: VPs of eng and principal engineers at startups are adopting the "Kappa Architecture" / "Turning the Database Inside Out", without realiz
88.
▲
by
georgewfraser
6y ago
That post explains that there are scenarios where it makes sense to store data permanently in Kafka. "Kafka is Not a Database" makes a different point, which is that Kafka doesn't solve any of the hard transaction-processing
89.
▲
by
georgewfraser
6y ago
It's funny that you use that example, we actually cited that in an earlier draft of this post. Despite the seemingly opposite title "Queues are Databases", that note actually makes many of the same arguments, that message bro
90.
▲
by
georgewfraser
6y ago
Joins were unavailable or subject to extreme limitations. Or just plain wrong!
More ›