Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
boredandroid
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
Sharing Is Caring: Multitenancy in Distributed Data Systems
(confluent.io)
2 points
by
boredandroid
10y ago
|
0 comments
32.
▲
Processing Tweets with Kafka Streams
(madewithtea.com)
2 points
by
boredandroid
10y ago
|
0 comments
33.
▲
Python Kafka Client Benchmarking
(activisiongamescience.github.io)
62 points
by
boredandroid
10y ago
|
18 comments
34.
▲
by
boredandroid
10y ago
Several people asked how this compares to Kafka (I'm one of the people who created Kafka at LinkedIn). Here's my take: I think the motivations they list are: 1. Different I/O model 2. Was started before Kafka had replication
35.
▲
by
boredandroid
11y ago
I'm one of the Kafka developers, happy to answer any questions.
36.
▲
Apache Kafka Security 101
(confluent.io)
7 points
by
boredandroid
11y ago
|
0 comments
37.
▲
by
boredandroid
11y ago
Thanks for clarifying! I think there are two separate things: 1. Is there a principled notion of when a write is "committed" and a leadership election algorithm that works with this to ensure committed writes aren't lost as l
38.
▲
by
boredandroid
11y ago
This post is a bit confused about how Kafka replication works. Replication in Kafka is always synchronous in the sense that the cluster internally has a strong notion of which messages are committed and no uncommitted message is handed out
39.
▲
by
boredandroid
11y ago
Here is how I think about this, there are three high level paradigms you see for processing: 1. Request/response (e.g. most UI actions, REST, etc) 2. Stream (e.g. subscribing to a Kafka topic) 3. Batch/periodic (e.g. Hadoop, DWH)
40.
▲
by
boredandroid
11y ago
One thing to realize is that a partitioned log is a generalization of an unpartitioned log (i.e. if you set # partitions = 1 in a partitioned log you have an unpartitioned log). In Kafka the purpose of partitions is to provide computational
41.
▲
Running Apache Kafka at Scale
(engineering.linkedin.com)
49 points
by
boredandroid
12y ago
|
9 comments
42.
▲
by
boredandroid
12y ago
I really think there are a couple of levels of immutability that it is easy to conflate. Specifically immutability for 1. In memory data structures...this is the contention of the functional programming people. 2. Persistent data stores. Th
43.
▲
by
boredandroid
12y ago
Anyhow in the Bay Area interested in learning more about Apache Samza should attend the meetup tonight in Mountain View: http://www.meetup.com/Bay-Area-Samza-Meetup/events/220354853...
44.
▲
by
boredandroid
12y ago
I agree with this comment, but I think this blog post is really addressing data flow at the company-wide or datacenter scale. Surprisingly there immutable event data hasn't really had much of a home at all beyond data warehousing.
45.
▲
Putting Apache Kafka to Use: A Guide to Building a Stream Data Platform
(blog.confluent.io)
33 points
by
boredandroid
12y ago
|
0 comments
46.
▲
What's Coming in Apache Kafka 0.8.2
(blog.confluent.io)
5 points
by
boredandroid
12y ago
|
0 comments
47.
▲
Apache Kafka for Beginners
(blog.cloudera.com)
1 points
by
boredandroid
12y ago
|
0 comments
48.
▲
by
boredandroid
12y ago
This is an excellent point, there IS a fundamental tradeoff between latency and throughput. Computers are much better a processing larger chunks of data linearly rather than small bits of data. However the question to ask is how much data d
49.
▲
by
boredandroid
12y ago
Yes, totally. In my definition the difference is that a stream processing system let's you define the frequency with which output is produced rather than forcing it to be "at the end of the data". This doesn't preclude b
50.
▲
by
boredandroid
12y ago
I am the author of the post. I think in cases where you are running totally different computations in different systems the Lambda architecture may make a lot of sense. However one assumption you may be making is that the stream processing
51.
▲
Questioning the Lambda Architecture
(radar.oreilly.com)
178 points
by
boredandroid
12y ago
|
12 comments
52.
▲
Concurrency is not about programming languages anymore
(blog.empathybox.com)
2 points
by
boredandroid
12y ago
|
0 comments
53.
▲
Benchmarking Apache Kafka: 2 Million Writes Per Second (On Three Cheap Machines)
(engineering.linkedin.com)
6 points
by
boredandroid
12y ago
|
0 comments
54.
▲
by
boredandroid
13y ago
Let me give several alternate reasons: 1. There are far fewer Germans. 2. The percentage of software engineers in Germany is lower. 3. Germany has fewer startups. Large companies typically take gender equality pretty seriously, if for no ot
55.
▲
by
boredandroid
13y ago
I am one of the Kafka committers. I had posted some notes on the Jepsen post which I think may give a more complete picture: http://blog.empathybox.com/post/62279088548/a-few-notes-on-k...
56.
▲
by
boredandroid
13y ago
Yup, agreed. Wanna help? :-)
57.
▲
by
boredandroid
13y ago
Yeah totally agree. I focused on real-time data because offline data processing is often able to be less principled about the usage of time and doesn't need an explicit log (e.g. many ETL pipelines work this way).
58.
▲
The Log: Real-time data's unifying abstraction
(engineering.linkedin.com)
271 points
by
boredandroid
13y ago
|
17 comments
59.
▲
by
boredandroid
13y ago
The rationale was that it should be a writer because we were building a distributed log or journal (a service dedicated to writing). The writer needed to be someone that (1) I liked, (2) sounded cool as a name, (3) was dead (because it woul
60.
▲
by
boredandroid
13y ago
Yup.
More ›