Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jeffffff
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
61.
▲
by
jeffffff
6y ago
this doesn't cut it. someone could take a list of email addresses, hash them, and then reidentify the dataset. hashing buys you nothing from a gdpr/ccpa compliance perspective, storing the hash is seen as no different from storing
62.
▲
by
jeffffff
6y ago
typically rehashing is done during insert. if you were going to do it while not doing anything else with the hash table you'd have to use a background thread, and then you'd have to do some form of synchronization, so there would
63.
▲
by
jeffffff
6y ago
iirc memcached either uses or used to use https://en.wikipedia.org/wiki/Linear_hashing to avoid rehashing induced pauses. it definitely makes sense for that use case but you're correct that it's not the right
64.
▲
by
jeffffff
6y ago
the size of the program is counted as part of the compressed size
65.
▲
by
jeffffff
6y ago
it's a real term but it's not a useful term to engineers. there is no such thing as a "data lake system". there are databases, filesystems, object stores, etc. where the term 'data lake' is actually useful is i
66.
▲
by
jeffffff
6y ago
this is about data for internal analytics purposes, typically meant to be queried using SQL. if you expect a typical data scientist or business analyst to pull together a bunch of data from a dozen microservices and join it themselves in or
67.
▲
by
jeffffff
6y ago
it's called 'tupperware'
68.
▲
by
jeffffff
6y ago
nevermind entity component system i knew that
69.
▲
by
jeffffff
6y ago
what is ECS in this context
70.
▲
by
jeffffff
6y ago
yeah redshift is not at all comparable to snowflake. big query is much closer, it's ahead in some areas and in the last year has closed some of the gaps where it wasn't. big query's biggest problem is that it's tied to g
71.
▲
by
jeffffff
6y ago
by that argument copyleft licenses aren't open source either? from a practical perspective this license is strictly more permissive than AGPL in that it only excludes cloud companies rather than all non-OSS companies.
72.
▲
by
jeffffff
6y ago
"There is a clear distinction between data I share publicly, data I share explicitly with US Facebook users, and data that I share in a limited setting not involving any US users." can you cite where the european laws and/or
73.
▲
by
jeffffff
6y ago
yeah i've learned the hard way not to give a customer an sla without a rate limit built into it
74.
▲
by
jeffffff
6y ago
at a certain scale you have to assume that at any given time at least one of your users is guaranteed to be doing something dumb that will cause performance issues. if you haven't implemented rate limiting or request prioritization at
75.
▲
by
jeffffff
6y ago
there isn't necessarily a fix. all compression algorithms expand the size of the data on most inputs in order to make certain inputs with sufficiently regular patterns a lot smaller. it just so happens that most of the images/vide
76.
▲
by
jeffffff
6y ago
if you can refactor without touching your tests and your tests still compile afterwards either the refactor was extremely trivial and didn't change any interfaces or you only had end to end tests.
77.
▲
by
jeffffff
6y ago
this sounds like a case where no amount of unit testing ever would've found the bug. someone found the bug either through reasoning about the implementation or using formal methods and then wrote a test to demonstrate it. you could spe
78.
▲
by
jeffffff
6y ago
apache hbase is full of awful state machine code implemented in a slight variant of "the wrong way" described in this paper. it is 100% unreadable spaghetti. here's one example: http://archive.cloudera.com/cdh
79.
▲
by
jeffffff
6y ago
a lot of the backlash is a byproduct of the name and the marketing. i don't think anyone is arguing that splitting things into services isn't ever a good idea. lots of companies were doing it before anyone had ever said the word &
80.
▲
by
jeffffff
6y ago
never worked at facebook but the answer is 1) testing 2) rolling/canary/blue-green deploys 3) feature flags 4) having a rollback plan
81.
▲
by
jeffffff
6y ago
it's not much more work but what's the point of it? the benefits of microservices seem to be the architectural separation of concerns and the downsides are from the distributed communication and deployment model, so in a lot of wa
82.
▲
by
jeffffff
6y ago
sounds kind of like this famous linus rant: https://lkml.org/lkml/2012/12/23/75
83.
▲
by
jeffffff
6y ago
assuming that data is being inserted continuously at the same rate, periodic full compaction of an lsm tree with an upper bound on time between full compactions is O(n^2). depending on the quantity and distribution of deletes, you can poten
84.
▲
by
jeffffff
6y ago
simd is a better fit than gpu for a lot of data warehousing workloads where you aren't doing much computation on each individual value and it's not worth the overhead of sending the data to the gpu. also true for search index perf
85.
▲
by
jeffffff
6y ago
game development seems to be moving away from OOP and towards entity component systems and data oriented design. i've heard arguments that OOP's killer application is desktop guis, which i could buy, but those become less and less
86.
▲
by
jeffffff
6y ago
i would make a distinction that data and state are two different things and argue that a lot of the mess people create with OOP is due to confusing the two. data is what exists externally to your program while state is strictly internal to
87.
▲
by
jeffffff
6y ago
while mandates to only use one tool can be overly restrictive, allowing a free for all is equally if not more dangerous. there are a lot of benefits to using a standard set of tools, and developers are for the most part really bad at calcul
88.
▲
by
jeffffff
6y ago
you need separate datastores for operational systems and analytics systems. operational datastores are optimized for point reads and low latency and analytical datastores are optimized for scans and high throughput. you also really don'
89.
▲
by
jeffffff
6y ago
most modern concurrent copying garbage collectors use memory protection and sigsegv handlers to avoid the need for locking
90.
▲
by
jeffffff
6y ago
thanks for your work on debezium! cdc to kafka has been a game changer for our data pipelines, so much easier to make generalized tooling at this level. our old pipeline requires either immutable rows with autoincrement primary keys or last
More ›