Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
voidmain
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
15 ms
·
211.
▲
by
voidmain
8y ago
"Anonymization" in the sense of transforming a dataset so that it's still useful but doesn't significantly reduce the privacy of the people it describes, is usually impossible, or at least beyond the state of the art. Pe
212.
▲
by
voidmain
8y ago
Find a way to extract enough energy from this reaction (it has to be at least slightly exothermic) to pay for it, and we're good.
213.
▲
by
voidmain
8y ago
Unless you have serious simulation or model checking tools, your ad hoc combination of etcd and ceph will almost certainly be buggy. I'm not 100% positive it is even possible in the asynchronous model. And when you decide you want, say
214.
▲
by
voidmain
8y ago
A sophisticated approach to block compression would be to build a dictionary using something like zstandard's "training mode" and a random sample of data blocks, store it (versioned) in FDB as well and reference the dictionar
215.
▲
by
voidmain
8y ago
Awesome. I actually think that there is a lot of potential for using block storage over FDB to bring extreme fault tolerance to legacy applications without rearchitecting them. Because block devices are unshared (and you can use FDB to ensu
216.
▲
by
voidmain
8y ago
I think maybe you could improve on this sort of approach by weakening the invariants. Instead of trying to make sure that only valid "collations" get on the main chain, let anyone willing to pay for it put collations on the main c
217.
▲
by
voidmain
8y ago
Yes, and that's the main advantage of choosing (a) or (b). But it's not quite as hard as it sounds; since all your state is safely in fdb you "just" have to worry about load balancing a stateless service.
218.
▲
by
voidmain
8y ago
Long before we acquired Akiban, I prototyped a sql layer using (now defunct) sqlite4, which used a K/V store abstraction as its storage layer. I would guess that a virtual table implementation would be similar: easy to get working, an
219.
▲
by
voidmain
8y ago
From the paper you link: "A history is serializable if it is equivalent to one in which transactions appear to execute sequentially, i.e., without interleaving... A history is strictly serializable if the transactions’ order in the seq
220.
▲
by
voidmain
8y ago
Update: Apparently Spanner is way slower than I thought in "slow" datacenters, doing a round trip for every transactional read. So this would absolutely stomp that. (Although spanner, as a higher level system, has the ability
221.
▲
by
voidmain
8y ago
I explain the basics of our concurrency control here: https://news.ycombinator.com/item?id=16877950 I guess textbook SSI is willing to "reorder" conflicting transactions if the result is still serializable, which
222.
▲
by
voidmain
8y ago
A "layer" uses the fdb client much as it might use rocksdb. The layer can be a library embedded in your application, or a network service, it's up to you.
223.
▲
by
voidmain
8y ago
Sequential writes should be a little faster than random at the individual storage node level, but if your entire write workload is a single ordered log scalability will suffer. It might be theoretically possible for fdb to scale in this s
224.
▲
by
voidmain
8y ago
Yes. Linearizable means both serializable and externally consistent (or is sometimes used as just a synonym for the latter), and FDB has these properties with respect to transactions.
225.
▲
by
voidmain
8y ago
"I wanted a house, but I got a pile of drywall and 2x4 framing studs." This is a totally legitimate complaint about FoundationDB, which is designed specifically to be, well, a foundation rather than a house. If you try to live in
226.
▲
by
voidmain
8y ago
It should probably be pointed out that atomic increment is in most situations a more efficient solution for high contention counters in modern FDB.
227.
▲
by
voidmain
8y ago
I'm not sure. The coordination consensus is (our own implementation of) disk paxos, which we liked for its operational properties in our context (the coordinators don't need to know about each other or communicate directly). An ea
228.
▲
by
voidmain
8y ago
Yes.
229.
▲
by
voidmain
8y ago
Before FDB was acquired by Apple, a lot of the engineers used Visual Studio on Windows. It's a good IDE for C++. You certainly don't need it to develop though! Windows was never the most important deployment platform for the produ
230.
▲
by
voidmain
8y ago
You can think of FDB as a distributed storage engine. It has the same low level data model as the engines you mention, but has distributed transactions, fault tolerance, automatic data partitioning, operational tooling, etc built in. So i
231.
▲
by
voidmain
8y ago
Because the data model is ordered, large blobs can and normally should be mapped to a bunch of adjacent keys and read with a range read, not a single huge value. That also allows you to read or write just part of one efficiently.
232.
▲
by
voidmain
8y ago
Versionstamped operations and transaction logging are fully transactional. Watches are asynchronous: they are used to optimize a polling loop that would "work" without them.
233.
▲
by
voidmain
8y ago
Yes. Or if it's not important for the indexer to process things chronologically, you could just have an index of the primary key (only) of records that haven't been indexed. If you are trying to make your external index MVCC, then
234.
▲
by
voidmain
8y ago
In addition to single-key asynchronous watches, there are also versionstamped ops (for maintaining your own, sophisticated log in a layer) and configurable key range transaction logging (but see the caveats in my other post on the topic). I
235.
▲
by
voidmain
8y ago
I'm not sure what you are asking, but depending on their individual performance and security needs layers are usually either (a) libraries embedded into their clients, (b) services colocated with their clients, (c) services running
236.
▲
by
voidmain
8y ago
Yes. I don't know how well documented it is, but there is an API (well, system keyspace) that can configure the database to log transactions for a selected key range (up to and including the whole database) into another selected key ra
237.
▲
by
voidmain
8y ago
In some distributed databases the client just connects to some machine in the cluster and tells it what it wants to do. You pay the extra latency as it redirects these requests where they should go. In FDB's envisioned architecture, th
238.
▲
by
voidmain
8y ago
I think you are on the right track. Storing every individual (term, document, ...) in the key value store will not be efficient, but you should be able to take Lucene's nice fast immutable data structure and stuff blocks of it (at the
239.
▲
by
voidmain
8y ago
If you expect to have lots of (hopefully very temporary!) node failures, I think FoundationDB has another trick up its sleeve. You can store (say) N+2 replicas of transaction logs , which are also relatively small and (since sequential) e
240.
▲
by
voidmain
8y ago
I'm speculating, but I think in this mode, from the "slow" datacenters you would see one round trip time to start a transaction, then reads will be fast (they can be done safely from your local datacenter because of MVCC), an
More ›