9 ms·
πfs – A data-free filesystem
- bmicraft 3y agoThat's just use he library of babel all over again http://libraryofbabel.info/ http://libraryofbabel.info/
- mal10c 3y ago"That's right! Every file you've ever created, or anyone else has created or will create! Copyright infringement? It's just a few digits of π! They were always there!" I didn't think that was actually mathematically proven yet. Was some proof accepted recently that makes that quoted sentence true?
- nivekney 3y agoThe proof exists in πfs, you just need to know the metadata to retrieve it!
- Solvency 3y agoBut so do all of the lawyers' rebuttals! And I'll bet my money they'll find the metadata for them first...
- WillAdams 3y agoThere was an examination of this sort of thing in the webcomic Freefall --- shortest story is 6 words, 100,000 words for English vocabulary, order matters --- it worked out to a data storage unit the size of a small moon.
- teddyh 3y agoLinks: <http://freefall.purrsia.com/ff2100/fv02058.htm http://freefall.purrsia.com/ff2100/fv02058.htm> <http://freefall.purrsia.com/ff2100/fv02059.htm http://freefall.purrsia.com/ff2100/fv02059.htm> <http://freefall.purrsia.com/ff2100/fv02060.htm http://freefall.purrsia.com/ff2100/fv02060.htm>
- WillAdams 3y agoThanks for posting those! I was off by a factor of 10.
- ravi-delia 3y agoNo one has proved that pi is normal, true. But pi is normal. I mean come on! Like just between us, two friends talking, we both know that pi is normal. It'll be a nightmare to prove but I'd put extremely good odds on it. Probably better odds than RSA being secure, which is also not strictly proven but pretty likely.
- kevincox 3y agoI love the idea of this, it has certainly boggled my mind a handful of times when thinking of content-addressed storage. Obviously it can't work, there is no loophole to infinite space or compression. Entropy is a very real thing. But it did lead me down an interesting train of thought. > maximise performance, we consider each individual byte of the file separately, and look it up in π. Ok, so we store a byte and look it up in π. Now we get an offset. The exact offset will depend on the byte of course. But to simplify let's assume that π is "optimal". We will assume that the fist 256 offsets contain the first 256 bytes. So our offset will be in the range 0-255. Storing our offset will then take 1 byte of storage. Oh, I have found the problem. So yes, you can find any data in π. But storing the location of that data will on average take the same amount of space as the data itself.
- ignoramous 3y ago> I love the idea of this, it has certainly boggled my mind a handful of times when thinking of content-addressed storage You only need 2 bits 1 and 0 to store every file there can ever be, and then store sequences in file metadata. Call it bifs.
- SirMaster 3y agoWhy are you storing an offset for each byte? I thought you are supposed to find the offset where your entire file is sequentially there in pi. Because if it's infinite and not repeating, then that string should be in there somewhere.
- kevincox 3y agoBecause that is what the project does. But I suspect that the analysis would look very similar for the entire file. Just calculate the average offset of the file of a given size. It won't be smaller than the file itself.
- kevincox 3y agoI guess that is mostly true. In theory you could think of a more efficient way to store the indexes such as run-length encoding. So you could store 1M of whatever the first byte of Pi is and then run-length encode the 1M zero addresses. You can also imagine a scheme such as 2 bit varint encoding the indexes or a tally system that only uses 1 bit per offset to store a zero. ...of course you are still better off to just compress the file consisting of a single repeated byte.
- r3trohack3r 3y agoThis is a fallacy. Infinity does not contain all possibilities. There are an infinite number of integers. You can start at 10 and count up forever, never running out of integers. But no matter how high you count, you’ll never count to “orange” - “orange” is not contained in the sequence of infinite integers. You’ll need to first prove that every sequence of integers is contained somewhere in pi, since the number of possible integer sequences grows faster than the “space” for sequences in pi. In other words, I can always pick a digit that creates a valid, non repeating, integer sequence from the pool of possible sequences while never creating the integer sequence “123456789123456789123456789.” You’d need to prove that pi doesn’t do this. Even if pi does contain every sequence of integers and you could map that to bytes which, in turn, maps to a file, this would not compress. Your metadata directory would be larger than the raw files unless you get very lucky and your file is very early in the sequence of pi. A byte can represent 256 unique values. 256 unique values can not compress to less than a byte. So if your index is a digit of pi where your file starts, your file starts after some other number of files. Your index is going to be the index inside of the address space of “all possible files.” This will get large very quickly.
- jermaustin1 3y ago> you’ll never count to “orange” - “orange” is not contained in the sequence of infinite integers. What is a string, but a sequence of bytes? What is a sequence of bytes, but a decomposed integer? "orange" = 111,494,907,916,911
- r3trohack3r 3y agoNo. Encoding is not a clever way around infinity. You can encode orange in RGB and HSL too, but the set of all integers still does not contain the concept of an orange. You’ve just assigned meaning to an integer. In the same way you can’t count to orange, you can never start at 1 and count up to -1. There are an infinite number of different integers greater than 1, and that infinity does not contain -1. SciFi likes abusing this. Just because there are an infinite number of universes doesn’t mean there is a universe that contains anything you can dream up Rick and Morty style.
- jakelazaroff 3y ago> You'll never run out of space again - π holds every file that could possibly exist! This isn't necessarily true, right? AFAIK this only holds if pi is normal, which we haven't proven.
- hackernewds 3y agoThought John Lamhert did that in 1728 https://www.sciencefocus.com/science/how-do-we-know-that-pi-is-infinite/ https://www.sciencefocus.com/science/how-do-we-know-that-pi-...
- constantcrying 3y agoNo, this proves pi can not be written as a fraction. But not being a fraction does not imply that a number contains all sequences of digits.
- dragonwriter 3y agoAs a fairly simple example, consider the number wirtten in decimal notation as: 0.101001000100001000001... (with an increasing-by-one sequence of zeroes between each successive 1.) It doesn't contain all sequences of digits, or even all digits, yet it cannot be written as a fraction.
- constantcrying 3y ago>yet it cannot be written as a fraction. Is that trivial?
- dragonwriter 3y ago> Is that trivial? My understanding is that it is a trivial (or at least well-known) result that a rational will always have a finite, infinitely-repeating terminal sequence in any base.
- xigoi 3y ago
- salgorithm 3y agoPeople will do anything nowadays to lower their AWS bill.
- warent 3y agoIn this implementation, to maximise performance, we consider each individual byte of the file separately, and look it up in π. LOL so you get to store your data for free, and all it takes is allocating like 8 bytes for every byte. This project is galaxy brain
- therobotking 3y agoHow am I reading so many comments not realising this repo is tongue-in-cheek?
- david422 3y agoOh, they probably do .. but then there's not much to discuss.
- jsdeveloper 3y agoI personally thought this during my college days. Never pursued this idea but kept it my mind as a way to compress files and share just index to first meta data, which will then link to next and next and so on.
- pk-protect-ai 3y agolmao :) Love it!!!!
- miahwilde 3y agoWake me up when it's self-bootstrapped. This repo should contain the metadata for the source code of itself in pi. Or at the very least the metadata for "Everything can be stored in pie" in pi.
- LoganDark 3y ago> Everything can be stored in pie Isn't this where the old meme of a hacksaw inside a cake came from?
- jazzsax 3y agoIf I were to dumb this down (so I can understand it), is this a fair analogy to the adage "give 1 million monkeys 1 million typewriters and they'll eventually type the entire works of William Shakespeare"? Somewhere in pi (at some insane offset) is the entire work of William Shakespeare. Is that the basic idea here?
- j_walter 3y agoBingo
- jadbox 3y agoTechnically 1 million monkeys with 1 million typewriters could type for a million years and yet still not produce a work of William Shakespeare. In theory (as I understand it), pi at some index contains all permutations of sets of values. I'm not a mathematician and only know this by hearsay.
- LoganDark 3y ago> Technically 1 million monkeys with 1 million typewriters could type for a million years and yet still not produce a work of William Shakespeare You have to calculate the probability of a monkey coming up with a work of Shakespeare, and then subtract that from 1 to get the probability of a monkey coming up with anything other than a work of Shakespeare, and then figure out what power N to raise that to (N attempts) for it to cross below some confidence value, like 0.05 ... then subtract that from 1 in order to get the probability of a monkey coming up with a work of Shakespeare after that many attempts. (no matter how high N is, you will never get 0, so no matter how many times you attempt, you always "could" still not produce a work of Shakespeare, because there's always a chance that none of those attempts worked) Then calculate how much time that might take (actually, you might even have to account for failed attempts that are shorter or longer than a work of Shakespeare, so you might even have to start with time rather than number of attempts, but I digress) and you have a possible figure for "how much time it would take to have a certain probability of a monkey having come up with a work of shakespeare" and oh dear I may have missed the joke
- paddim8 3y ago
- letsawaytoday 3y ago[flagged]
- adamgordonbell 3y agoI had an idea for a file transfer system based on digits of pi. You'd give everybody DVDs of digits of pi (or they could calculate them themselves) , and then transfer files faster by just sending them the offset into pi. At the time I thought it could work with a big enough bank of digits of PI on both sides. If transfer was expensive, and calculating digits was cheap then you could give everyone an infinite supply of digits of pi and have a nearly infinite compression system. I discovered that often the offset into pi is much larger than the data you are sending. Turns out it's an expensive way to sent things. Also, it turns out that this area was already well understood. There are no free lunches with entropy. But it was a fun idea to kick around.
- SillyUsername 3y agoI think a lot of people had similar ideas, I know I did. The idea of using Pi is essentially the same as having a shared dictionary as used in both regular compression and in lossy as waveforms (in layman's, terms - technically it's a best match) both ends have the dictionary and you simply index it. In fact whole network protocols use this concept too such as protocolbuffers which are part of grpc. The difference here being the dictionary is infinite and until we get quantum computers on the desktop indexing into pi will always be slow. Some of the address space may not be feasible too, e.g. your file may require a billion bit address/index. It may still be feasible to do partial matches for a file, 50% one index, 50% another for performance improvement. I guess the trade off is a balance between performance and number of indexes Vs the original file length, and where in Pi the address space becomes unfeasible.
- adamgordonbell 3y agoIt's interesting how zstd can use a shared dictionary to compress small amounts of data. It's a more practical ( then my idea ) shared dictionary.
- wahern 3y agozlib also supports dictionaries, as does the deflate algorithm. The problem is that zlib doesn't provide any tooling to generate an optimal dictionary from a large sample, whereas zstd does provide such a utility.
- bzmrgonz 3y agoSo Pi=akashic records??
- myaccount80 3y agoYou can encode every book and information in a small rod. Just take a 1 meter long rod and encode your book into binary, eg 01011110111. Now, take your binary string, and transform it into a decimal number in base 2 by prepending 0. , E.g. x=0.01011110111. As this number is finite and smaller than 1, you can just take your rod and cut it at length x. Now when you want to retrieve your information you just need to measure x from your rod with a ruler, take the decimal part and convert it into binary, and voila. You can encode almost infinite amount of data into a simple rod. Assuming you can measure and cut very precisely
- capitainenemo 3y agoSmall problem of how to make cuts at a subatomic level ;) (edit for pedantism and type of atom) at an optimistic 100 billion carbon atoms in length for a metre stick, the maximum distinct cuts you can make are 100 billion, or 100 gigabytes of info. We do much better these days with thumb drives.
- chrisshroba 3y agoThat 's only if you can make multiple cuts (or slits), and it's actually 8x less because we're talking about bits, not bytes. With only one cut as the parent comment describes, you can only store the log base 2 of 100 billion, which is 36, so about 4 bytes of info, or one long integer.
- capitainenemo 3y agoAh true. 100 gigabits.. can't edit my comment now. But oh well, it was pretty rough anyway, so "on the close order" ;)
- jclarkcom 3y agoNot to mention thermal expansion which doesn’t happen evenly and trying to cut or measure would have an effect
- 3y ago
- CuriousSkeptic 3y ago> In this implementation, to maximise performance, we consider each individual byte of the file separately, and look it up in π. Had me laughing out loud. Priceless!
- b33j0r 3y agoI irrationally think that I go to too much effort to make a point in our field with over-the-top clever stuff. Nope, I’m undertraining. Get ready for TauOS.
- ftxbro 3y agoIn this quintessential hacker news post and discussion we see: eating the onion anecdotes of having the same idea https://en.wikipedia.org/wiki/Normal_number https://en.wikipedia.org/wiki/Pigeonhole_principle
- dang 3y agoRelated: πfs – A data-free filesystem - https://news.ycombinator.com/item?id=28699499 https://news.ycombinator.com/item?id=28699499 - Sept 2021 (30 comments) PiFS – The Data-Free Filesystem - https://news.ycombinator.com/item?id=26208704 https://news.ycombinator.com/item?id=26208704 - Feb 2021 (1 comment) Πfs: Never worry about data again - https://news.ycombinator.com/item?id=21359338 https://news.ycombinator.com/item?id=21359338 - Oct 2019 (1 comment) The π Filesystem for FUSE: Store Your Data in π - https://news.ycombinator.com/item?id=19223032 https://news.ycombinator.com/item?id=19223032 - Feb 2019 (1 comment) pifs - Avoid disk space usage by saving your files in the digits of Pi - https://news.ycombinator.com/item?id=18687275 https://news.ycombinator.com/item?id=18687275 - Dec 2018 (1 comment) πfs – A data-free filesystem - https://news.ycombinator.com/item?id=13869691 https://news.ycombinator.com/item?id=13869691 - March 2017 (105 comments) Πfs: Stores your data in π - https://news.ycombinator.com/item?id=10856108 https://news.ycombinator.com/item?id=10856108 - Jan 2016 (1 comment) Πfs: Never worry about data again - https://news.ycombinator.com/item?id=10847693 https://news.ycombinator.com/item?id=10847693 - Jan 2016 (1 comment) File system that stores location of file in Pi - https://news.ycombinator.com/item?id=8018818 https://news.ycombinator.com/item?id=8018818 - July 2014 (98 comments) 100% Compression Using Pi - https://news.ycombinator.com/item?id=6698852 https://news.ycombinator.com/item?id=6698852 - Nov 2013 (32 comments)
- based2 3y agohttps://en.wikipedia.org/wiki/The_Library_of_Babel https://en.wikipedia.org/wiki/The_Library_of_Babel