3 ms·
One of the biggest problems with extremely large files is how a simple insert or delete near the front of the file causes all bytes following it to be shifted a
by didgetmaster 4y ago
One of the biggest problems with extremely large files is how a simple insert or delete near the front of the file causes all bytes following it to be shifted and re-written to disk. Add 3 bytes to the beginning of a 50 GB file and you are writing 50 GB to disk.
I have been implementing a file system replacement project for several years. It is designed to handle hundreds of millions of files within a single container; put contextual meta-data tags on them; and enable lightning fast searches for things based off file type and/or tags. (https://www.youtube.com/watch?v=dWIo6sia_hw https://www.youtube.com/watch?v=dWIo6sia_hw)
One of the ideas (not yet fully implemented) was to break up large files at the file system level. You might have a 50 GB file of data that looks exactly like a normal file to any application accessing it, but in reality it might be 10 separate chunks of 5 GB each. If you add or delete any bytes within any individual chunk, it only adjusts that specific chunk. For example, deleting 100 bytes at offset 6 GB causes the second chunk to shrink by 100 bytes. All the chunks following it are unaffected. The file still looks to be 100 bytes smaller to the application, but it doesn't realize that a chunk in the middle was just reduced in size instead of all the bytes after the change being shifted down.
This feature would also make it easier to copy large files from one system to another. Data could be transferred one chunk at a time. If the copy was interrupted, only missing chunks would need to be copied when the process was restarted.
- themaninthedark 4y agoI know that BitTorrent has to allocate all the space before it downloads but would that be a good base to start with?
- nicolaslem 4y agoThis is also how Restic works, changing a few bytes in a 50 GB file won't reupload 50 GB of data unlike most backup solutions that work primarily with files.
- didgetmaster 4y agoThis change would not be just be for backups or file transfers. It would fundamentally change how big files are stored by the file system. Much like the way fragmentation is handled solely by the file system and an application accessing the file is completely unaware if the file is stored within a single fragment or multiple fragments; the file system would manage the individual 'chunks'. For example, a 6 GB file might be made up of 3 separate 2 GB chunks. An application might delete 20 bytes from the front of the file. This causes the first chunk to now be 2 GB - 20 bytes. The other 2 chunks are unchanged. Current file systems do not allow this where a file can have a block somewhere in its interior that is just a partial block.
- riceart77 4y ago> Current file systems do not allow this where a file can have a block somewhere in its interior that is just a partial block. Because it would be costly for uncertain benefit?
- rshaban 4y agocan I ask when this might need to happen ?
- kalleboo 4y agoEvery file system I know (even FAT32) supports file fragmentation and could do this (give or take block boundaries), but I don't know if there are any OS APIs to take advantage of that to actually let applications insert or remove data in the middle of a file. I'm assuming it's not in POSIX.