6 ms·
The issue with a client app backing up dropbox and onedrive folders on your computer is the files on demand feature, you could sync a 1tb onedrive to your 250gb
by Neil44 6mo ago
The issue with a client app backing up dropbox and onedrive folders on your computer is the files on demand feature, you could sync a 1tb onedrive to your 250gb laptop but it's OK because of smart/selective sync aka files on demand. Then backblaze backup tries to back the folder up and requests a download of every single file and now you have zero bytes free, still no backup and a sick laptop.
You could oauth the backblaze app to access onedrive directly, but if you want to back your onedrive up you need a different product IMO.
- danpalmer 6mo agoThis is a complexity that makes it harder, but not insurmountable. It would be reasonable to say that if you run the file sync in a mode that keeps everything locally, then Backblaze should be backing it up. Arguably they should even when not in that mode, but it'll churn files repeatedly as you stream files in and out of local storage with the cloud provider.
- bayindirh 6mo ago> Arguably they should even when not in that mode, but it'll churn files repeatedly as you stream files in and out of local storage with the cloud provider. When you have a couple terabytes of data in that drive, is it acceptable to cycle all that data and use all that bandwidth and wear down your SSD at the same time? Also, high number of small files is a problem for these services. I have a large font collection in my cloud account and oh boy, if I want to sync that thing, the whole thing proverbially overheats from all the queries it's sending.
- vladvasiliu 6mo agoBut if the files are only on the remote storage and not local, chances are they haven't been modified recently, so it shouldn't download them fully, just check the metadata cache for size / modification time and let them be if they didn't change. So, in practice, you shouldn't have to download the whole remote drive when you do an incremental backup.
- bayindirh 6mo agoYou can't trust size and modification time all the time, though mdate is a better indicator, it's not foolprooof. The only reliable way will be checksumming. Interestingly, rclone supports that on many providers, but to be able to backblaze support that, it needs to integrate rclone, connect to the providers via that channel and request checks, which is messy, complicated, and computationally expensive. Even if we consider that you won't be hitting API rate limits on the cloud provider.
- NetMageSCW 6mo agoIf you can’t trust modification time you are doing something so unusual that you probably need to be handling your backups privately anyway.
- bayindirh 6mo agoI don't think so. Sometimes modification time of a file which is not downloaded on computer A, but modified by computer B is not reflected immediately to computer A. Henceforth, backup software running on computer A will think that the file has not been modified. This is a known problem in file synchronization. Also, some applications modifying the files revert or protect the mtime of the file for reasons. They are rare, but they're there.
- jtbayly 6mo agoReading your comments, it sounds like you are arguing it is impossible to backup files in Dropbox in any reasonable way, and therefore nobody should backup their cloud files. I know you haven’t technically said that, but that’s what it sounds like. I assume you don’t think that, so I’m curious, what would you propose positively?
- bayindirh 6mo ago> I know you haven’t technically said that, but that’s what it sounds like. Yes, I didn't technically said that. > It sounds like you are arguing it is impossible to backup files in Dropbox in any reasonable way, and therefore nobody should backup their cloud files. I don't argue neither, either. What I said is with "on demand file download", traditional backup software faces a hard problem. However, there are better ways to do that, primary candidate being rclone. You can register a new application ID for your rclone installation for your Google Drive and Dropbox accounts, and use rclone as a very efficient, rsync-like tool to backup your cloud storage. That's what I do. I'm currently backing up my cloud storages to a local TrueNAS installation. rclone automatically hash-checks everything and downloads the changed ones. If you can mount Backblaze via FUSE or something similar, you can use rclone as an intelligent MITM agent to smartly pull from cloud and push to Backblaze. Also, using RESTIC or Borg as a backup container is a good idea since they can deduplicate and/or only store the differences between the snapshots, saving tons of space in the process, plus encrypting things for good measure.
- nine_k 6mo agoThis. You should not try to backup your local cache of cloud files as if those were your local files. Use a tool that talks to the cloud storage directly. Use tools with straightforward, predictable semantics, like rclone, or synching, or restic/Borg. (Deduplication rules, too.)
- mroche 6mo agoMy understanding of Backblaze Computer Backup is it is not a general purpose, network accessible filesystem.[0] If you want to use another tool to backup specific files, you'd use their B2 object storage platform.[1] It has an S3 compatible API you can interact with, Computer Backup does not. But generally speaking, I'd agree with your sentiment. [0]: https://www.backblaze.com/computer-backup/docs/supported-backup-data https://www.backblaze.com/computer-backup/docs/supported-bac... [1]: https://www.backblaze.com/docs/cloud-storage-about-backblaze-b2-cloud-storage https://www.backblaze.com/docs/cloud-storage-about-backblaze...
- Chaosvex 6mo agoThen do it in memory, assuming those services allow you to read the files like that. It sounds like they do based on your other comments.
- bayindirh 6mo agoThe problem is, downloading files and disk management is not in your control, that part is managed by the cloud client (dropbox, google drive, et. al) transparently. The application accessing the file is just waiting akin to waiting for a disk spin up. The filesystem is a black box for these software since they don't know where a file resides. If you want control, you need to talk with every party, incl. the cloud provider, a-la rclone style.
- NetMageSCW 6mo agoWhy would they do new backups of old files all the time? They would just skip those.
- Dylan16807 6mo agoUnless it does something very weird it won't trigger all those files to download at the same time. That shouldn't be a worry. And, as a separate note, they shouldn't be balking at the amount of data in a virtualized onedrive or dropbox either considering the user could get a many-terabyte hard drive for significantly less money.
- bayindirh 6mo ago> Unless it does something very weird it won't trigger all those files to download at the same time. That shouldn't be a worry. The moment you call read() (or fopen() or your favorite function), the download will be triggered. It's a hook sitting between you and the file. You can't ignore it. The only way to bypass it is to remount it over rclone or something and use "ls" and "lsd" functions to query filenames. Otherwise it'll download, and it's how it's expected to work.
- Dylan16807 6mo agoWhy would it use either of those on all the files at once? It should only be opening enough files to fill the upload buffer.
- bayindirh 6mo agoMaybe it'll, maybe it won't, but it'll cycle all files in the drive and will stress everything from your cloud provider to Backblaze, incl. everything in between; software and hardware-wise.
- Dylan16807 6mo agoThat sounds very acceptable to get those files backed up. It shouldn't stress things to spend a couple weeks relaying a terabyte in small chunks. The most likely strain is on my upload bandwidth and yeah that's the cost of cloud backup, more ISPs need to improve upload.
- 6mo ago
- appreciatorBus 6mo agoShoutout to Arq backup which simply gives you an option in backup plans for what to do with cloud only files: - report an error - ignore - materialize Regardless, if you make it back up software that doesn’t give this level of control to users, and you make a change about which files you’re going to back up, you should probably be a lot more vocal with your users about the change. Vanishingly few people read release notes.
- decadefade 6mo agoLove Arq!
- Lord_Zero 6mo agoWhy no linux support?
- CamperBob2 6mo agoIf it's open-source, Linux support is only a few hours with Claude away. If it's not open-source, but the protocol is documented, see above. If it's not open-source, and the protocol isn't documented, well... that makes the decision easy, doesn't it?
- cobertos 6mo agoBackups software written by Claude? No thanks. I've used enough Claude coded applications that I wouldn't trust that with a backup, unless it had extensive tests along with it.
- bastawhiz 6mo agoThat doesn't really make a lot of sense, though. Reading a file that's not actually on disk doesn't download it permanently. If I have zero of 10TB worth of files stored locally on my 1TB device, read them all serially, and measure my disk usage, there's no reason the disk should be full, or at least it should be cache that can be easily freed. The only time this is potentially a problem is if one of the files exceeds the total disk space available. Hell, if I open a directory of photos and my OS tries to pull exif data for each one, it would be wild if that caused those files to be fully downloaded and consume disk space.
- bombcar 6mo agoIt's generally now handled decently well, but with three or four of these things it can make backups take annoying long as without "smarts" (which are not always present) it may force a download of the entire OneDrive/Box each time - even if it never crashes out.
- bastawhiz 6mo ago> it may force a download of the entire OneDrive/Box each time - even if it never crashes out. I am not aware of any evidence supporting this.
- jrmg 6mo agoRight, but even if that’s working it breaks the user experience of services like this that ‘files I used recently are on my device’. After a backup, you’d go out to a coffee shop or on a plane only to find that the files in the synced folder you used yesterday, and expected to still be there, were not - but photos from ten years ago were available!
- NetMageSCW 6mo agoThere’s no reason to think that would happen - files you had from ten years ago would have been backed up ten years ago and would be skipped over today.
- thecapybara 6mo agoThat would make sense for online-only files, but I have my Dropbox folder set to synchronize everything to my PC, and Backblaze still started skipping over it a few months ago. I reached out to support and they confirmed that they are just entirely skipping Dropbox/OneDrive/etc folders entirely, regardless of if the files are stored locally or not.
- catlikesshrimp 6mo agoat least you can fix (bandaid) that by cloning your Dropbox folder to anotherfolder. Double taken space for the greater goodx
- whimblepop 6mo agoThe whole "just sync everything, and if you can't seek everything, pretend to sync everything with fake files and then download the real ones ad-hoc" model of storage feels a bit ill-conceived to me. It tries to present a simple facade but I'm not sure it actually simplifies things. It always results in nasty user surprises and sometimes data loss. I've seen Microsoft OneDrive do the same thing to people at work.
- carefulfungi 6mo agoI’ve lost data not realizing I was backing up placeholder files (iCloud). Hiding the network always ends in pain. But never goes out of style.
- whimblepop 6mo agoMy own approach to simplicity generally means "hide complexity behind a simple interface" rather than pushing for simple implementations because I feel that too much emphasis on simplicity of implementations often means sacrificing correctness. This particular example is a useful one for me to think about, because it's a version of hiding complexity in order to present a simple interface that I actually hate. (WYSIWYG editors is another one, for similar reasons: it always ends up being buggy and unpredictable.)
- eblume 6mo agoSame. I lost a lot of photos this way. I've recently moved over to Immich + Borg backup with a 3-2-1 backup between a local synology NAS and BorgBase. Painful lesson, but at least now I feel much more confident. I've even built some end-to-end monitoring with Grafana.
- chrisweekly 6mo agoCareful w that Synology NAS, mine's now a brick that may also have led to permanent data loss.
- 6mo ago
- signorovitch 6mo agoThe primary trouble I have with backblaze was that this change was not clearly communicated, even if perhaps it could be justified.
- ineedasername 6mo agoThat seems like a pretty straightforward issue to solve, to simply backup only those files that are actually on the system, not the stubs. If it's on your computer, it should able to get backed up. If it's just a shadow, a pointer, it doesn't. Making the change without making it clear though, that's just awful. A clear recipe for catastrophic loss & drip drip drip of news in the vein of "How Backblaze Lost my Stuff"
- lazide 6mo agoThe stubs are the thing on your computer?
- coldtea 6mo agoImagine if they could detect stab or real file huh? Space technology, I know! Or just fucking copy them as stubs and what's actually downloaded as actually downloaded! Boggles the mind! Or maybe just do what they do now, but WARN about that in HUGE RED LETTERS, in the website and the app, instead of burying it in an update note like weasels!
- wrs 6mo agoThe OP’s complaint is that the files were not backed up. If they had discovered that only stubs were backed up, I don’t think they’d be any happier.
- ineedasername 6mo agoNot what I meant: The other cloud storage services connected to the computer, eg OneDrive. Those files, when they are just stubs. I'm saying that Backblaze could simply not backup stubs, if the person isn't syncing the actual file to their drive. If they are, backblaze should back it up.
- wrs 6mo agoThe point of the stubs is so you don’t have to know which cloud files are actually on your device at any given moment, because they will be fetched automatically. Marking files as “always local” is a niche feature, and in any case has nothing to do with whether you want those files backed up. To have this thing you’re not supposed to need to worry about affect whether your files got backed up is exactly the problem here. The goal is to back up your files, whether they’re in the cloud or not. I sympathize with Backblaze’s problem with their file change monitor, but then they should considee implementing connectors for OneDrive, Dropbox, etc. and back up files directly from the cloud.
- bombcar 6mo agoThe issue really isn't that it's not backing up the folder (which I can see an argument for both sides and various ways to do it) - it's that they changed what they did in a surprising way. Your backup solution is not something you ever want to be the source of surprises!
- downrightmike 6mo agoThe fault is with the PC manufacturers screwing you on disk space claiming 1TB, when its only 256gb. bait and switch
- tencentshill 6mo agoCloud placeholders have been a feature for years, plenty of programs have mitigations for this behavior.