7 ms·
Comparing dd and cp when writing a file to a thumb drive
- bravetraveler 3y agoOne of my favorite party tricks is using shell redirects to do this. Using coreutils ('cp') is a nice approach I'm all for anything not 'dd' or cargo-culted like balenaEtcher This all depends on a certain format of ISO that I can't recall There's no magic. Checksum the drive and the file after
- mort96 3y agoUsing shell redirects sucks though, you need to run the shell as root. Using cp seems overkill? It's really really designed to copy files between file systems and has a lot of logic to handle different cases and different optimisations which don't matter when writing to block devices. As you say, there's no magic. I want a program which just calls 'open' on a path I give it and then uses 'write' to write my data to it. cp does so, so much more related to file systems, shell redirects aren't a separate program I can run with sudo, but I can trust dd to do the job. And it has a (kinds bad) progress monitor to boot. The default 512 byte block size is unfortunate though.
- GuB-42 3y agoAs far as I know, you can't copy using shell redirects. At least not with a simple POSIX shell. Redirection only open file descriptors, but you need a command to actually make the copy from one to the other. It can be "cat", which is often cargo culted in it own right. Or even "dd", which can work with stdin/out. I usually prefer not to use shell redirects if there is a command that take filenames. That's because it gives more control to the app. The app knows how the file will be used, the shell doesn't, so it can open it the most appropriate way, avoid overwriting the output file if something goes wrong, output better error messages, etc... Now, if you don't trust the app (for example if you fear it will modify your input file), then shell redirects may be the better option.
- Brian_K_White 3y agoYou can absolutely read and write by just redirects. No cat or dd or anything needed.
- LegionMammal978 3y agoCould you give an example? Like GP, I'm also not aware of any way to do such a thing without using shell-specific extensions.
- Brian_K_White 3y agoI confess it's not as simple as $ <a >b Sorry about that, because it means the ansdwer while technically true in some weird case where you really need it, isn't exactly convenient like you'd actually use it. I made it sound trivial and obvious and direct and it's not. Use read in a loop, with special care with LANG and IFS to make all bytes meaningless. Except there is no way to avoid null being special, but you can handle null by making null the delimiter for read, and only reading one byte at a time. So even though you can't store an actual null in a variable, you can still detect that there was a null and print a new one back out, and since you only read one byte at a time, you do that for each individual input byte and strings of nulls are not collapsed. It looks like a lot, but, read is a builtin, and at least in bash and ksh and zsh so is printf, and although this is a loop, it's actually not even a sub-shell. If you edit variables inside the loop, they are still there after the loop, ie, you never forked a child. while LANG=C IFS= read -d '' -r -n 1 x ;do printf '%c' "$x" ;done <junk1.rnd >junk2.rnd
- bravetraveler 3y agoThis works in ZSH but not BASH... I have a feeling you just helped me see why :D
- js2 3y agoYou're right. You'd have to use the `read` built-in and `echo` or `printf` in a while loop combined with redirection, but the POSIX-specified `read` built-in is intended for text and isn't going work correctly with binary files. With bash you could maybe use the non-standard `read -N` option but I'm not sure about null bytes.
- tadfisher 3y agocp queries the preferred block size of the destination file in 'struct stat', and has specific tweaks for certain filesystems. As far as I can tell, dd does not do this as it calls through to 'write' directly. In any case, the tests in ioblksize.h indicate that bs=4M is far too large and may perform worse than the default for cp/cat (128KiB). There is a script there that should clear things up for more modern systems. The point about fdatasync is superfluous as you can run 'sync' yourself, or unmount the filesystem.
- NoZebra120vClip 3y agoThis example was performed on the block-special device, so there is no filesystem. dd uses a default block size of 512 bytes, according to the manual page: calling write(2) directly means you need to choose a buffer of some size. "bs=" sets both input and output block sizes, which probably isn't the best idea in this case. Block sizes are a tricky subject: https://utcc.utoronto.ca/~cks/space/blog/unix/StatfsPeculiarities https://utcc.utoronto.ca/~cks/space/blog/unix/StatfsPeculiar... https://utcc.utoronto.ca/~cks/space/blog/tech/SSDsAnd4KSectorsII https://utcc.utoronto.ca/~cks/space/blog/tech/SSDsAnd4KSecto...
- rollcat 3y agoI've been re-implementing a bunch of coreutils as an exercise, and got stuck on dd input/output block sizes, AND disk/partition block sizes for a while. (As far as I understand it, for dd I need a ring buffer the size of max(ibs, obs), and then some moderately clever book-keeping to know when to trigger the next read/write, perhaps with code specific to ibs>obs, ibs<obs, etc; partitioning on the other hand is plainly stupid, there's decades of hardware and software just lying to each other and nothing makes sense.) Thank you and everyone else in this thread for the know-how and references! I would like to eventually write an article (or at least heavily commented source) to hopefully explain all this nonsense for other people like me.
- mort96 3y agoTFA mentions that you can use sync after cp to do the same thing as fdatasync. You can't "unmount the file system" because you're writing directly to the block device, the thumb drive isn't mounted.
- codetrotter 3y agoSpeaking of thumb drives. In the case of copying files to a mounted file system, I’ve sometimes found it faster to use a tar pipeline than cp when copying data to an USB stick or SD/microSD card. Instead of: cp -r ~/wherever/somedir/ /media/SOMETHING/ I would do cd ~/wherever/ tar cf - somedir/ | ( cd /media/SOMETHING/ && tar xvf - ) And it would be noticably faster. Not the same use case as linked article, but wanted to bring this up since it’s somewhat related.
- xorcist 3y agoThere's also rsync which works perfectly with local paths and can resume from interruptions (by default), with or without crc checking the material ("-c"). That can be useful for removable storage which can sometimes be a bit unreliable. Just take care that cp, tar and rsync each have slightly different handling of extended attributes and sparse files. (By the way, I believe "tar -C /path" is the canonical way of doing "cd /path ; tar" without resorting to subshells.)
- codetrotter 3y agorsync is a strange beast to me First thing that makes it so weird is how it assigns different meaning between paths including vs not including trailing slash. Completely different from how most command line tools I am used to behave in Linux and FreeBSD. That alone is enough to remind me every time I try to use rsync why I don’t like and generally don’t use rsync.
- pastage 3y agoLost too much data that way, trailing slash with delete. I still feel bitter, the UX is terrible considering the effect are so different from copy and delete.
- pessimizer 3y agoIf it didn't do that, it would have to add a switch that you'd still have to look up. It's a tool that's most often doing a merge and update, but looks like a copy command. I think that made it friendlier. How would you separate "merge this directory into that one and update files with the same name" from "copy this directory into that directory"?
- BenjiWiebe 3y agoI've switched to the following: <infile pv >outfile And it seems to work very well plus gives a good progress indicator.