6 ms·
Most of the time it's much better (as in, faster) to just use cat (or pv, to get a nice progress bar) for writing a file to a block device, because it streams,
by voidz 10y ago
Most of the time it's much better (as in, faster) to just use cat (or pv, to get a nice progress bar) for writing a file to a block device, because it streams, and lets underlying heuristics worry about block sizes and whatnot.
So:
cat foobar.img > /dev/sdi
will stream the file rather than what dd does, i.e. read block, write block, read block, write block and so on.
Usually I also lower the vm.dirty_bytes and vm.dirty_background_bytes to 16 resp. 48 MB (in bytes) beforehand, which limits the buffer sizes to those amounts. Else it will seem that the progress bar indicates 300MB/s is written, and when it completes you still have to wait a really long time for things to have been written out.
Afterwards I restore back vm.dirty_ratio and vm.dirty_background_ratio to respectively 10 and 5 - the defaults on my system.
I wish that all of those projects, tutorials etc. that explain how to write their image to a block device, like an sdcard, would start advise using cat, because there is no reason to use dd, it's just something that people stick with because others do it too.
I only use dd for specific blocks, like writing back a backup of the mbr, or as a rudimentary hex editor.
- RJIb8RBYxzAMX9u 10y ago> I wish that all of those projects, tutorials etc. that explain how to write their image to a block device, like an sdcard, would start advise using cat, because there is no reason to use dd, it's just something that people stick with because others do it too. I'd wondered whether dd or cat were faster, and indeed cat is faster, but not by much. Also, for some embedded devices, you have to write to specific offsets, so dd is more convenient and explicit. Lastly, cat composes poorly with sudo. $ sudo dd if=foobar.img of=/dev/sdi # works $ sudo cat foobar.img > /dev/sdi # fails unless root b/c redirection is done by shell > I only use dd for specific blocks, like writing back a backup of the mbr, or as a rudimentary hex editor. xxd / xxd -r is much nicer, but I suppose sometimes vim is not available...
- greggyb 10y agosudo (cat foobar.img > /dev/sdi) No?
- nerdponx 10y agoWould sudo (cat foobar.img > /dev/sdi) Work?
- RJIb8RBYxzAMX9u 10y agoNo, unless you have a shell where sudo is a built-in, and I don't know of any.
- c0l0 10y agoThis is actually impossible (at least with UNIX shells), as shell builtins cannot have the SUID bit set - and your shellitself also shouldn't ;)
- JoshTriplett 10y agoYou could have a shell with a "sudo" builtin that knew how to invoke a separate "sudo" program with the right shell syntax and quoting, such that "sudo somecommand > /path/to/root-writable-file" did the right thing.
- abofh 10y agoNot really - to sudo, your shell would have to be setuid - and constantly fork stuff as you to get user permissions. Alternatively your shell could maintain a separate process for privileged access, but that puts a whole lot of your security on the assumption that your shell has no bugs that might allow escalation. In short, you could do it, but it'd be ripped out of every server that's been hardened, and for users that don't want to care - they're just running 'sudo su' anyhow. Speaking only for myself, the thought of my shell having a magical escalation process would scare the bejeezus out of me - and I'm supposed to have root on our boxes!
- TheCoelacanth 10y agoThe shell wouldn't need to be setuid if it just performed a syntax transformation and then called the existing sudo binary.
- mercora 10y agoone could use: cat foobar.img | sudo tee /dev/sdi > /dev/null
- JoshTriplett 10y agoI've used that to write short files (such as settings in /sys or /proc), but for large files, tee has the disadvantage of writing everything twice, and the pipe adds another write and read of every byte.
- mercora 10y agoI just realized this too and tried to close stdout instead. tee complains about it but goes on with its business: cat foobar.img | sudo tee /dev/sdi >&-
- RJIb8RBYxzAMX9u 10y agoWhy's everyone so against dd? :-) If you're going to use tee, might as well not bother with cat at all[0]: $ sudo tee /dev/sdi >&- < foobar.img Or better yet, pv[1]; you even get a progress bar that way! [0] https://en.wikipedia.org/wiki/Cat_%28Unix%29#Useless_use_of_cat https://en.wikipedia.org/wiki/Cat_%28Unix%29#Useless_use_of_... [1] http://www.ivarch.com/programs/pv.shtml http://www.ivarch.com/programs/pv.shtml
- mercora 10y agoThanks for the reminder :) I do fall quite often for the useless use of cat. But most of the time i also do not really care about it much. I did omit pv in believe tee will always be available but it is great and absolutely preferred when available.
- JoshTriplett 10y ago> I'd wondered whether dd or cat were faster, and indeed cat is faster, but not by much. That performance difference often comes from block size; "dd bs=1M" typically runs much faster than the default block size of 512 bytes.
- carussell 10y ago> xxd / xxd -r is much nicer, but I suppose sometimes vim is not available `od` is ubiquitous--it's POSIX and a requirement of the Single UNIX Specification
- na85 10y agoSudo is anachronistic anyways. Laptops don't need to follow multiuser best practices, and frankly a complex root password offers little.
- deleted 10y ago[deleted]
- deleted 10y ago[deleted]
- fragmede 10y ago> So: > cat foobar.img > /dev/sdi > will stream the file rather than what dd does, Sorry, that's not quite right. `cat` (and your shell, presumably bash) does the same fundamental thing as `dd`, ie, read block, write block. There's not really an underlying 'stream' primitive that `cat` (or `bash`, as you're using redirection to write the file) is using compared to `dd` What `cat` does do, is it does a better job of trying to find an optimal blocksize than a naive `dd` call does. `dd` simply defaults to a 512-byte block size, which is just inherited from history when 512-bytes was the alignment for, well, everything. There are numerous optimizations upon the fundamental read-block, write-block primitives to make it go faster (`cat` makes use of some of these). The linux kernel actually has a "stream file to socket" syscall to avoid the copy to user-land and back to "stream" a file out to the network, but that's not happening here, and there's still reading and writing of blocks in the kernel happening. See also: http://git.savannah.gnu.org/gitweb/?p=coreutils.git;a=blob;f=src/cat.c;h=001408576c4c17a156c1e8761ed2d9c96aa3d0cf;hb=HEAD#l167 http://git.savannah.gnu.org/gitweb/?p=coreutils.git;a=blob;f...
- LukeShu 10y agoThat's still not quite right. Once cat is spawned, the shell gets out of the way, and has nothing to do with it. (Exception: zsh will "helpfully" insert itself with `tee`-like operation in certain situations.) What `cat` doesn't do is finding an optimal block size for writes. It does a good job of finding optimal block size for reads, but not for writes. Many older block devices perform very poorly when the write block size is not optimal.
- omribahumi 10y agoThe fact that zsh is a parent doesn't necessarily hurt performance. The parent and child can share the stdin fd, and the child could be the only one reading from it.
- LukeShu 10y agoThat's not what I was referring to. Though I should have been more clear: I don't believe that zsh inserts itself in the mentioned situation; just that it does insert itself in some situations, and that makes it an exception to my general "once the process is spawned, the calling shell doesn't matter" statement. As an example of what I was referring to, in zsh, this a >&2 | b is equivalent to Bourne shell: a | tee /dev/stderr | b except that `tee` is implemented as part of zsh, rather than a separate program (for no difference in performance). That is, zsh inserted itself into the middle of the pipeline. There are two pipes instead of one; zsh/tee reads from the a pipe, and then writes that data to the b pipe (and stderr). This does hurt performance.
- mbrumlow 10y agoThis is all wrong. I am not even sure what you mean by stream or how it would be any difference than "read", "write", "read", "write". Because that is literally what cat does, and so does dd. Many people feel cat is faster than dd because dd's default block size is 512 bytes as you can see from a simple dd command. strace dd if=10M of=junk read(0, "\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0"..., 512) = 512 write(1, "\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0"..., 512) = 512 read(0, "\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0"..., 512) = 512 write(1, "\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0"..., 512) = 512 read(0, "\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0"..., 512) = 512 write(1, "\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0"..., 512) = 512 strace cat 10M > junk2 read(3, "\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0"..., 131072) = 131072 write(1, "\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0"..., 131072) = 131072 read(3, "\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0"..., 131072) = 131072 write(1, "\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0"..., 131072) = 131072 read(3, "\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0"..., 131072) = 131072 write(1, "\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0"..., 131072) = 131072 So when benchmarking dd vs cat make sure you are comparing apples to apples. The beauty of dd is you don't have to tinker with things like vm-.dirty_bytes and vm.dirty_background_bytes before using a command. Those should not be messed with on most systems and should definitely not be messed with for the sake of running a single command. When you use cat, it makes some decisions for you. Most of the time makes good choices but it happens to make not so optimal choices for working with large files being written to other devices or file systems. To avoid the pitfalls of using cat to copy image files (and not have to change global system settings) you use dd. With dd you can use O_DIRECT to bypass the VFS caching layer so your vm.dirty_* settings are ignored in general. You can also specify the block size optimal for the device you are writing to and even reading from. While dd may not stand for disk or drive or its name have anything to do with disk it is a powerful tool that is the appropriate tool to use when working with large files, precise movement or data and yes copying image files to devices.
- quotemstr 10y agoAlmost, but not quite. You really want buffer in there to smooth out IO time variances. <foo.img pv -cNsrc | buffer -p75 -s256k -b128 | pv -cNdst >/dev/whatever buffer lets the rest of the pipeline proceed while pv is blocked writing to the output file. Yes, conventional filesystem readahead helps to some extent, but IME, not enough, especially when foo.img is a block device itself.
- wpietri 10y agoHoly moly. I have spent a lot of years typing Unix commands, but it never occurred to me to put the input redirection first. But that's much more pipeline-ish, so I like it a lot.