19 ms·
Useless Use of "dd" (2015)
- booi 4y agoMind.. blown..
- LeoPanthera 4y agoOccasionally I run a dd on a loop with increasing block sizes to see what is actually fastest. I regularly see instructions on the web saying you should use "1M" or even "4M", but in my tests, smaller block sizes are often faster. A few years ago, "128K" was usually the fastest choice. Today, on faster systems, "512K" has a slight edge. I could not tell you why, though. Try it for yourself.
- usr1106 4y agoIf you wirte to USB sticks / SD cards you should use (a multiple of) of the erase block size in oder not wear out the flash. I don't know what somewhat current sizes are. But 1M sounds likely to be a reasonable multiple. For USB also make sure to power it off before removing otherwise you might lose data. Eg. by udisksctl power-off -b /dev/sdX
- yjftsjthsd-h 4y agoHuh. I always run `sync` (twice, because Tradition™) before yanking a USB stick, and I've never noticed issues; any idea if powering off is better?
- zamadatix 4y agoSync just syncs but it doesn't prevent anything from starting to write while you go to pull it. "udisksctl power-off" checks nothing is using the drive, commits buffers to storage, deconfigures the drive, and powers it off. Unfortunately it may kill more devices than you want in some scenarios due to the way killing the port works. Also for normal USB drives I don't think® poweroff is any different than an immediate physical pull post unmounting. The upside is it'll be completely disabled the instance the command completes, i.e. it can't even be written to raw as it is no longer powered even though plugged in. Not sure how helpful that is in reality though. I usually just umount the drive which syncs and prevents further writes (well, to the mount at least) but doesn't try to kill it for me. Most likely sync is "good enough" for the vast majority of use though.
- usr1106 4y agoI learned in the 1980s that it's sync; sync; sync I feel more comfortable with just unmounting and powering off nowadays. But I have no hard evidence. Hardly ever user more than one USB drive same time, so not sure whether there could be side effects with powering off one of them. Well, USB being USB, nothing would surprise me too much.
- viraptor 4y ago> Want to simulate a lseek+execve? Use dd! I can't find anything about it in the man page - does anyone know the options? I'm assuming it doesn't mean simply running dd with seek to output another file and then running that.
- koala_man 4y agoOriginal author here. An example would be `{ dd bs=1 count=0 skip=1234; myprogram; } < file`. This is similar to C `lseek(0, 1234, SEEK_CUR); execve("myprogram", ...);` It's not something you'd realistically use, but it's better to be limited by imagination than tooling.
- viraptor 4y agoAh, I get it! What I initially thought you meant was creating a process from an ELF embedded further into the file. For example lseek-ing into a tar file and somehow execve-ing an included binary. Not I wonder if Linux you allow me to run /dev/stdin...
- jlmorton 4y agodd if=/dev/zero of=/proc/$(pidof nginx)/fd/0
- guerrilla 4y agocat /dev/zero > /proc/$(pidof nginx)/fd/0
- naikrovek 4y agoOh thank you for submitting this to HN. I’ve been telling people not to use dd for years and everyone looks at me like I just gave birth to a full grown dinosaur or something. “Well why does the entire internet say to use dd then?” Because they copy from each other just like you copied from them. Just use cat.
- HomeGear 4y ago>“Well why does the entire internet say to use dd then?” Because they copy from each other just like you copied from them. Great statement. This brings me some much needed internal clarity on my own thoughts and actions.
- chubot 4y agoThis is basically the parable of "Grandma's Ham" https://www.executiveforum.com/cutting-off-the-ends-of-the-ham/ https://www.executiveforum.com/cutting-off-the-ends-of-the-h... tl;dr nobody in 2 generations knows why they cut the ends off the ham before cooking it, until they talked to grandma, who said her pan was too small A Unix thing that's been posted to HN for a decade, that's almost literally the same story: Understanding the bin, sbin, usr/bin , usr/sbin split https://news.ycombinator.com/item?id=3519952 https://news.ycombinator.com/item?id=3519952 tl;dr /usr/bin is separate from /bin because someone had a small hard disk once
- ahartmetz 4y agoSo Grandma's Ham is a variant of Chesterton's Fence... When you do or don't do something for reasons of tradition, find the real reasons for doing / not doing it.
- soneil 4y agoYou’d be amazed how many people I’ve found that do ‘tar xzf filename’ without knowing what xzf is doing.
- jthrowsitaway 4y ago
- aib 4y agoI use it for the O_DIRECT flag (oflag=direct). Better progress report, and I know when to pull the stick. Edit: Looks like O_SYNC (oflag=sync) is needed, too. Should update my 'sddd' alias.
- Retr0id 4y agoIs it? I thought O_DIRECT implied O_SYNC, in practical terms?
- aib 4y agoSo did I, but my man page (open(2)) cast doubts: "The O_DIRECT flag on its own makes an effort to transfer data synchronously, but does not give the guarantees of the O_SYNC flag that data and necessary metadata are transferred." While I don't think this is relevant to block devices, I see no harm in including the flag in my alias, either.
- enasterosophes 4y agoSo I should stop using dd as a text editor, is that what you're trying to tell me?
- sramsay 4y agoYou move like them. Are you the chosen one?
- yjftsjthsd-h 4y ago> using dd as a text editor Bravo, you've found something (marginally) more upsetting than using ed:)
- hyperdimension 4y agoyes 'y' | dd of=response bs=1 count=1 yes 'e' | dd of=response bs=1 count=1 seek=1 yes 's' | dd of=response bs=1 count=1 seek=2
- jws 4y agodd lets you specify the write block size. This is essential when writing to 9 track tape. Everything is a byte stream… except for the things which are not.
- naikrovek 4y agoI don’t think the author is intending to tell everyone to always use something else, I think the author is trying to tell most people who use dd that there are easier ways to do what they are trying to do, and that things are files, which is something that a lot of people here seem to be forgetting.
- tjoff 4y agoWhich is like saying vim is useless, 90% of what you'll do is covered by nano. Well, it pays off a lot to know a complex tool and using it for easy stuff gets you in the habit of using it, reminds you of the arguments etc. dd is really really useful.
- naikrovek 4y agoyes, it is. and if you know you need it, you know you need it. and if you don't know that, you don't need it.
- mark-r 4y agoI had almost managed to completely forget about 9 track tapes. Now you've made my head hurt.
- throwawayboise 4y agoI used to used dd to convert EBCDIC to ASCII reading tapes from a 1/2" reel-to-reel tape drive. The "convert" capability of dd is what differentiates it from utilities such as cat.
- spudlyo 4y agoIn the bad old Linux days, if you read a huge amount of data off the disk (like if you were doing a backup) Linux would try to take all that data and jam it in the page cache. This would push out all the useful stuff in your cache, and sometimes even cause swapping as Linux helpfully swaps out stuff to make room for more useless page cache. One of the great things about `dd` is that you have a lot of control how the input and output files are opened. You can bypass the page cache when reading data by using iflag=direct, which would stop this from happening.
- guerrilla 4y agoI think using a small bs also determines the size of the cache you use, as its the buffer.
- KMag 4y agoGP is talking about the Linux kernel's buffer cache. Unless you tell the kernel to operate directly on the disk, your reads come from and your writes go to pages within the kernel's buffer cache. Using a small bs probably results in a buffer of only bs bytes inside dd's address space, but the buffer cache is completely different and resides in the kernel. That is, without iflag=direct, dd will repeatedly ask the kernel to copy bs bytes from the kernel's buffer cache into dd's address space, and then ask the kernel to copy bs bytes from its address space into the kernel's buffer cache.
- cosmotic 4y agoIt's great to have control over this but I suspect most users never knew this was happening, had no idea dd could bypass that behavior, nor knew which argument to pass to dd to accomplish this. It's like saying 'what makes 3d printers so great is you can make anything!' but you'd be way better off with an industrially forged object than the 3d printed object.
- birdyrooster 4y agoright and good luck doing injection or vacuum molding for a few copies of a niche design Edit: correcting for nitpicker
- indigodaddy 4y agoNot exactly dd, but in my NOC tech days, ddrescue was _indispensable_ for cloning drives with pending sectors/errors.
- mgerdts 4y agoWhile there’s a lot of truth here, there are times when you can do much better than cat. A while back I tweeted: Today I found the magical dd command that causes an NVMe drive to run at almost full speed: # dd if=/dev/nvme3n1 bs=4096k iflag=direct of=/dev/null status=progress 959090524160 bytes (959 GB, 893 GiB) copied, 178 s, 5.4 GB/s The trick to getting this throughput is telling Linux to do an insanely large IO (4 MiB). The drive can't do 4 MiB reads - the largest IO it can handle is 2 MiB. # nvme id-ctrl /dev/nvme3 | grep mdts mdts : 9 # echo '4 \* 2^9' | bc -l 2048 More in the tread starting here: https://twitter.com/OMGerdts/status/1514376206082269191?s=20&t=6YcJZRCQIgeV4m1DU5OvRQ https://twitter.com/OMGerdts/status/1514376206082269191?s=20...
- jagrsw 4y agoMy SSD - apparently one of the fastest on the market - KINGSTON SKC3000D2048G - rated 7GB/7GB read/write, after writting data to it, stopped being very fast. No idea, if it's something related to how nvme/ssd-s work, or maybe I have a broken unit. $ sudo dd if=/dev/nvme0n1 bs=4096k iflag=direct of=/dev/null status=progress ... 9667870720 bytes (9,7 GB, 9,0 GiB) copied, 20,2714 s, 477 MB/s But when reading from empty space (as in, non-written yet blocks), it goes at full PCIE x4 Gen4 speeds. $ sudo dd if=/dev/nvme0n1 bs=4096k iflag=direct of=/dev/null status=progress skip=450000 ... 16710107136 bytes (17 GB, 16 GiB) copied, 2,42802 s, 6,9 GB/s I have another nvme drive - Force MP510 - and it doesn't care if data was previously written or not. When reading from it, I get ~full x4/Gen3 speeds of 3.5GB/s PS: nvme smart-log shows 100% available spare, and 0% percentage_used, so it doesn't seem to be wear-related.
- dijit 4y agocopying zeroes is much less work for the controller and SLC cache. is it possible that the controller is over-heated? Do you have the NVME drive under a GPU?
- nousermane 4y ago> copying zeroes is much less work Yes, reading "evicted" or "trimmed" space [0] specifically, is much less work - only flags are read, not actual media. [0] https://en.wikipedia.org/wiki/Trim_(computing) https://en.wikipedia.org/wiki/Trim_(computing)
- dredmorbius 4y agoUseful uses of dd(1). There remain some useful applications of dd. These may of course be achieved by other mechanisms, but typically less conveniently. 1. Read a specific number of blocks or bytes from a source: dd if=/dev/hda of=/root/mbr bs=512 count=1 This will make a copy of, say, your master boot record (first 512 bytes of your first disk drive) and stash it in your /root directory. 2. Read from specific bytes of a file dd if=mydata skip=1k bs=32 count=1 Reads 32 bytes after the first 1024 (1k) bytes of "mydata". 3. Write to specific bytes of a file dd if=source of=target seek=10k bs=512 count=1 conv=notrunc That should write 512 bytes from "source" beginning 10k into "target". (I've not tested this, you should verify.) 4. Create a sparse file. Sparse files appear to have a nonzero size, but take up no space on disk, until data is actually written to them. These are often used as "inflating" dynamic filesystem images for virtual machines. dd if=/dev/zero of=sparsefile bs=1 count=0 seek=20000M # Create 20 GB sparse file 5. Case conversions. Sure, you could use tr(1), but where's the sport? dd if=MixEdCaSE of=lcase conv=lcase # Convert to lower case dd if=MixEdCaSE of=ucase conv=ucase # Convert to upper case 6. ASCII / EBCDIC conversions dd if=ebcdic of=ascii conv=ascii # ebcdic -> ascii dd if=ascii of=ebcdic conv=ebcdic # ascii -> ebcdic When reading to or from IBM data tapes, you might find blocking / unblocking conversions useful. I've done this, but it's so long ago that I don't trust my memory on that any more. Odds are good you'll not have to worry about this. There are other useful applications as well, though these are not typically encountered very often. Do feel free to explore and attempt these on safe media.
- minedwiz 4y agohttps://web.archive.org/web/20220523020952/https://www.vidarholen.net/contents/blog/?p=479 https://web.archive.org/web/20220523020952/https://www.vidar... Archive link of this; at least for me, the original site isn't responding.
- cc101 4y agoBack in the Dark Ages (1968) we had pre-written JCl scripts. I don't think many people knew JCL. We just appended a script to the front of our Fortran card decks. A DD card was the final card. After the frustrating work of finally getting the JCL right, the DD card was always just the "Do it Damn it" card in my mind.
- liftm 4y agoI see some useless use of pv there. dd status=progress is neat.
- Nux 4y agoThat's a relatively new feature btw. The lines of rhel 6 did not have that for example, afaicr.
- aftbit 4y agoI mostly use `gzip -d image.gz | dd of=/dev/sda bs=1M oflag=direct status=progress` to write a compressed disk image to a slow USB stick with progress bar and no disk caching. This avoids the lengthy wait at the end of a typical `cp` for the sync before unplug.
- salmo 4y agoIt’s funny (given that the name likely is inspired by Useless Use of ‘cat’) that examples here have actual useless uses of ‘cat’. ‘cat file | pv > disk’ can just be ‘pv < file > disk’. But whatever. I’m sure 90% of code stems from something someone read and copied or came most easily to them. It works. ‘dd’ was really useful for finicky media like tapes and doing EBCDIC translation. It’s still great when you combine bs and count. Blow away an MBR, make an xGB file, etc. It’s a Swiss Army knife. It can do a lot, but isn’t the best tool for most things. I still love it. Probably just muscle memory.
- hackmiester 4y agoOr just `pv file > disk` ... and then pv will give you a progress bar automatically
- salmo 4y agoOh, good point. When pv sees the actual file vs stdin, it can give % progress without hints. Which is nice. And it’s nicer that the recent ‘dd status=progress’ (if I remember that option right).
- koala_man 4y agoOriginal author here. The intended point was that `dd if=.. | something | dd of=..` is just as useless as the two cats in `cat .. | something | cat > ..`. People sometimes mock the latter while still believing that the former is necessary.
- salmo 4y agoSorry, that came off snarkier than I meant rereading it. You’re 100% correct. I’ve grown more tolerant of these ceremonial uses. There really aren’t many folks that understand shell. I see it in CICD, etc all the time, too. The one left that kills me that I see in vendor scripts all the time is ‘command; ret=$?; if [ “$ret” -ne 0 ]; then foo; fi’ If you’re handling return code cases, then cool. But folks don’t realize ‘test’ aka ‘[‘ is just another command. But then again, so much Java, etc. I read is the same copy/paste. Those can be worse because example code isn’t prod ready, where bad shell usually does the right thing, just awkwardly or inefficiently. But as you point out, so much isn’t thinking about what you’re doing, just mimicking.
- yuuta 4y agoUseless use of dd: dd if=/path/to/file of=/dev/stdout
- rsync 4y ago'dd' over ssh: mysqldump -u mysql db | ssh user@rsync.net "dd of=db_dump" pg_dump -U postgres db | ssh user@rsync.net "dd of=db_dump" (not useless, though ...)
- augusto-moura 4y agoWhy not use `tee`? I regularly do `cat ~/.ssh/id_rsa.pub | ssh foo@bar tee -a ~/.ssh/authorized_keys` without repercussion. For larger files I usually use rsync as you can resume interrupted transfers
- EnigmaCurry 4y agowhy not ssh-copy-id?
- thenickdude 4y agoYou can use cat for that: mysqldump -u mysql db | ssh user@rsync.net "cat > db_dump"
- a-dub 4y agoi'm not 100% on this, but i seem to have foggy memories of seeing it used in bizarre ways like this by scripts written to be ultraportable (output of early versions of autotools maybe?).
- deleted 4y ago[deleted]
- theonemind 4y agoCult of dd : https://news.ycombinator.com/item?id=13896675 https://news.ycombinator.com/item?id=13896675 / https://eklitzke.org/the-cult-of-dd https://eklitzke.org/the-cult-of-dd
- ggm 4y agouse of cat needs to be discussed too: cat < file | somecmd | cat > file2 #vs somecmd < file > file2
- donio 4y agoThat particular example is an overly silly but I often write cat file | grep ... or cat file | awk '{blah}' for ad-hoc stuff even though I am fully aware that I could easily avoid the cat simply because having the regex or awk code at the very end makes it easier to read, edit or extend with additional pipeline elements.
- billpg 4y agoI prefer the first way, as you can read it left to right and each step is meaningful. "Take the contents of file, process it through somecmd, then save the result to file2." I would hope both commands boil down to exactly the same action. Any overhead from invoking cat (isn't that a shell built-in?) should be negligible.
- ufo 4y agoThe difference with cat is that at least it is looks nicer in some cases. For example, if I want to put the input file at the start of the pipeline, the cat-less version looks weird and confusing. <file1 somecmd >file2 I wish that shells would just optimize the useless cat behind the scenes, so I could use it without soliciting complaints :)
- __del__ 4y agoit's a short walk from shells that optimize "known commands" to signed apps and an ecosystem you can't contribute to unless you've paid the appropriate gatekeeper
- Maursault 4y agodd found fame about 17 years ago when supposedly low-level copying gui applications would consistently not succeed in some cases for unknown reasons. dd worked where fancier copy applications were hit or miss. So while other methods will work, dd consistently also works. UNIX was designed redundantly intentionally. And I'm pretty sure the U stands for 'useless.'(yeah, jk)
- pseudostem 4y agoI first used it to dual boot windows (98 or NT4 I don't remember) and slackware. I can't remember specifics, but I dd ed the first 512B to a file and pointed the windows OS loader to that file. The specifics I don't remember are why I wanted to use the windows loader instead of LILO.
- Maursault 4y agodd is pretty old software. I just meant it entered the mainstream consciousness of self-techs (non-pros) and pirates simultaneously ~2004-2007.
- yrro 4y agoIndeed, most users are better off by launching GNOME Disks, going to the kebab menu, choosing Restore Disk Image & selecting the image and target disk interactively. Or running: $ gnome-disks --restore-disk-image=/etc/motd Behind the scenes it calls udisks, which uses polkit to ensure that the user is authorized to write to the target disk. Overall, the chances of a user typoing and wiping out the wrong disk are greatly reduced. They also get to see progress etc.
- eternityforest 4y agodd is one of THE most common ways people accidentally wipe disks. All it takes is typing sda when you meant sdb. The by-label trick helps a bit, but I still don't like it. Etcher exists. Etcher will give an alert when it's done flashing. Etcher will verify what it wrote. Etcher will predict how much time is remaining. The raspi imaging utility goes even farther and gives you configuration options. There are many other special purpose flasher utilities like that. Where dd really shines is in a script, for making empty images of a certain size. But even then... there are tools like truncate to make a sparse file instead. The only other time I ever need the CLI is for directly creating a compressed image, but that's a somewhat uncommon task for me. And I would not be surprised if one of the GUIs had it by now.
- superdisk 4y agoDownloading a giant electron application which bundles Chrome and Xbox joystick drivers just to write to a flash drive feels downright criminal when there's a 200kb program that will do the same job already on your computer.
- Kaze404 4y agoI personally use dd for such cases as well, but I don't mind downloading a "giant electron application which bundles Chrome and Xbox joystick divers" (I just downloaded Etcher and it takes 40MB on my machine) if that's what it takes to lower the risk of accidentally wiping my hard drive when I want to write an ISO to a USB drive.
- eternityforest 4y agoElectron makes the problem of flashing totally solved. I have a 1TB SSD for exactly this reason, so I don't have to worry about what's light and what isn't.
- eggsome 4y agoMy best useless use of dd was in the Ubuntu 16.04 days. I was at a satellite office with all windows PCs for the day, so used a live disk to get a decent environment to get things done. Only problem was that the DVD drive kept spinning down and every time I did something that was not cached it made me wait forever. nohup + while loop + sleep 4s + raw dd read from CD for the win :) EDIT: Reading this article it sounds like dd has no "special" ability to access the disk in a raw way. But surely that's what the nocache option is for...
- ace2358 4y agoI’ve only used dd once in my life (not much of a hacker!) but it was mostly useless. Dropbox had some promotion for their new photo storage service, where they were giving away free extra space up to 10gb to encourage you to store your photos. Some clever cookie told me you could make 10gb of ‘empty’ jpeg files and put them in your Dropbox folder. The Dropbox app would compress this 10gb of files for upload (usual behaviour for the app in those days) and increase your storage permanently for free. I can’t remember the command but it was something like dd if=/dev/zero of=/path/to/dropbox/1.jpg bs=10000000 Bam! 1 10gb jpeg file full of zeros that compress down to a few bytes. Instant 10gb free. Now uploading required! I think I still have that dropbox account, I learnt about dd and dev/{null,zero,one} that day and never used them again.
- orkj 4y agoFor those not familiar, there is a "useless use of cat" (https://porkmail.org/era/unix/award https://porkmail.org/era/unix/award) which I have a feeling this title is referring to
- geogra4 4y agolovingly referred to as "disk destroyer"
- AlchemistCamp 4y agoI'm disappointed this wasn't about VIM.
- MertsA 4y agoI think this has a bit of bad advice for using cp or shell redirection to read / write to raw block devices but dd isn't necessarily the best either. Personally any time I'm trying to image some hard drive, damaged or otherwise, I'll just about always jump straight to ddrescue (not dd_rescue). It's similar to dd, surprise surprise, but it keeps a log of which parts of the input and output have been copied / had errors / skipped. Nothing is more annoying than waiting an hour for some large copy to make progress and then run into an error or get interrupted for whatever reason. Using ddrescue because it keeps a log of the status of the operation you can resume it with the same command and it will pick back up right where it left of instead of having to start all over. It's also intelligent enough to not fail the first time and skip over some bad region of the disk on error and come back and reattempt it using various strategies once it's already copied the low hanging fruit. There's very little reason not to use it, even if it's just to get a nice progress view instead of just the current amount of data copied.
- gbrown_ 4y agorw is a nice dd like tool with a unix syntax https://sortix.org/rw/ https://sortix.org/rw/
- kazinator 4y ago> Usage of dd in this context is so pervasive that it’s being hailed as the magic gatekeeper of raw devices. That's the thing; it isn't. Author forgot to explain (if he knows that at all) that /dev/sda2 on Linux is not a raw device. It's a block device. So if dd is hailed as something to use on /dev/sda, that's not an example of being hailed for a raw device. dd's capability to control the read/write size is needed for classic raw devices on Unix, which require transfers to follow certain sizes. E.g if a classic Unix tape needs 512 byte blocks, but you do 1024 byte writes, you lose half the data; each write creates a block. The raw/block terminology comes from Unix. You have a raw device and a block device representing the same device. The block device allows arbitrarily sized reads and writes, doing the re-blocking underneath. That overhead costs something, which you can avoid by using the raw device (and doing so correctly).
- trasz 4y agoAt least this used to be the case. Nowadays FreeBSD doesn’t implement block devices at all - there are only raw disk devices.
- p_l 4y agoSomething that, IIRC, came from Linux and it's allowance on at least some block devices to support "character" style access. Tapes are still annoying on both in their block-ness, iirc?
- trasz 4y agoIn a way, but in Linux block devices can still be accessed as block devices, while in FreeBSD (since around 1999) they can't - there's no caching at that level anymore; raw devices (which Linux got a few years before that) are the only kind of devices. (If you do "ls -al /dev", you'll still see block devices, but it's maintained only to pretend to userspace, just like major/minor numbers.) Tapes aren't block devices at all - that's why you can't mount them :-)
- p_l 4y ago
- adrusi 4y agoI always assumed that dd was preferred when writing disk images to flash media to ensure each hardware block is written to exactly once, extending the live of the device and increasing performance substantially. I never checked if that's actually a problem with general-purpose file IO, nothing stops the kernel from noticing when a pipe is connected to a block device and configuring an appropriate buffer. status=progress is also quite handy. You can splice a "| dd status=progress |" in the middle of a pipeline to get a sense of what's happening.
- lgeorget 4y agoOne thing dd does for me, that cp and cat do not, is that it forces me to read and check the command at least three times before pressing enter, which is a very good thing when messing with raw devices. When I first learned that dd was not magical, I started using cp but I made some mistakes with partitions number and whatnot (nothing serious). Maybe it's just the wierd syntax or the fact that I treat dd differently, I'm just more cautious and don't press enter automatically. Of course it's silly but to me at least, that's a good reason to keep using dd.
- teddyh 4y ago> in the same way we ended up with a Window system called X There actually was a Window system named “W”; it was the window system for the “V” operating system (IIRC). “X”, therefore, was the successor to “W”.
- GekkePrutser 4y agoI understand his point, but using 'dd' allows you to set a buffer size which can make cloning a bunch faster. It also has great progress reporting (status=progress) which is really useful for the things dd is usually used for. And even if you use it without it being needed, it's not a big deal. It doesn't add much overhead, if any.
- GekkePrutser 4y agoBy the way, on versions that don't support status=progress (like busybox and the BSD versions), you can periodically send a USR1 signal to get a progress update.
- trasz 4y agoNot sure about other BSDs, but FreeBSD and MacOS already support status=progress. Also ^T is much more convenient than SIGUSR1.
- GekkePrutser 4y agoI thought FreeBSD didn't but indeed it was Alpine (which uses busybox instead of gnu-utils) indeed. What I normally do is just do a 'watch -n 30 killall -USR1 dd' in another window which triggers regular progress updates :) That why I don't use ^T I use FreeBSD also but indeed it supports it now.
- messe 4y agoOn BSDs (and IIRC macOS) you want a SIGINFO not SIGUSR1, which can be sent by ^T.
- licebmi__at__ 4y ago>cat /dev/cdrom > myfile.iso Heh, I remember finding out this on the good old days when I was finding out how to rip cds and shitting bricks. I mean I read it on a BBS and thought "that can't be right"; I was expecting to find something like the Nero suite back in windows. Much more recently, I enjoyed the same kind of amazement on the bash tcp pseudo devices.
- bcook 4y agoI was much more confused by your post than I should've been because I overlooked that ">" is how HackerNews prefixes quoted text.
- mdp2021 4y ago> dd if=/dev/sda | gzip > image.gz Last time I researched this, a long time ago: `lz4` is (was?) your best friend. Oh, but - to follow the submitted article's "rant" - if one wants to use `cat` instead of `dd`, shall there be freedom... Sometimes people like some kind of form-aided "readibility and clarity and method", which leads you to forms like `cat filename | grep` .
- gabrielblack 4y agoI don't think so, this command could be significantly slow: cp myfile.iso /dev/sdb compared with this one: dd if=myfile.iso of=/dev/sdb bs=32M because implementation of cp have a fixed buffer, so if the amount of data is big and the disks fast, using cp you are calling more read() and write() syscalls than necessary, slowing down the copy process.
- chronogram 4y agoOn the left side of a typical Linux hobbyist experience graph you probably have "this is a disk image file, and this is a disk, I'll just copy the disk image file onto the disk!", then you have a period of "I'll use the cool dd tool (without oflag=direct and/or sync because by this point you know people use it but you don't know why people use it) like I saw on the internet!", then when you understand that everything is a file you have "this is a disk image file, and this is a disk, I'll just copy the disk image file onto the disk!" again. I personally suggest recommending people the Disks application included with Ubuntu Desktop and Fedora/Centos Workstation. It shows icons representing internal disks, SD cards or flash drives so they know what device they want to work with. If they want to take their time they can see all the information about the drives and partitions, they can start discovering and asking questions and reading up on how computers use disks right from there if they want to. And if they don't want to, it's just extra confirmations that it's the correct disk or DVD that they want to put their image onto. Then when they're sure about the device they can create a disk image and restore a disk image in that same application!
- klibertp 4y agogparted is also nice, although Disks seem to display more information by default.
- flohoff 4y agoThe issue is that the author makes a lot of use of pipes e.g. "|" which has a single page as its buffer and immediatly makes the task pretty much CPU bound. (Because of context switches etc). This is why people use "dd" - It does not need pages/pipes/context switches to stream data in large chunks from a device to a file or vice versa.
- mkup 4y agoIn Linux, raw disk devices like /dev/sda1 are cached by kernel (unless opened with O_DIRECT flag). In FreeBSD (and presumably other UNIX implementations) they aren't: https://docs.freebsd.org/en/books/arch-handbook/driverbasics/#driverbasics-block https://docs.freebsd.org/en/books/arch-handbook/driverbasics... So, in FreeBSD "dd if=/dev/ada0p1 of=/dev/null bs=1 count=1" will fail: disk driver will return EINVAL from read(2) because I/O size (1) is not divisible by physical sector size (usually 512). "cat" with buffer size X (which depends on the implementation) will either work or not depending on divisibility of X by physical sector size, and other random factors, like short file I/O caused by delivery of a signal. Summary: dd(1) still has its place and author of original article is getting it wrong.
- rawoke083600 4y agoStill my fav way to create a linux-boot-usb and a poorman's disk-benchmark tool :)
- nikau 4y agoGood time to mention "sdd" which can handle faulty disks. It can be set to skip over bad blocks to copy as many healthy blocks as possible, and then go back and retry the bad blocks over and over until they either read successfully or you stop it.
- bayindirh 4y agoIt has a more advanced brother: gnu_ddrescue, ddrescue short. It can nudge and prod disks with severe bad sectors to give up the secrets they hold. I used it once to recover a 1TB NTFS drive which can't be mounted by Windows. It only failed to recover 12kB in 1TB, which corrupted an MP3 file, which can be easily replaced.
- nikau 4y agoAhh nice, it's been a while since I've had to recover data from a disk, will keep ddrescue in mind.
- felixhammerl 4y agoBuT YoU ArE UsInG cat AnD dd wRoNg! So?
- sitkack 4y agoThe article might educate, but that doesn't appear to be its primary purpose. The problem with this article and the useless-cat article is that 1) they create a false right vs wrong dichotomy and 2) it gets amplified and used as a tool by the small minded as a citation for "the correct way". A post on learning your tools is fine. But calling something useless and needlessly adding branches to one's decision tree doesn't empower anyone except jerks. Stay useless. Were any of the cited examples harming anything? No they weren't.
- G3rn0ti 4y ago> The fact of the matter is, dd is not a disk writing tool. Neither “d” is for “disk”, “drive” or “device”. Where does dd's name come from then? Its man page does not tell.
- OJFord 4y agoSee footnote: > dd actually has two jobs: Convert and Copy. A post on comp.unix.misc (incorrectly) claimed that the intended name “cc” was taken by the C compiler, so the letters were shifted in the same way we ended up with a Window system called X. A more likely explanation is given in that thread as pointed out by Paweł and Bruce in the comments: the name, syntax and purpose is almost identical to the JCL “Dataset Definition” command found in 1960s IBM mainframes.
- nyuszika7h 4y agoIronically, the post contains a useless use of `cat`: cat /dev/sda | pv | cat > /dev/sdb This could be replaced by just: pv < /dev/sda > /dev/sdb Or in this case even: pv /dev/sda > /dev/sdb
- yakubin 4y agoIf you want to write the input file at the beginning and write the command later, then you may use: </dev/sda pv >/dev/sdb
- bee_rider 4y agoJust a note -- despite the title, the article eventually presents a nuanced view and points out that "dd" has some uses. > If an alias specifies -a, cp might try to create a new block device rather than a copy of the file data. If using gzip without redirection, it may try to be helpful and skip the file for not being regular. Neither of them will write out a reassuring status during or after a copy. > dd, meanwhile, has one job*: copy data from one place to another. It doesn’t care about files, safeguards or user convenience. It will not try to second guess your intent, based on trailing slashes or types of files. > However, when this is no longer a convenience, like when combining it with other tools that already read and write files, one should not feel guilty for leaving dd out entirely.
- t312227 4y agohmmm ... i have to admit, i really don't get this article, imho. its not very well written: why using these "oldschool" if / of parameter how its meant to be used dd < /dev/zero > /dev/sdx like others already mentioned: parameter "bs" ~ block-size dd < /dev/zero bs=1M > /dev/sdx or parameter "count" ~ number of blocks of specified size dd < /dev/zero bs=1M count=100 > /dev/sdx show progress ~ utility "pv" dd < /dev/zero bs=1M | pv > /dev/sdx etc.etc. why do people write such articles w/o at least consulting the manpages!? and why is such a mediocre article on the frontpage of HN!? just my 0.02€
- StephenSmith 4y agoOkay, so what's the best way to image an SD card then? I use the following on a 64gb SD card -- is there a better way? dd if=/dev/sda | gzip > img.gz dd if=img.gz | gunzip | dd of=/dev/sda
- infinityio 4y agothis is not the point of the article, but you may want to specify blocksize in those commands? ymmv tho
- sfink 4y agoDo these not work?: gzip < /dev/sda > img.gz gunzip < img.gz > /dev/sda Those should give you better block sizes than default dd without bs=128k or whatever.
- usui 4y agoI would install pv (optional) and do cat /dev/mmcblk0 | pv -abtrN Processed | gzip > img.gz cat img.gz | gunzip | pv -abtrN Processed > /dev/mmcblk0
- bzbarsky 4y agoFor your first command: gzip -c /dev/sda > img.gz For your second command: gunzip -c img.gz > /dev/sda are what I suspect the blog post would recommend.
- yakubin 4y agoIs there a tool other than dd that would allow me to skip n bytes from a file and then read m bytes? I mean one tool, not a combination of head and tail, which is too elaborate for my taste.
- abotsis 4y agodd indeed has advantages: skip, offset, being able to ignore errors to name a few.
- dmuth 4y agoThat's interesting about the blocksize. I was never impacted by that because I mostly use dd for writing large files, and a large block size always made more sense to me. For example, if I wanted to test writing a 1 GB file to a flash drive, I'd do something like: dd if=/dev/zero of=/mnt/test.data bs=1M count=1024 Overall, a very informative article though!
- FunnyBadger 4y agoThere was a time when "dd" was the only game in town for deep copies on Unix systems. E.g. HPUX in the 1980s. It always had issues of course: e.g. if the source was a 100MB drive and you dd'ed to a blank 200MB, you'd get a new 100MB drive and all the rest was lost because it was a low level copy. Similarly any bad sectors on the old drive became bad sectors on the new and any bad sectors on the new were simply used as if they were OK.
- alanh 4y agoyou know you’re a web dev if `dd` makes you think of "definition definition", a child of `dl` and sibling to `dt` :)
- krylon 4y agoDid anyone else expect something along the lines of dd if=/dev/urandom of=/dev/null ?
- anfractuosity 4y agoI didn't think 'cat' itself could force a sync to disk? When writing to a USB sd card reader, sometimes the data wouldn't be written entirely without calling 'sync' but dd can perform the sync itself with 'fsync'
- viddi 4y agoTwo years ago, I had to clone a disk image from one UEFI notebook over to another identical one. Using gzip with output redirection to the device file (might also have been zcat or cat with gzip) did just not work. After days of unsuccessful attempts, I resorted back to dd with gzip. And believe it or not: Only with dd the disk clone was successful. I came as far as comparing the image files, which were identical, but I did not have the time to compare the disks. So, this still is a mystery for me. On a side note: If you ever need to do a full disk clone nowadays, I can only recommend clonezilla. Using FAI (Fully Automatic Installation) would be even better instead of doing image clones, but sometimes, that's not an option.
- a1369209993 4y ago> dd if=/dev/sda | gzip > image.gz This actually serves at least two approximately legitimate purposes: firstly, it ensures that reads from sda are aligned to a (at least nominal, ie 512-byte) disk block, which doesnt matter for normal, kernel-supported drives like IDE/SATA/most USB (which is almost certainly what sda is), but avoids bespoke devices (or their drivers) trying to do something clever when gzip asks for 1 or 7 or 17 bytes. (And writes to poorly-designed devices/drivers can be even worse.) More importantly, like useless use of cat, it prevents gzip from trying to delete sda when it's done, which is something it will in fact do: $ echo test > /tmp/sda $ gzip /tmp/sda $ cat /tmp/sda cat: /tmp/sda: No such file or directory (For gzip specificially, you can also prevent this by writing `gzip </tmp/sda`, but I've occasionally run into tools that try to 'intellegently' handle stdin file descriptors that point at 'real' files, so I feel better having a separate process blocking the way.)
- smcameron 4y agogzip doesn't do a stat(2) and detect that /dev/sda is a device node? looks at gzip source code -- sure looks like it will delete device nodes. And seems like the wrong thing for it to be doing.
- a1369209993 4y agoThe underlying question, when dealing with stupid DWIM logic, isn't "Does it do that?"; it's "Can I be sure that it does that, and will continue to do that, even on old versions on legacy systems that a unknown amount of random crap depends on to not fall over because I accidentally the hard drive?". And there's also the fact that deleting the source file by default is always the wrong thing for it to be doing. If I want to deal with corner cases like being almost out of disk space, I can pass --delete-source explicitly.