11 ms·
I find sponge really useful. Have you ever wanted to process a file through a pipeline and write the output back to the file? If you do it the naive way, it won
by HellsMaddy 4y ago
I find sponge really useful. Have you ever wanted to process a file through a pipeline and write the output back to the file? If you do it the naive way, it won’t work, and you’ll end up losing the contents of your file:
awk '{do_stuff()}' myfile.txt | sort -u | column --table > myfile.txt
The shell will truncate myfile.txt before awk (or whatever command) has a chance to read it. So you can use sponge instead, which waits until it reads EOF to truncate/write the file.
awk '{do_stuff()}' myfile.txt | sort -u | column --table | sponge myfile.txt
- jamespwilliams 4y agoAre there any legitimate reasons to have a particular file both as an input to a pipe and as an output? I wonder whether a shell could automatically “sponge” the pipe’s output if it detected that happening.
- teawrecks 4y agoYeah, it seems like the kind of command that you only need because of a quirk in how the underlying system happens to work. Not something that should pollute the logic of the command, imo. I would expect a copy-on-write filesystem to be able to do this automatically for free.
- paulmd 4y ago> I would expect a copy-on-write filesystem to be able to do this automatically for free. this is an artifact of how handles work (in relation to concurrency), not the filesystem. copy-on-write still guarantees a consistent view of the data, so if you write on one handle you're going to clobber the data on the other, because that's what's in the file. what you really want is an operator which says "I want this handle to point to the original snapshot of this data even if it's changed in the meantime", which a CoW filesystem could do, but you'd need some additional semantics here (different access-mode flag?) which isn't trivially granted just by using a CoW filesystem underneath.
- caymanjim 4y agoDo people really use copy-on-write filesystems though? I mean it'd be great if that were a default, but I rarely encounter them, and when I do, it's only because someone intentionally set it up that way. In 30+ years of using Unix systems, I can't even definitively recall one of them having a copy-on-write filesystem in place. Which is insane considering I used VAX/VMS systems before that and it was standard there.
- mjochim 4y agoHuh? Btrfs is copy on write and it's definitely being used.
- r3trohack3r 4y agoFreeBSD uses ZFS by default, which is copy on write post-snapshot.
- Arnavion 4y agoThe shell can't detect that unless the file is also used as input via `<`. So it couldn't do that in HellsMaddy's example, since the filename is given as an arg to awk instead of being connected to its stdin.
- dejj 4y agoAnd detecting that would actually implement move semantics in the shell. Isn’t there already a project to bring Rust to the shell?
- mjochim 4y agosed --in-place has one file as both input and output. It's not really different from any pipe of commands where the input and output files are the same. But sed also makes a copy of the file before overwriting it - per default.
- indigodaddy 4y agoNever heard of sponge. Isn’t it just the same as tee ?
- halostatue 4y agoIt’s part of the moreutils collection and is better considered a buffered redirect than anything similar to `tee`.
- indigodaddy 4y agoAh yes I glanced too quickly over the surface here. It does look more like redirection. Will have to look at it more. Appreciate your helpful response vs the downvoters and the one unhelpful/snarky response.
- AshamedCaptain 4y agoThe point is not to truncate the file immediately when it is open for output, and before the input side has had time to slurp it.
- binwiederhier 4y agoI urge you to actually read the link before commenting. sponge is the front and center example of the link you are commenting on.
- indigodaddy 4y agoThe downvoting here is the equivalent of getting shamed for “asking a stupid question at work.” Yes I should have done a bit more homework but shaming for asking a clarifying question is unreasonable. Those of you who have the downvote trigger-finger can and should do better.
- HellsMaddy 4y agoI agree with you, the downvotes are unnecessary, it was actually a good question. tee actually does sorta work for this sometimes, but it’s not guaranteed to wait until EOF. For example I tested with a 10 line file where I ran `sort -u file.txt | tee file.txt` and it worked fine. But I then tried a large json file `jq . large.json | tee large.json` and the file was truncated before jq finished reading it.
- alerighi 4y agoIt would also be useful in cases when the file must be written with root privileges (but the shell is run as a normal user) Now you have to use `tee` for that, that is fine but if you don't want to echo the file back to the terminal you have to do command | sudo tee file > /dev/null with this you can simply do command | sudo sponge file
- leni536 4y agoCan't you just use `cat`?
- ben0x539 4y agoYou cannot use `sudo cat` to open a file with root privileges, because `sudo cat > foo` means "open file foo with your current privileges, then run `sudo cat` passing the file to it", and the whole root thing only happens after you already tried and failed to open the file. I seem to get this wrong roughly once a week.
- gorgoiler 4y ago… | sudo tee foo
- readme 4y agosudo sh -c "cat > file"
- deleted 4y ago[deleted]
- icedchai 4y agoI knew a person who would give something like this as a sysadmin / devops interview question. It was framed as "'sudo cat >/etc/foo' is giving a permission denied error! what's wrong?" Usually the interview candidate would go off on a tangent...
- justinsaccount 4y agoprotip: don't run this awk '{do_stuff()}' myfile.txt | sort -u | column --table | sponge > myfile.txt
- RulerOf 4y ago>I find sponge really useful. When I found `sponge`, I couldn't help but wonder where it had been all of my life. It's nice to be able to modify an in-place file with something other than `sed`.
- caymanjim 4y agoI can see how this is handy, but it's also dangerous and likely to bite you in the ass more than not. I think sponge is great, but I think your example is dangerous. If you make a mistake along the way, unless it's a fatal error, you're going to lose your source data. Typo a grep and you're screwed.
- HellsMaddy 4y agoRight. You should always test the command first. If the data is critical, use a temporary file instead. I usually use this in scripts so I don’t have to deal with cleanup.
- sedatk 4y ago> If the data is critical, use a temporary file instead Use a temporary file always. Sponge process may be interrupted, and you end up with a half-complete /etc/passwd in return.
- rcoveson 4y agoCouldn't `mv` or `cp` from the temp file to `/etc/passwd` be interrupted as well? I think the only way to do it atomically is a temporary file on the same filesystem as `/etc`, followed by a rename. On most systems `/tmp` will be a different filesystem from `/etc`.
- AdamJacobMuller 4y agomv can't, or, more correctly the rename system call can not. rename is an atomic operation from any modern filesystem's perspective, you're not writing new data, you're simply changing the name of the existing file, it either succeeds or fails. Keep in mind that if you're doing this, mv (the command line tool) as opposed to the `rename` system call, falls back to copying if the source and destination files are on different filesystems since you can not really mv a file across filesystems! In order to have truly atomic writes you need to: open a new file on the same filesystem as your destination file write contents call fsync call rename call sync (if you care about the file rename itself never being reverted). This is some very naive golang code (from when I barely knew golang) for doing this which has been running in production since I wrote it without a single issue: https://github.com/AdamJacobMuller/atomicxt/blob/master/file.go https://github.com/AdamJacobMuller/atomicxt/blob/master/file...
- thanatos519 4y ago'sort -o file' does the same thing but I like the generic 'sponge'. Not sure why it's a binary, since it's basically this shell fragment (if you don't bother checking for errors etc): (cat > $OUT.tmp; mv -f $OUT.tmp $OUT) Hmmm ... "When possible, sponge creates or updates the output file atomically by renaming a temp file into place. (This cannot be done if TMPDIR is not in the same filesystem.)" My shell fragment already beats sponge on this feature!
- kevincox 4y agoIt would be nice to update to use anonymous files where supported (Linux does). This allows you to open an unnamed file in any directory so that you can do exactly this, write to it then "rename" it over another file atomically.
- mzs 4y agoIf you're using awk anyway: awk 'BEGIN{system("rm myfile.txt")}{do_stuff()}' <myfile.txt | sort -u | column --table > myfile.txt In general though: <foo { rm -f foo && wc >foo }
- dan-robertson 4y agoI usually used tac | tac for a stupid way of pausing in a pipeline. Though it doesn’t work in this case. A typical use is if you want to watch the output of something slow-to-run changing but watch doesn’t work for some reason, eg: while : ; do ( tput reset run-slow-command ) | tac | tac done
- caymanjim 4y agoThe easy way to do this is with 'tee /dev/tty' in the middle.
- deleted 4y ago[deleted]
- dan-robertson 4y agoI don’t understand why that would work? The goal is to buffer the output of the commands and then send it all at once.
- caymanjim 4y agoI misunderstood the goal.
- yjftsjthsd-h 4y agoHow is that different from | cat?
- efreak 4y ago`cat` will begin output immediately. `tac` buffers the entire input in order to reverse the line order before printing it. Piping through tac ensures EOF is reached, while piping through tac a second time puts the lines back in order
- pixelbeat__ 4y agoUsing sponge to replace a file is equivalent to method 6 at: https://www.pixelbeat.org/docs/unix_file_replacement.html https://www.pixelbeat.org/docs/unix_file_replacement.html Note the caveats listed there
- 1vuio0pswjnm7 4y ago"Have you ever wanted to process a file through a pipeline and write the output back to the file?" Personally, no. I prefer to output text processsing results to a different file, maybe a temporary one, then replace the first file with the second file after looking at the second file. awk '{do_stuff()}' 1.txt | sort -u | column --table > 2.txt less 2.txt mv 2.txt 1.txt With this sponge program I cannot check the second file before replacing the first one. The only reason I can think of not to use a second file is if I did not have enough space for the second file. In that case, I would use a fifo as the second file. (I would still look at the fifo first before overwriting the first file.) mkfifo 1.fifo awk '{do_stuff()}' 1.txt | sort -u | column --table > 1.fifo & less -f 1.fifo awk '{do_stuff()}' 1.txt | sort -u | column --table > 1.fifo & cat 1.fifo > 1.txt
- ikiris 4y agoYeah they should have called that tool yolocat
- pletnes 4y agoIf 2.txt is tracked by git, there’s no need to go through these hoops. Then, sponge starts to make sense. Except from that case, I agree totally.
- 1vuio0pswjnm7 4y agoDid git exist when this sponge program was written?
- varenc 4y agoYou’re implying that all uses of `sponge` are interactive with a user actively involved. I used sponge a lot in scripts where I’ve already checked and validated its behavior. When I’m confident it works I can keep just use sponge for simplicity.
- deleted 4y ago[deleted]
- matheusmoreira 4y ago> Have you ever wanted to process a file through a pipeline and write the output back to the file? I don't think overwriting the input data is a good idea due to risk of data loss.
- chakkepolja 4y agoThis was such a footgun! This may be fairly intuitive if you know how shell redirection is implemented. But hard to think of that during the time you write a command.
- deleted 4y ago[deleted]