4 ms·
I carefully considered this before I optimized GNU yes. The reason it is useful is because yes can output anything, and so is useful to produce any repeated da
by pixelbeat__ 4y ago
I carefully considered this before I optimized GNU yes.
The reason it is useful is because yes can output anything, and so is useful to produce any repeated data for test files etc.
You can see the justification detailed in the original optimization commit:
https://github.com/coreutils/coreutils/commit/35217221 https://github.com/coreutils/coreutils/commit/35217221
- dncornholio 4y agoSo you actually wrote a new program. Yes was made for those pesky installers, not for producing large amounts of data. This would be my approach. Keep yes simple and create a new program that does spilling data well.
- jraph 4y agoAnd then you have a possibly confusing situation where you have two programs that do essentially the same thing, but one is faster, and the other one is possibly not provided by default. As a user, in cases it matters, you'd have to know the issue and bother installing the new program. This is worse. As developers, I think it's our duty to make life of users simpler, even if it makes our lives a bit more complicated. I'd argue that's what we are here for. I guess there's no ideal solution. But I think the "new" program does what the first one did better, and does not do anything worse. We are talking about a program that still under 1000 lines of code and that's not getting new features every month, or at all anyway, so maintainability does not seem to be a big issue? I see the reasoning but I don't see any actual practical drawback to have improved the original program directly in this specific case. I don't see any advantages of keeping "yes" dead simple neither. The new version is still pretty much readable and the extra time it takes to read it and modifying without doing mistakes seems worth the advantages.
- dncornholio 4y ago> And then you have a possibly confusing situation where you have two programs that do essentially the same thing My point is, I'd never thought of using yes for this purpose. So in this case, you could make a command called 'outputsomethingfast' and you could make a command called yes, that internally calls 'outputsomethingfast --output=yes' or something like that. To me this is way more logical, and more in line with the Linux philosophy, right?
- jraph 4y agoI guess. I don't like the name "yes" and I think we would have been better off with a more general name since the command is general, but now it's there, so… However, this is independent from this optimization, "yes" already had this feature of outputting anything, I think? But I expect this kind of accident to happen in any working system that has been long enough. This seems unavoidable. So we'd better put up with this kind of mess probably.
- jcelerier 4y agoi really don't want to have to learn 12342384 programs. it's much less discoverable than having a few programs with a --help (and more generally a tree-based organization of functionality on your computer). also, if there's a new program, say, "fastrepeat" wouldn't that be a duplication of functionality between "yes" which just outputs "y" and "fastrepeat 'y'" which is, like, even more bloat since now you need both ?
- jacquesm 4y agoI would much rather have a very easy to remember command that does one thing and one thing only (namely, what it says on the tin) than to have to remember or dig through a whole slew of command line options in order to get 'yes' to become the equivalent to 'no' or 'cat'.
- mavhc 4y agoreally you want one program that outputs data, and yes is an alias to outputdatafast --data="y"
- jcelerier 4y ago> I would much rather have a very easy to remember command that does one thing and one thing only well, I definitely don't. I don't want to encumber my mind with a name for every single of the 25000 "one thing" things I have to do.
- remix2000 4y agoTouche, and there is already a program for the exact purpose of generating large amounts of data: jot(1) https://manpage.me/?q=jot https://manpage.me/?q=jot
- dncornholio 4y agoAnd here is the problem. Now we have 2 programs that do essentially the same thing.
- xphx 4y agoOr if homogeneous data is fine, just cat or dd from a source like /dev/zero.
- drewzero1 4y agoI think I've seen /dev/random used for generating large amounts of garbage data as well. (Though usually my problem is too much garbage data, not too little.)
- endgame 4y agoI have a history question. I've seen this link a few times: https://www.gnu.org/prep/standards/html_node/Reading-Non_002dFree-Code.html#Reading-Non_002dFree-Code https://www.gnu.org/prep/standards/html_node/Reading-Non_002... Did this advice from GNU inspire you to optimise `yes`, did your optimised `yes` inspire GNU to write this, or is there no historical connection between your optimised `yes` and this advice?
- pixelbeat__ 4y agoTrying to differentiate GNU implementations had nothing to do with it. This is never a consideration for me. It was worth the slight increase in complexity for the reasons stated in the commit message. Also an unmentioned point is that the coreutils code is very often referenced, so should be as robust and performant as possible, so those properties may percolate elsewhere.
- endgame 4y agoThanks, that makes a lot of sense.