8 ms·
A surprise with how '#!' handles its program argument in practice
- tombert 11mo agoHuh. I wish I had known this before. NixOS is annoying because everything is weird and symlinked and so I find myself fairly frequently making the mistake of writing `#!/bin/bash`, only to be told it can't find it, and I have to replace the path with `/run/current-system/sw/bin/bash`. Or at least I thought I did; apparently I can just have done `#!bash`. I just tested this, and it worked fine. You learn something new every day I guess.
- miffe 11mo agoSeems like it only works in zsh, not bash or fish
- tombert 11mo ago[tombert@puter:~/testscript]$ ./myscript.sh bash: ./myscript.sh: bash: bad interpreter: No such file or directory You are right. Appears to only work with zsh. I will resume being annoyed then.
- adastra22 11mo agoIs this UNIX?
- tombert 11mo agoThis is NixOS, so no, it's Linux. I guess I just hoped it would work on Linux as well.
- adastra22 11mo agoLinux is UNIX in the context of my question. On Linux the shebang is actually handled by the kernel. When you load a binary and attempt to execute it with a syscall, the kernel reads the first few bytes of the binary. If it is an ELF header, it executes the machine code as you would expect. If the first two bytes are "#!", then it interprets it as a shebang header and loads the specified binary to interpret it. Again, this is kernel code. I admit I'm a bit confused as to why it didn't work on your system. This shouldn't be handled at the shell level.
- adastra22 11mo agoThe kernel interprets the shebang line, not the shell.
- jolmg 11mo agoIt is possible for the shell to handle it. From zshall(1): > If the program is a file beginning with ‘#!', the remainder of the first line specifies an interpreter for the program. The shell will execute the specified interpreter on operating systems that do not handle this executable format in the kernel. Taking a quick look at the source in Src/exec.c: execve(pth, argv, newenvp); // [...] if ((eno = errno) == ENOEXEC || eno == ENOENT) { // [...] if (ct >= 2 && execvebuf[0] == '#' && execvebuf[1] == '!') { // [...] (pprog = pathprog(ptr2, NULL))) { I guess at some point someone added that `|| eno == ENOENT` and the docs weren't updated.
- opello 11mo agoI did a little digging and found that the `|| eno == ENOENT` was added quite a bit earlier[1] than the actual pathprog lookup[2]. While I could find the "issue discussion" for the pathprog change[3] I wasn't able to find it for the ENOENT addition, which was kind of interesting and frustrating--[4] is the `X-Seq` mentioned in the commit but that seems to be inconsistent or incorrect for the actual cross-reference, and nearby in time wasn't helpful either. [1] https://sourceforge.net/p/zsh/code/ci/29ed6c7e3ab32da20f528aeba4bb76635f1a3b06/ https://sourceforge.net/p/zsh/code/ci/29ed6c7e3ab32da20f528a... [2] https://sourceforge.net/p/zsh/code/ci/29ed6c7e3ab32da20f528aeba4bb76635f1a3b06/ https://sourceforge.net/p/zsh/code/ci/29ed6c7e3ab32da20f528a... [3] https://www.zsh.org/mla/workers/2010/msg00522.html https://www.zsh.org/mla/workers/2010/msg00522.html [4] https://www.zsh.org/mla/workers/2000/msg01168.html https://www.zsh.org/mla/workers/2000/msg01168.html
- tombert 11mo agoI'm not sure the reason then, but they're definitely right; it works fine with zsh, doesn't work with bash. I wrote a test script to try it myself. I don't have fish installed and can't be bothered to go that far, but I suspect they're right about that as well.
- Crestwave 11mo ago`#!/usr/bin/env bash` is the most portable form for executing it from $PATH
- hamandcheese 11mo agoIs this meaningfully more portable than #!bash though?
- tombert 11mo agoIn a sibling thread someone pointed out that #!bash doesn't actually work if you're calling it from bash, and appears to only work with zsh. I just tried it and they were absolutely right, so `#!/usr/bin/env bash` is definitely more portable in that it consistently works.
- saurik 11mo agoThis mechanism doesn't do a PATH lookup: #!bash would only work if bash was located in your current working directory.
- Ferret7446 11mo agoI raise you a https://www.felesatra.moe/blog/2021/07/03/portable-bash-shebang https://www.felesatra.moe/blog/2021/07/03/portable-bash-sheb... Pedantic, but "#!" and "portable" don't belong in the same sentence
- Crestwave 10mo agoThe script you posted is only portable in theory and not in practice. Executing it while using a non-POSIX shell like Elvish (even without having it as a default shell) makes it immediately fail. Meanwhile, `#!/usr/bin/env` is completely portable in the practical sense. Even systems with non-standard paths like NixOS, Termux and GoboLinux patch in support for it specifically.
- nflekkhnnn 11mo agoAnything other than ”#!/usr/bin/env bash” is doomed to fail at some time.
- normie3000 11mo agoAnd is this shebang guaranteed to work always? Why isn't it more common?
- eyelidlessness 11mo agoIt’s quite common, although I probably see it used more frequently to invoke other (non-shell) scripting languages.
- int_19h 11mo agoIt's guaranteed to work provided that Bash is in the path. It's very common for Python. Less so for Bash for two reasons: because the person who writes the script references /bin/sh instead (which is required to be there) even when they are writing bash-isms, or because the person who writes the script assumes that Bash is universally available as /bin/bash.
- jolmg 11mo agoBecause /bin is the standard location for bash. The only one that breaks that expectation is NixOS (and maybe GuixSD?), apparently. I'm surprised they didn't symlink /bin or put a stub. Last time I tried NixOS was like 10 years ago. I thought there was a /bin/bash, but maybe it was just a /bin/sh? Other interpreters like python, ruby, etc. have more likelyhood of being used with "virtual environments", so it's more common to use /usr/bin/env with them.
- tombert 11mo agoThey do symlink /bin/sh to be fair, and that's very often good enough for a lot of scripts. That's what I usually do if I don't need anything bash offers.
- 11mo ago
- saintfire 11mo agoYou can use `#!/usr/bin/env bash` on NixOS
- tombert 11mo agoI didn't know that actually. I'll start using that from this point forward.
- kreetx 11mo ago/usr/bin/env and /bin/sh are part of the POSIX standard, this is why NixOS has those available.
- jlokier 11mo ago> /usr/bin/env and /bin/sh are part of the POSIX standard, this is why NixOS has those available. Contrary to popular belief, those aren't in the POSIX standard. The following are not in the POSIX standard, they are just widely implemented: - "#!" line. - /bin/sh as the location of a POSIX compliant shell, or any shell. - /usr/bin/env as the location of an env program. - The -S option to env.
- Guvante 11mo agoI think you are using "not required by the POSIX standard" when you say "not in" which is not an accurate shorthand. #! is certainly in the POSIX standard as the exact topic of "is /bin/sh always a POSIX" shell is a discussion point (it is not guaranteed since there were systems that existed at the time that had a non-POSIX shell there)
- johnisgood 11mo agoAre they in POSIX? I do not think they are. All of them is a convention from what I remember. Shebang is a kernel feature, for example, and POSIX does define the sh shell language and utilities, but does not specify how executables are invoked by the kernel. Similarly, POSIX only requires that sh exists somewhere in the PATH, and the /bin/sh convention comes from the traditional Unix and FHS (Filesystem Hierarchy Standard), but POSIX does not mandate filesystem layout. ... and so on. Correct me if I am wrong, perhaps with citations?
- teo_zero 11mo ago> apparently I can just have done `#!bash` I think you're mixing two concepts: relative paths (which are allowed after #! but not very useful at all) and file lookup through $PATH (which is not done by the kernel, maybe it's some shell trickery).
- aaronmdjones 11mo ago> and file lookup through $PATH (which is not done by the kernel, maybe it's some shell trickery) It's libc. Specifically, system(3) and APIs like execvp(3) will search $PATH for the file specified before passing that to execve(2) (which is the only kernel syscall; all the other exec*() stuff is in libc).
- jolmg 11mo ago> Although this is probably the easiest way to implement '#!' inside the kernel, I'm a little bit surprised that it survived in Linux (in a completely independent implementation) and in OpenBSD (where the security people might have had a double-take at some point). But given Hyrum's Law there are probably people out there who are depending on this behavior so we're now stuck with it. I don't see what there would be to gain in disallowing the program path on the shebang line to be relative. The person that wrote the shebang can also write relative paths in some other part of the file.
- saurik 11mo agoOr, like, if you aren't reading and caring about what the interpreter is--as that's the only time this can burn you: it isn't doing a PATH lookup, so you can't walk into this one on accident--then it could literally be something like /bin/rm on some key file. This entire article is based on an assumption that this is somehow so obviously bad that there doesn't even need to be an explanation or defense of any kind of that idea.
- jolmg 11mo ago> This entire article is based on an assumption that this is somehow so obviously bad that there doesn't even need to be an explanation or defense of any kind of that idea. I'm not reading it like that. The tone is just one of surprise, since this isn't something that one typically sees. Since it's obscure, it leads one to wonder if it can be bad, and I don't see how it could be. I think it survived in the independent Linux because it's the simple and obvious way to do things, and it doesn't lead to any exceptional power of misuse one didn't already have with writing the rest of the file.
- saurik 11mo agoRight: I agree with you. I'm saying the article is making an unfounded assumption and am providing more reasoning for why you are correct.
- jibal 11mo agoThere's no security issue here. Certainly the OP hasn't explained why there is one.
- jmclnx 11mo agoThere could be. Say you have '#!sh', someone could have put a broken 'sh' somewhere, that could be the first one the '#!' sees.
- jibal 10mo agoNo, this is about relative paths, not searching $PATH ..., which it doesn't do. (That happens when people do '#!/usr/bin/env sh' ... the security issues with that are a separate discussion.)
- somat 11mo agoThe thing that surprised me was that you can't write an interpreter in an interpreted language, at least not in obsd. It is possible if you jump through a few hoops but you can't directly call it. An example: if you made a language in python /bin/my_lang: #does nothing but pretend it does #!/usr/local/bin/python3 import sys print('my_lang args', sys.argv) for line in sys.stdin: print('invalid_line:', line) my_script: #!/bin/my_lang line of stuff another line of stuff chmod u+x my_script ./my_script Probably for the best, but I was a bit sad that my recursive interpreter scheme was not going to work. Update: looks like linux does allow nested interpreters, good for them. https://www.in-ulm.de/~mascheck/various/shebang/#interpreter-script https://www.in-ulm.de/~mascheck/various/shebang/#interpreter... really that whole document is a delightful read.
- jolmg 11mo agoWorked for me, but the way you described it has issues: 1. You chmod my_script twice. 2. Did you chmod u+x /bin/my_lang too? Since you put it in /bin, are you sure the owner isn't root?, in which case your user wouldn't have execute permission. Try +x instead of u+x. 3. Do you have python in that path? Try `/usr/bin/env python` instead. 4. In case you expected otherwise, my_script wouldn't be passed through stdin. It's just provided as an argument to my_lang.
- somat 11mo agoI am on openbsd. which does not allow it, it looks like nested interpreters are s supported on linux. So my loss there.
- teo_zero 11mo agoWait... your interpreter reads from stdin. Shouldn't it read its first arg, instead?
- somat 11mo agoI think that was what I was trying to figure out, how the program was passed. but OpenBSD does not do nested interpreters, it looks like if I had tried Linux it would have worked.
- rmunn 11mo agoHit the link expecting to read about UTF-8 Byte Order Marks at the top of the file, so that the first few bytes aren't actually #! but 0xEF 0xBB 0xBF #! instead. Ran into this one just a few months ago when a coworker who uses Windows had checked a Bash script into the Git repo. His editor was configured to save files as "UTF-8 with BOM" and so we were getting errors that looked like "./doit.sh: line 1: #!/bin/bash: No such file or directory". Can you see the invisible BOMb in that line? It's there, I promise you. That's not what the article was actually about, as it turned out. The surprise in the article was about relative paths for script shebang lines. Which was useful to learn about, of course, but I was actually surprised by the surprise.
- deleted 11mo ago[deleted]
- cm2187 11mo agotbh it is lame for any program reading a text file to not support BOM. It's just one if.
- deleted 11mo ago[deleted]
- theblazehen 11mo agoThere isn't really any one "text file" though, the kernel looks for the first two bytes to match what "#!" corresponds to in ASCII. https://www.youtube.com/watch?v=J8nblo6BawU https://www.youtube.com/watch?v=J8nblo6BawU is some great watching on how "Plain text isn't that simple"
- SAI_Peregrinus 11mo agoUTF-8 is a text format with no BOM. Just like ASCII doesn't support a BOM. The BOM is a UTF-16 or UTF-32 thing, so "UTF-8 with BOM" is a binary file that happens to contain some UTF-8 strings as well. Since it's not a text file, it makes sense that utilities expecting text files don't handle it.
- 11mo ago
- userbinator 11mo agoI wonder what the reason was for having the kernel handle this, instead of the shell? To allow programs besides the shell to execute interpreted scripts as if they were actual binaries? This is of course in stark contrast to dynamic linking, which is performed by a userspace program instead of the kernel, and much like the #!, this "interpreter"'s path is also hardcoded in dynamically linked binaries.
- stabbles 11mo agoIn a call `execve("/my/script", ...)` of course the kernel has to figure out how to run it, there is no shell involved. As for scripts vs elf executables, there's not much of a difference between the shebang line and PT_INTERP, just that parsing shebangs lines is simpler.
- ajb 11mo agoRemember how slow early machines were. An unnecessary additional 'exec' would have seemed very wasteful, especially as most of the initial uses would have been shell scripts, so you would be execing the shell twice.
- alexey-salmin 10mo ago> This is of course in stark contrast to dynamic linking, which is performed by a userspace program instead of the kernel I'm not sure what the contrast is. In both cases the interpreter lives in userspace, in both cases kernel finds it via the path hardcoded in the file: shebang for text, .interp for ELF
- coppsilgold 11mo agoThere is no security issue here. The file with the '#!' needs to be executable, and at that point it doesn't matter what it invokes because you made it executable. It could have shellcode in it or it could call python3 which can also execute shellcode. Or more likely, it would just be a malware binary which you deliberately gave permissions to and executed.
- NegativeK 11mo agoIt's a vulnerability via pathing, not a worry that the shebang script could be malicious. Someone may have dropped a malicious executable somewhere in the user's path that the shebang calls. The someone shouldn't be able to do that, but "shouldn't" isn't enough for security. Or maybe the relatively pathed executable has unexpected interactions with the shebanged script, compared to what the script author expected. Etc.
- 1718627440 11mo ago> You're using a suspiciously old browser Does the author not like Firefox or what?
- zahlman 11mo agoI'm on FF 145 and don't see this message.
- extraduder_ire 10mo agoThat page[0] should explain the reasoning behind the redirect. There's a separate redirect[1] if you have a never version reported but are missing headers. Seems to be a fairly custom attempt at keeping poorly written scrapers off the site. curl works just fine for me. (I found these redirect URLs by manually setting the user agent header in curl) Fun idea. What's your browser's user agent? You might have an extension that alters yours to not report your browser's actual version number. 0: https://utcc.utoronto.ca/~cks/cspace-old-browser.html https://utcc.utoronto.ca/~cks/cspace-old-browser.html 1: https://utcc.utoronto.ca/~cks/cspace-no-sec-fetch.html https://utcc.utoronto.ca/~cks/cspace-no-sec-fetch.html
- 1718627440 10mo agoYeah, I kind of understand the author, it just sucks to be on the receiving end. I use Firefox ESR 115, with this UA header: "Mozilla/5.0 (X11; Linux x86_64; rv:109.0) Gecko/20100101 Firefox/115.0"
- rurban 11mo agoThis insight came from perl5 btw.
- gbacon 11mo agoSee also the perlrun documentation[0]. This example works on many platforms that have a shell compatible with Bourne shell: #!/usr/bin/perl eval 'exec /usr/bin/perl -wS $0 ${1+"$@"}' if 0; # ^ Run only under a shell The system ignores the first line and feeds the program to /bin/sh, which proceeds to try to execute the Perl program as a shell script. The shell executes the second line as a normal shell command, and thus starts up the Perl interpreter. On some systems $0 doesn't always contain the full pathname, so the "-S" tells Perl to search for the program if necessary. After Perl locates the program, it parses the lines and ignores them because the check 'if 0' is never true. If the program will be interpreted by csh, you will need to replace ${1+"$@"} with $*, even though that doesn't understand embedded spaces (and such) in the argument list. To start up sh rather than csh, some systems may have to replace the #! line with a line containing just a colon, which will be politely ignored by Perl. Other systems can't control that, and need a totally devious construct that will work under any of csh, sh, or Perl, such as the following: eval '(exit $?0)' && eval 'exec perl -wS $0 ${1+"$@"}' & eval 'exec /usr/bin/perl -wS $0 $argv:q' if 0; # ^ Run only under a shell [0]: https://perldoc.perl.org/perlrun#-S https://perldoc.perl.org/perlrun#-S
- crest 11mo agoHistorically FreeBSD used to split the #! argument list fully which meant you could put oneline scripts in there. At some point too many ports had to struggle with Linuxisms that this was changed to only split it into command and first argument. The old behaviour can be accessed with a flag in the env wrapper.
- nmz 11mo agoWhat's suprised me is that there is a line length limit on #! for 256 bytes. You can have spacing after #! for some reason? POSIX does not mandate -S, which means any script that uses it will only work on freebsd/linux
- dhosek 11mo agoFor personal use, I used to write #!perl to avoid the hassle of keeping track of the location of the binary across differing systems (although now I wonder if this only worked in the shell I ran under WinNT4—it’s been ages since I’ve done that sort of thing and these days I tend not to use #! at all to run scripts.)
- p4cmanus3r 11mo agoThe authors previous article about "...(not) using #!/usr/bin/env whatever" doesn't sit well with me. "The only reason to start your script with '#!/usr/bin/env <whatever>' is if you expect your script to run on a system where Bash or whatever else isn't where you expect (or when it has to run on systems that have '<whatever>' in different places, which is probably most common for third party packages)." His very first point is how you should only use it don't know where to expect bash binary, when I feel like, while it's probably safe in most nix os', assuming it limits future enhancements by requiring that binary be in place. However unlikely it would need to or someone would want to move it.
- pta2002 11mo agoI've encountered systems that only have bash in /bin/bash, or in /usr/bin/bash, and it's a hell of a pain to have to fix every script when using different distros (I think it must've been an old Fedora and Ubuntu?). Nowadays, most distros are moving towards having /bin be a symlink to /usr/bin, so it's mattering less and less, but I see no reason not to just do /usr/bin/env which is supposed to be on the same place on every distro.
- b33j0r 11mo agoNever thought of it this way; isn’t it always safe to assume env is in PATH? Maybe `#! env <shell>` could be considered a DSL for hashbangs. My reasoning is that `/usr/bin/env` is the thing that seems to be hard-coded to a system path, in most cases.
- mbreese 11mo agoI do this all the time for python virtualenvs. Instead of calling ‘#!/usr/bin/env python3’ I will call ‘#!venv/bin/python3’. This way I don’t have to worry about whether the environment was activated or not.
- zahlman 11mo agoIf you're okay with running the script only from the parent directory of the venv (I guess you have things set up so that's the project root), fine. You never have to "worry about whether the environment was activated", unless your code depends on those environment variables (in which case your shebang trick won't help you). Just specify the path to the venv's Python executable. You aren't really intended to put your own scripts in the venv's bin/ directly, although of course you can. An installer will create them for you, from the entry points defined in pyproject.toml. (This is one of the few useful things that an "editable install" can accomplish; I'm generally fairly negative on that model, however.) If you have something installed in a venv and want to run it regardless of CWD, you can symlink the wrapper script from somewhere on your PATH. (That's what Pipx does.)
- mbreese 11mo agoI use venv's to keep library installs out of the global space. My python scripts won't run without those libraries, so running from the project root is pretty much required. I don't use the variables set by activating the venv, so that's not a real concern. I've done this for years to keep my individual projects separate and not changing activations when switching directories. I also make sure to only call `venv/bin/pip` or `venv/bin/python3` when I'm directly calling pip or python. So, yes -- you have to be in the root project directory for this to work, but for me, that's a useful tradeoff. Even when running code from within a docker container, I still use this method, so I make sure that I'm executing from the proper work directory. If I think that I need to run a program (without arguments), I'll have a short shell wrapper that is essentially: #!/bin/bash cd $(dirname $0) venv/bin/python3 myscript.py As far as running a program that's managed by venv/pip, symlinks are essentially what I do. I'll create a new venv for each program, install it in that venv, and then symlink from venv/bin/program to $HOME/.local/bin/. It works very well when you're installing a pip managed program.
- mzs 11mo agomore/older https://www.in-ulm.de/~mascheck/various/shebang/ https://www.in-ulm.de/~mascheck/various/shebang/
- tracker1 10mo agoKind of an adjacent aside, I kind of love how many languages now work with shebangs for creating user scripts... Go, Rust, C#, Deno (TS/JS) etc. The vast majority of my shell/user scripts are now in TS with Deno... The only thing I do wish worked better is that VS Code would look at the shebang to determine a file's language without a file extension. Also works nicely with project automation and CI/CD integration as well... You can do a lot more, relatively easier than having to have knowledge of another language/tool you aren't as comfortable with.
- tylergetsay 10mo agocombined with nix shebangs, it's amazing to just code and go little scripts with zero dependency setup