22 ms·
Not knowing the /proc file system
- 9dev 3y agoTangent: Looking at those code samples, I wonder whether coming up with the shortest possible, most cryptic variable names is somewhat of a sport amongst C developers. Are you guys still coding in Notepad and need to conserve keystrokes, or where does that reluctance to use proper names come from?
- mauvehaus 3y agoIt's party cultural. Look at the NT kernel mode documentation. In the Windows side of the world, it's perfectly normal for functions to take eight to a dozen parameters, each with a fairly verbose name. Example: https://learn.microsoft.com/en-us/windows-hardware/drivers/ddi/fltkernel/nf-fltkernel-fltcreatefile https://learn.microsoft.com/en-us/windows-hardware/drivers/d...
- mixmastamyk 3y agoThe early nineties were a lot different than the early seventies. One poster already mentioned the difference between paper and "glass" teletypes.
- dale_glass 3y agoJust inertia from the times where teletypes were a thing. As in, a mechanical printer that served as your console. When you have that as your interface you want to keep things terse. Even after that you had compilers with limits like 6 character names for a symbol. Those constraints went away, but names made to be comfortable for the users of teletypes and ancient compilers stuck around, and people made more of them because it fit the preexisting theme. And now it's just plain inertia, where it keeps on going because that's what it looks like in the books, so students imitate it.
- mhitza 3y ago> Just inertia from the times where teletypes were a thing. As in, a mechanical printer that served as your console. That's more than inertia, that seems to cross the threshold for ceremony. When was the last time that actual printers/teletypes were used, the 60s? The inwrtia part of it is people imitating what they already see on a project (cause why would you disrupt something as a newcomer on a project) and line length rules which I've seen come up from time to time along the years (not sure if a 120 character line length is enforced nowadays within the kernel)
- MobiusHorizons 3y agoThere are benefits to keeping code less far to the right on the line for easy reading. I review a fair bit of code these days, and I’ve never once wished someone used more characters for their variable names. Long variable names have a tendency to make the code they are part of wrap a lot and become hard to read. I think a lot of people think that a verbose variable name is always better, so they end up lazy and don’t try to figure out what makes the variable interesting in context.
- RhysU 3y agoAs a lingua franca, it's culturally reasonable. After you've grokked it the first time, e.g. "dgemm" is much more convenient than "double precision general matrix multiplication". This is akin to speaking aloud "DC" versus "The District of Columbia". Uses far outweigh learning. Consider even Python calls it a "dict" not a "dictionary" because the latter is a mouthful. Though what was wrong with "map" I often wonder.
- dale_glass 3y ago> Consider even Python calls it a "dict" not a "dictionary" because the latter is a mouthful. Though what was wrong with "map" I often wonder. A conflict with the map() function maybe?
- arp242 3y ago> Just inertia from the times where teletypes were a thing No. I like these names as it's just easier to read. fp or fd is short and to the point, and "filePointer" or variants thereof add noting except more characters to read, more "wall of text"-y code that's harder to scan, etc. And I don't really want to have a discussion about it as such; whatever your preference is that fine. I'll be happy to adjust to whatever works well within a team. But I wish people would stop spreading nonsense like this to invalidate other people's preferences.
- lifthrasiir 3y agoMaybe your tangent is pointing at a wrong direction, because I think there are also a sizable portion of such programmers in many other languages. It seems that anOverlyLongAndInformationFreeIdentifier is widely despised, but once names do have enough information contents, an exactly preferred name wildly varies even in a single code base. For example induction variables in Python generators tended to be shorter than average in my experience.
- cdogl 3y agoThis is particularly true in statically typed languages that require a type declaration. Here’s some go: func copy(f *os.File) { … } I think f is more than clear enough as a parameter name. The same can be said of variables where the type is easily inferred from the declaration or initialization.
- rascul 3y ago> anOverlyLongAndInformationFreeIdentifier Apple has one 82 characters. https://developer.apple.com/documentation/contacts/cnlabelcontactrelationyoungercousinmotherssiblingsdaughterorfatherssistersdaughter https://developer.apple.com/documentation/contacts/cnlabelco...
- xyzelement 3y agoAnd I would say that’s a good case where the long name is meaningful and necessary. On the other hand when you are dealing with a file pointer in a language that deals with file pointers constantly, “fp” is meaningful enough.
- School-Cotton 3y agoCNLabelContactRelationBiaoMei would have been a lot better. People can look up what that means when they need to.
- xyzelement 3y agoI haven't been a dev in a bit, but I'd say between having a longer variable name, and having to crack open the fucking dictionary.... I have a clear preference.
- jstanley 3y agoWhat specifically are you referring to? Do you mean like using `fp` for a file pointer or `fd` for a file descriptor? That is idiomatic, and I would consider calling them `filePointer` or `fileDescriptor` to be an obvious smell that the developer doesn't know what they're doing.
- williamcotton 3y agoI beg to differ... slightly! I'll do something like: int epollFd = epoll_create1(0); So yes, I don't think I have ever typed out "fileDescriptor" but I do label these things to make it more legible! OP has a point. We could do with more discipline when it comes to naming conventions in C. We're not using punch cards any longer. Somewhat related, there's another comment about the 80 char max... clang-format keeps to this. I'm kind of OK with this particular remnant of punch cards because I can get four editor windows open side by side on my ultra wide monitor. The nginx codebase is pretty freaking fantastic and I've learned a lot from just randomly browsing through the source, but I mean, come on: https://github.com/nginx/nginx/blob/master/src/core/ngx_array.c https://github.com/nginx/nginx/blob/master/src/core/ngx_arra... if ((u_char *) a->elts + a->size * a->nalloc == p->d.last) { What's p? Oh, ngx_pool_t. Object pool? Memory pool? What's a again? elts? Is that elements? If that read: if ((u_char *) array->elements + array->size * array->nalloc == memoryPool->d.last) { It's time for the old school conventions to change a little bit. C ain't going anywhere. Let's embrace a world where descriptive variable names are not subject to cost/benefit analysis like they were in 1977.
- arein2 3y agoIf you are working within a code base that uses fd and fp, then it's expected to preserve the style. However if it's a new project I would use filePointer and fileDescriptor to have consistent naming scheme along all variables.
- veqq 3y agoTo some extent, I like when people do it in other languages. It exhibits strong exposure to older programming material (where overshort vars abound), evincing great passion for programming, computer science etc. A lot of lisp material features single letter vars, for example. (Actually, var length in prod code bases seems related to scope. So toy examples with obvious context see single letters strewn about. But digging through the classic books will rub off on the budding programmer...
- dezgeg 3y agoBecause for some reason many C programmers still insist on 80-character line length limit and 8-wide tab indentation, so something has to give to fit all on the screen
- IggleSniggle 3y ago8-wide tab is still just 1 char of your 80 allotted. Just sayin'.
- MayeulC 3y agoNah, kernel coding style counts tabs as 8 characters for indentation purposes. It is also discouraged to nest conditional structures too deeply. Here's the very opinionated documentation: https://www.kernel.org/doc/html/v6.5/process/coding-style.html https://www.kernel.org/doc/html/v6.5/process/coding-style.ht... I also want to note that the 80 columns limit was bumped to 100, and is no longer strictly enforced: https://www.phoronix.com/news/Linux-Kernel-Deprecates-80-Col https://www.phoronix.com/news/Linux-Kernel-Deprecates-80-Col
- IggleSniggle 3y agoI wasn't entirely convinced by the opinionated style doc that 80 meant columns and not characters (of COURSE it means columns, the opposite would be nonsense, but sometimes I get stubborn), so I looked up the kernel repo and indeed, my comment didn't pass the sniff test OR the scripts/checkpatch.pl test, where the max_line_length is enforced using calls to expand_tab, which converts tabs into 8 spaces before checking the length of the line.
- enriquto 3y ago> Tangent: Looking at those code samples, I wonder whether coming up with the shortest possible, most cryptic variable names is somewhat of a sport amongst C developers. I definitely make a point of using only single-letter variables in most of my C and python programs. It is a very common usage in scientific computing. In math, all variables are single letters. Always. If a variable has more than one letter, you read it as the product of several variables, one for each constituent letter. When you are translating a formula that reads "y=Ax", you want to write something like y = A * x Writing this formula as output = operator * input or, god forbid, as something like output = linalg.dot(operator, input) is completely ridiculous to any mathematician. Mathematics itself used to be like that in ancient times. But after centuries of distillation, we arrived to the modern efficient notation. Some "programmers" want us to go back to the ancient ways, writing simple formulas as full-sized English sentences. But they will take single-letter variable names from our cold, dead hands! Of course, the first appearance of each single-letter variable must be accompanied by a comment describing what it is. But after this comment, you can use that letter as many times as you want. Encoding that information in the variable name itself would be disturbingly redundant if you use the variable more than once (which will be always the case).
- Terr_ 3y agoThat one lands near the margin-of-error for my sarcasmometer, but either way I'd like to emphasize (non-ironically) that the terse form is indeed very efficient... for someone repeatedly writing/copying it down by hand using a quill and ink. However that particular use-case has become dramatically less significant.
- enriquto 3y agoHeh. No sarcasm at all in my comment! (but I like to write in an over the top way...) Terse notation has nothing to do with manual handwriting. Modern math books and articles are still written by computer using a very terse symbolic notation, which has been developed during the last six centuries. Originally, the symbols +, = were shorthand abbreviations of Latin words. I guess computer scientists want to re-invent everything from scratch. How long will they need to evolve from "LinearAlgebra.matrixVectorProduct(,)" to the empty string? I hope it's less than six centuries!
- crabbone 3y agoYou should try Prolog. That's where single-letter variables are really, really common. C - not so much. In C you are more likely to see acronyms. It's also interesting that for some reason, macro names are spelled out in full, but function names are abbreviated. So, typical C looks like: LONG_SCREAMING_MACRO_NAME_(uh, oh);
- kjs3 3y agoSure...all us oldsters should be using things like ThisIsAVariable_i_like_it_very_much and spend a few hours making sure our bloated IDE highlights the variables we really like in mauve. Bonus points for using ChatGPT to generate the variable names.
- Animats 3y agoThe Linux /proc "file system" is kernel to user space communication hammered into the wrong form because "everything is a file". /proc is a system call with a fake file system API, and this matters. The sample code won't work reliably, because it assumes that the "files" won't change while being read. If you read /proc, you must "read" each file with one unbuffered kernel read to be free of race conditions. See [1]. The Linux kernel has no standard mechanism for delivering a variable-sized result from a system call. I/O comes closest to that, so this was hammered into the file system API. The dents show. /proc does not have standard file semantics. [1] https://stackoverflow.com/questions/5713451/is-it-safe-to-parse-a-proc-file https://stackoverflow.com/questions/5713451/is-it-safe-to-pa...
- Galanwe 3y agoI generally agree with your comment, especially coming from OpenBSD where /proc does not exist, for reasons you mentioned, as well as a purely practical one: let's not pressure the VFS for no reason. I think the biggest issue with /proc at the moment is that it leaks in core Linux components, and getting rid of it is becoming impossible. Some years ago I worked on adding `$ORIGIN` rpath support in OpenBSD and looked at how Linux did it. Mind you, the dynamic linker (ldlinux.so) grabs the running executable path from /proc... One thing though, you focus a lot on syscalls, but I think the challenge in recent years has been to find new ways to have kernel <-> userspace interactions _outside_ of syscall, which are cumbersome to use, rigid in structure, and practically speaking, there are only so many entries you can store in an IDT. The work on netlink is going in the right direction IMHO. It allows userland to communicate bidirectionally with the kernel, register on specific events, etc; while using a familiar tooling with bsd sockets that is convenient from both shell and programming languages.
- o11c 3y ago> Some years ago I worked on adding `$ORIGIN` rpath support in OpenBSD and looked at how Linux did it. Mind you, the dynamic linker (ldlinux.so) grabs the running executable path from /proc... `getauxval(AT_EXECFN)` is "better" but there are weird edge cases with `fexecve` where you only get `AT_EXECFD`, even ignoring the inevitable TOCTOU.
- 0mp 3y agoHyperfine is a really nice tool for this kind of benchmarking.
- tjungblut 3y agoWhen parsing the proc files, many people forget that process names can have spaces in them and that causes some very funny outputs. The Python script in this posts handles those correctly, kudos.
- eternityforest 3y agoThat's one thing I don't like about Linux. There are too many times where one has to parse things that are not a proper serialization format, and the things that ARE proper formats are a actually a mix of multiple. Sometimes you even see binary files being used to store less than 100 bytes. It would be cool if they just said "Everything here will be TOML" or something. But I like to stick with higher level tools anyway and avoid touching the low level stuff on Linux, so it's fine in practice.
- dadie 3y agoProbably off-topic: I may miss something but I have the feeling the shown C program should segfault a lot as the variable `fname` in `char* make_filename(const char*)` does not seem to be initialized.
- ab71e5 3y agoYeah it should be something like char fname[BUFSIZE]; Also better to pass that in as a variable I don't think you can return a char array allocated on the stack like that
- linuxftw 3y agoI use the /proc file system all the time. Often containers don't have ps or several other tools installed, so you can use /proc to find out how a process started, open file descriptors, open sockets, and a bunch of other information. In fact, many user-space tools like netstat are just purpose-built readers of things in /proc.
- tanelpoder 3y agoSame here, you can get pretty far with just catting, grepping, awk/sed/sort/uniq'ing all the various /proc entries. And not only /proc/PID/stuff, but also each individual thread (task) state from /proc/PID/task/TID/* too. Initially I wrote a python program for flexible querying & summarizing of what the threads of interest are doing (psn) and then wrote a C version to capture & save a sampled history of thread activity (xcapture) [1]. I ended up spending too much time optimizing the C code - as just formatting strings taken from /proc pseudofiles and printing them out took very little time compared to the kernel-dives when extracting things like WCHAN and kernel stack via the proc intereface. That's why I've since built an eBPF prototype for sampling the OS thread activity. The old approach still works even on RHEL5 machines with 2.6.x kernels without root access too :-) [1] https://0x.tools/#usage--example-output https://0x.tools/#usage--example-output
- williamcotton 3y agoThis is running in a barebones Ubuntu container on my MacBook as BSDs don't use /proc: root@74c03a282fbe:/# ed a ls /proc/[0-9]/status | xargs -n 1 cat | awk '/^Name:/ { name = $2 } /^Pid:/ { pid = $2 } END { print "cmd: " name ", pid: " pid }' . w prc.sh 132 q root@74c03a282fbe:/# chmod +x ./prc.sh root@74c03a282fbe:/# hyperfine --warmup=100 "./prc.sh" Benchmark 1: ./prc.sh Time (mean ± σ): 2.2 ms ± 0.3 ms [User: 1.3 ms, System: 2.7 ms] Range (min … max): 1.8 ms … 5.0 ms 880 runs Warning: Command took less than 5 ms to complete. Results might be inaccurate. Warning: Statistical outliers were detected. Consider re-running this benchmark on a quiet PC without any interferences from other programs. It might help to use the '--warmup' or '--prepare' options.
- aunderscored 3y agoI think you may have convinced me to learn more awk with this. I'd have constructed something horrible with sed or so to do this. And this is just much cleaner
- williamcotton 3y agoThe AWK Programming Language, Second Edition was just released. I'd say it's an instant classic! CSV support is slowly rolling out... you can compile Kernighan's nawk from source right now, gawk has it in trunk I think, and GoAwk has had support and is now also following the --csv design decision made by Mr. Kernighan.
- mzs 3y agoI like your style, but it's not quite right: $ echo /proc/[0-9]/status | wc 1 7 105 $ echo /proc/[1-9]*/status | wc 1 483 8778 It also will sporadically print error messages due to all the race conditions. Here's my stab at it: $ cat dumbps #!/bin/sh case $# in 1) ;; *) echo "Usage: dumbps user" >&2 exit 2 ;; esac 2>&- find /proc -mindepth 2 -maxdepth 2 -type f -name status -user "$1" | awk '{ while ((getline li < $0) > 0) { if (li ~ /^Name:/) { split($0, fn, "/") print fn[3], substr(li, 6) break } } close($0) }' $ ./dumbps "$USER" | grep -w 'bash$' 1616678 bash $ ./dumbps 0 | grep -w 'systemd$' 1 systemd $ It's pretty fast too: $ ls -ld /proc/[1-9]*/status | awk '$3 == "root"' | wc 445 4005 29367 $ time ./dumbps >/dev/null Usage: dumdps user real 0m0.001s user 0m0.000s sys 0m0.001s $
- mananaysiempre 3y ago> I’m using a custom function for reading lines from a file. [Listing: fgetLine()] It’s probably easier to just use getline(), it does basically the same thing and is in POSIX.1-2008[1]. [1] https://pubs.opengroup.org/onlinepubs/9699919799.2018edition/functions/getline.html https://pubs.opengroup.org/onlinepubs/9699919799.2018edition...
- spacechild1 3y agoAFAICT, getline() is not available on Windows, neither with MSVC nor MinGW. The actual solution is to use C++ instead of C :)
- toast0 3y agoThat shouldn't be a big deal for accessing the /proc filesystem as that's also not available on Windows, AFAIK.
- spacechild1 3y agoFair enough!
- deleted 3y ago[deleted]
- adamcc 3y agoYeah, so you are probably right. The custom definition is something that grew out of my response to the following two posts: A general review of the stdlib https://nullprogram.com/blog/2023/02/11/ https://nullprogram.com/blog/2023/02/11/ A review of scanf https://sekrit.de/webdocs/c/beginners-guide-away-from-scanf.html https://sekrit.de/webdocs/c/beginners-guide-away-from-scanf.... Originally I implemented `getLine` (terrible name!) as a way to get multiple lines of input from stdin in a relatively safe way. The implementation borrows almost all of its ideas from the `sekret.de` post. Because I knew the implementation and it was close to hand on my machine, I used it for this version and just swapped out `stdin` for just any old file. Edit: here's the implementation for reference, https://github.com/adammccartney/algorithms/blob/master/libs/adio/adio.c https://github.com/adammccartney/algorithms/blob/master/libs...
- chasil 3y agoHere is a little shell script to print out what is easily seen in /proc/*/cmdline. This requires a GNU xargs that supports NULL termination (or something compatible); as I understand it, this cannot be done with POSIX tools. $ cat shps #!/bin/dash for path in /proc/*/cmdline do p=${path#*/} p=${p#*/} p=${p%/*} case "$p" in *[!0-9]*) continue;; esac c="$(xargs -0 echo < "$path")" [ -n "$c" ] && printf %6d\ %s\\n "$p" "$c" done*
- tczMUFlmoNk 3y ago> This requires a GNU xargs that supports NULL termination (or something compatible); as I understand it, this cannot be done with POSIX tools. You can do this with `tr`, right? #!/bin/dash for path in /proc/*/cmdline do p=${path#*/} p=${p#*/} p=${p%/*} case "$p" in *[!0-9]*) continue;; esac c="$(tr '\0' ' ' < "$path" | sed '$s/ $//')" [ -n "$c" ] && printf %6d\ %s\\n "$p" "$c" done Those escape sequences are covered in the standard: https://pubs.opengroup.org/onlinepubs/9699919799/utilities/tr.html https://pubs.opengroup.org/onlinepubs/9699919799/utilities/t... (But please correct me if I'm missing something!) This also avoids (POSIX-)undefined behavior with `echo` in case one of the processes' `argv[0]` happens to begin with `-n` or any argument contains a backslash.
- chasil 3y agoI had read elsewhere that POSIX tr was not capable of this, but it appears that I am wrong.
- Calzifer 3y agoint s_isdigit(const char* s) { Name does not fit implementation. Should be "contains_digit". Also I would say almost any "issomething" method involving a loop has an opportunity for an early return.
- deleted 3y ago[deleted]
- a1369209993 3y agoAs written, it should be `s_hasdigit`, but it's actually wrong - it should be: int s_isnum(const char* s) { int result = (*s != '\0'); while (*s != '\0') { if ((*s < '0') || ('9' < *s)) { result = 0; } s++; } return result; } otherwise, it will choke on directories like `/proc/etc64/` or `/proc/net6/` if any such directory is added. (Plus other bits like early return, but that's not a correctness bug.)
- adamcc 3y agoThanks for this! Had not been aware of files such as `/proc/net6/` or `/proc/etc64/`. Will include this as an edit to the original post.
- a1369209993 3y agoAFAIK, no such directories exist in the linux proc code at this time. (Which means you won't get bug reports even in principle until, possibly years later, some such directories are added. At which point the code will start to to malfunction, possibly in bizzare and inconsistent ways, despite having worked fine for years by that point. This is a design principle I learned by experience.)
- deathanatos 3y ago> One thing that stands out is the call to ioctl - I have no idea why this is happening and it also appears to be causing an error. As far as I understand, the ioctl call signifies that the program is trying to do a terminal control operation. Dunno. It's not an error, per se. (The ioctl is literally erroring, but that's an expected possibility for the calling code, and it handles that.) The reason there's an ioctl is documented in the docs for `open`: > buffering is an optional integer used to set the buffering policy. Pass 0 to switch buffering off (only allowed in binary mode), 1 to select line buffering (only usable in text mode), and an integer > 1 to indicate the size in bytes of a fixed-size chunk buffer. You're not passing the `buffering` arg, so the subsequent text applies: > When no buffering argument is given, the default buffering policy works as follows: > * Binary files are buffered in fixed-size chunks; […] (That doesn't apply, as you're not opening the file in binary mode, so it's the next bullet that applies) > * Interactive” text files (files for which isatty() returns True) use line buffering. Other text files use the policy described above for binary files. That ioctl is the underlying syscall that isatty() is calling. It's determining if the opened file is a TTY, or not. The file isn't a TTY, so the ioctl returns an error, but to our code that just means that "no, that isn't a TTY". (And thus, your opened file will automatically end up buffered. The flow here is a good default, for each of the cases it is sussing out.)
- viraptor 3y ago> lseek(3, 0, SEEK_CUR) This part can determine if the descriptor is seekable or not. It's a noop if it is and returns an error if it isn't. > It was quite clear from the strace output that the overhead of the interpreter costs practically the same as running the program itself. This kind of idea keeps repeating and... it's misplaced. You can't use a high level API like glob and expect it will do the same minimum of work as your trivial implementation. This has nothing to do with the interpreter itself.
- chaps 3y ago/proc is amazing once you get the hang of it and get a good understanding of what's all in there. Especially if you're doing low level performance tuning. It's particularly helpful in larger infrastructures where tool variability means differences in available commands, their output, and cli options. I'm sure /proc iteration has its own issues of variability across large infrastructres, but I haven't seen it. It's a fairly consistent API. Or at least it was, since I haven't touched a large infrastructure in some time. When I got tired of `lsof` not being installed on hosts (or when its `-i` param isn't available) I ended up writing a script [1] that just iterates through /proc over ssh and grabs all inet sockets, environment variables, command line, etc from a set of hosts. Results in a null-delimited output that can then be fed into something like grafana to create network maps. Biggest problem with it is the use of pipes means all cores go to 100% for the few seconds it takes to run. [1] https://github.com/red-bin/lsofer https://github.com/red-bin/lsofer