14 ms·
Cheating at a company group activity using Unix tools
- smitty1e 5y agoFor doing work with JSON data, I'd add: https://stedolan.github.io/jq/ https://stedolan.github.io/jq/
- b6z 5y agoI don't understand what you mean. Half of the article is using jq.
- dsr_ 5y ago"After some digging, it was easy to find the HTTP request that pulled this information from the server. And it even had all the birthdates in the JSON!" HR needs to know this, but it shouldn't be available to random employees.
- devenvdev 5y agoWhy though? We use hibob, and anyone can find anyone's full name and birthday via UI anyway. Are there any compliance issues with this?
- dsr_ 5y agoDo you live in a place where there are anti-discrimination laws concerning employee ages? If not, perhaps there is no issue for you.
- actually_a_dog 5y agoAnybody who can effectively discriminate against an employee based on age probably has access to that info directly via some HR system anyway.
- toomuchtodo 5y agoIt’s PII and should be restricted to only those who require the data for their job (HR).
- mattrighetti 5y agoFor those interested in this topic I would suggest these incredible lectures by MIT [0], especially the data wrangling one. Lectures are hosted on YouTube, they are extremely valuable and easy to follow and they give a pretty good insight on a lot of Unix topics. [0]: https://missing.csail.mit.edu/2020/ https://missing.csail.mit.edu/2020/
- tzs 5y agoThe 'comm' command should be in there. With no options 'comm' takes two files, F1 and F2, which should be lexically sorted, and produces 3 columns of output. The first column consists of lines that are only in F1, the second column consist of lines that are only in F2, and the third column consists of lines that are common to both files. The option -1 tells it to not print column 1, -2 tells it not to print column 2, and -3 does the same for column 3. These can be combined, so -12 would only print column 3 (the lines that are in both files) and -13 would only print column 2 (the lines that are in F2 but not F1).
- rolandog 5y agoThis is new to me, but really useful. Thanks for sharing!
- unixbane 5y ago> Was it worth it? > 1 minute to do this > 1 minute to do that and 1 minute to introduce RCE vulns into company #589179283672's pipeline due to the "you don't understand the security implications of using fragile UN*X tools" problem which applies to anyone actually learning something from this article DAY OF THE SEAL SOON,
- Chris2048 5y ago> into company #589179283672's pipeline The article isn't describing such a scenario (load-bearing script).
- perryizgr8 5y ago> regex is so ubiquitous and valuable that if you don’t know it yet, you should learn it) Regex is one of those things I have to learn every single time I need to use it. I just can't seem to force myself to remember.
- amtamt 5y agoExtension of classic problem from "Programming Pearls" by John Bentley. Nice to see such pragmatism for one time problems.
- carapace 5y agos/John/Jon/
- unixbane 5y agojq ... | sed -E 's/([0-9][0-9]).([0-9][0-9]).[0-9]*$/\2_\1/' this fails for me since the jq output lines are surrounded by quotes. had to remove $. did i do something different or are we running different jq versions?
- ccalloway 5y agoMost of the justifications for using collections of command-line Unix tools are no longer valid today. Instead you should be using a proper programming language. Note that people who still do use complex solutions built from cat, head, cut, etc, and who know what they're doing, will typically either write a shell script (which won't be structured particularly differently from the equivalent Python or whatever) or will rely heavily on awk (itself a full-featured programming language, no easier to learn than any other scripting language), or both. One-liners which pipe text between four or five different commands are the equivalent of hand-soldered boards or bitwise arithmetic. Interesting to learn about for historical reasons but of no practical utility. The use of things like xargs and jq in this solution, difficult to invoke Unix utilities for doing things that are trivial in any reasonable language, makes this even more clear.
- 3np 5y agoI guess it depends on what you do. For a subset of tasks, shell scripting is a lot faster to implement than the equivalent python/js/go/ruby/rust.
- Kinrany 5y agoNo "proper programming language" is capable of ergonomically piping between programs. Shell is indeed very old and it's time for a replacement, but it's not there yet. Oilshell might get there eventually or at least spark interest in this area.
- ilyash 5y ago> No "proper programming language" is capable of ergonomically piping between programs. Solved. https://github.com/ngs-lang/ngs https://github.com/ngs-lang/ngs I'm the author. Frustrated with exactly this situation I created Next Generation Shell. It's a "proper programming language" on one hand but domain-specific for "DevOps"y scripting on another. So sane syntax, data structures, error handling, multiple dispatch on one hand but also syntax for running external programs, pipes and redirects. You are welcome!
- ccalloway 5y ago
- l0b0 5y agols | grep '.csv$' | xargs cat | grep 'cake' | cut -d, -f2,3 > cakes.csv That's quite a few antipatterns in one go. Unless you have a bajillion files the `xargs` is unnecessary, the `cat` and `ls` are unnecessary (and `ls` in shell scripts is a whole class of antipatterns by itself). You might want to use something like this instead: grep cake *.csv | cut -d, -f2,3 > cakes.csv
- XorNot 5y agoI'd go further and say don't parse CSV with plaintext tools because it's barely a plaintext format. Use a CSV library and save yourself heartache when someone drops a quoted string in somewhere.
- oblio 5y agoI don't understand why you're being downvoted. Parsing CSV with simple text-oriented tools is bad of an idea as parsing HTML with regexps.
- devenvdev 5y agoMost likely because it contradicts most people's experience, some CSVs can't be parsed with cli tools, but most of them can be, and it's much easier than writing code that does the same. So what the parent commenter says is true, just not pragmatic.
- oblio 5y agoIt is pragmatic with a minor tweak. https://csvkit.readthedocs.io/en/latest/ https://csvkit.readthedocs.io/en/latest/ https://github.com/BurntSushi/xsv https://github.com/BurntSushi/xsv Heck, even sqlite has some stuff to help with processing CSV files: https://www.sqlite.org/csv.html https://www.sqlite.org/csv.html
- philwelch 5y agocsvkit and SQLite were always too slow whenever I tried to use them for this sort of thing. I don’t remember if I tried xsv. I do think I ended up converting most of my CSVs to TSV, or exporting them as TSV in the first place.
- pkrumins 5y agoThe first example is super super bad here. Never pipe `ls`. When you feel like you need to pipe `ls`, then you know you want to use `find`.
- nixpulvis 5y agoI would recommend `find ... -exec`, but I still haven't figured out how to make it compose properly with other UNIX tools.
- revscat 5y agoYou may want to use file globbing instead. This is one I just used yesterday afternoon. I needed to search for a string in every .js or .jsx file in my project, but didn't want to include specs in the search. rg 'MySearchString' **/*.js[x]#~*spec* Voila. Note that this is for zsh, and you need to set the EXTENDED_GLOB option. But once you do you'll find yourself rarely needing to reach for `find`.
- nicce 5y agoIs it really better option? You can find "find" from every Linux system. For this sample, you would need to install zsh and enable globbing manually as well. It works only on your machine, but on the other hand there is chance learn something which applies everywhere.
- oweiler 5y agoTo be even more pedantic, you probably want to use a glob
- tzs 5y agoThere are a lot of times one only wants non-dotfiles in the current directory. The find would be something like find . -not -path '*/\.*' -type f -depth 1 What advantages does that have over 'ls' for that case?
- jasode 5y ago