5 ms·
> The pipe operator, |>, passes the output of a function to the first argument of the next function. So, instead of writing something like: foo(bar(baz(new_fun
by marhee 4y ago
> The pipe operator, |>, passes the output of a function to the first argument of the next function. So, instead of writing something like:
foo(bar(baz(new_function(other_function(my_input)))))
You have the option to write:
my_input
|> other_function()
|> new_function()
|> baz()
|> bar()
|> foo()
The readability benefit is clear - you don't have to read the code "backwards" to understand what it really does.
—- (end of snippet)
This is a nice example of fixing a problem that does not need be there at all.
Just use intermediate variables instead of doing clever stuff. Benefits:
- 1 way of doing stuff
- variable names will help readers understand the intention of the function calls
- you can do explicit error handling for each call
Note: in contrast, in shell prompts the pipe operator is useful because:
- less typing counts here
- prompts are use-once mostly, maybe repeated (history & scripts), but readability is not a mayor concern
- commands deal with large amount of data often, hence storing in intermediate variables can be expensive or even impossible.
None of these points are valid for a programming language. (Unless of course you just fancy the syntax)
Btw I find the joke about “bearded wizards” not very inclusive towards women. It feels assuming that only men (can) be that wizardry.
- UlisesAC4 4y agoIf you need temporal variables to give context you have other problems with the naming of the function that you are calling in the first place. We have used temporal names between function calls not because it was more readable but because chain of methods was not an option as you have said.
- josevalim 4y agoI would find such principled approach to be unproductive. The same could be said about Python, the `.` is equally syntax sugar for passing the object as first argument. So if we take this Pandas code: df.sort_values('dep_date') .groupby('name')['duration'] .transform('cumsum') If we were to pass df as first argument in nested calls, I would find it less readable. But I would also find using variables in this case to be more noise than helpful: sorted_df = df.sort_values('dep_date') grouped_duration_df = sorted_df.groupby('name')['duration'] grouped_duration_df.transform('cumsum') The variables are just repeating some of the information found on the right side and IMO they end up getting in the way of understanding the whole pipeline.
- marhee 4y agoYou could reuse df (as the method presumably change df), in addition: - you have more room to comment what the intention of the code/call is - you have room to handle / check for errors For example, the "sort_values" and "groupby" in you example are obvious most readers. But "transform("cumsum") is probably obvious to you but I don't know the intention of the code. By reassigning to df it makes it clear that the returned value from each of the functions is a `df` (some kind of query builder I think). Actually rewriting you r code I discovered your groupby did a groupby and then selected a result column I think. So we would get: // sort result on dependency date df = df.sort_values('dep_date') // ... check if sort_values worked (i.e `dep_date` is valid column) // group result by name and select duration column durations = df.groupby('name')['duration'] // compute columns cumulative sum of durations sum = durations.transform('cumsum') I agree its much more verbose, so this won't work if you (just) want conciseness.
- josevalim 4y agoIf you reuse the variable, then you are not getting any of the alleged benefits of using variables. The same name tells little except it is perhaps the same data type and in Elixir you would also get this information from the module you are invoking: text |> String.split(“,”) |> Enum.join(“ “) You could also equally add comments between the lines in the dot example: df.sort_values(“dep_date”) # group and select .groupby(…)[…] Although most of the comments above are discardable (IMO) because they are restating the code. My point is: sometimes I will break out into variables to get some of the benefits you mention. But forcing all intermediate steps to assign to variables is as harmful as using “|>” or “.” exclusively and forgetting about variables altogether. If you need to add error handling, code comments, etc, you can break out of the pipeline as needed.
- h0l0cube 4y ago> won't work if you (just) want conciseness I think Jose's example provides the best of both worlds, conciseness and readability. > you have more room to comment what the intention of the code/call is There's nothing stopping each pipe expression from having a comment of its own if you lean towards literate programming, or you feel that the pipe function + arguments begs further explanation to unfamiliar developers. > you have room to handle / check for errors You can pipe your results into a validation function. Ecto, the go-to database mapper, has exactly this pattern. Failing fast is idiomatic Erlang/Elixir (rather than let the process live on with corrupt state) which means if the validation fails, it ought to raise an exception so the external caller can fix their call/request, or if it's a bug, the developer can be alerted to fix the code.
- h0l0cube 4y agoNo downvotes from me btw, but I think the author and just about most people using Elixir (which tends to not be a first language for most in the community) have used the classic-style of temporary variables for lack of any other option. Certainly in my own experience the pipe operator has eliminated those redundant expressions rendering the code more concise, expressive, and hence more readable. Many APIs in other languages emulate something similar with method chaining for this very reason (see Fluent APIs). Extracting variables out of if-expressions can still be useful for improving readability of long boolean expressions though.
- deleted 4y ago[deleted]