20 ms·
Go+: Go designed for data science
- carbocation 6y agoI write a lot of Go, and I spend most of my time doing analysis (usually in R, occasionally in python). I'm interested to understand whether there was a specific motivating example that drove the creation of this new go-like language. This is Hacker News, so there definitely doesn't need to be anything beyond "I could, so I did." But if this actually solves some problem better than existing solutions, it would be cool to read about. Edit: Without a motivating example, it's hard to imagine that people will want to pickup a Go-like (but not exactly Go) language for data science.
- iujjkfjdkkdkf 6y ago> Without a motivating example, it's hard to imagine that people will want to pickup a Go-like (but not exactly Go) language for data science Exactly. I use almost exclusively python (including for data science- or ML really). I've been wanting an excuse to learn Go by doing a project with it. But learning some third Go-like language would be a tougher sell for me, unless there is really something it does better than python, because it still doesnt give me the benefit of learning Go. But like someone else said, "because you can" is usually a good enough reason to build or learn a new language, so I'm sure it's still worth it for many.
- resonantjacket5 6y agoIt seems it compiles down to golang usually? As in the read me they run ``` gop go tutorial/ # Convert all Go+ packages in tutorial/ into Go packages go install ./... ``` It's more like typescript for javascript than a completely separate language.
- CameronNemo 6y agoIf you are looking to learn a language specific to data science, Julia is fairly mature.
- joppy 6y agoThe first things I would look for in a data science language are multidimensional arrays, linear algebra packages, data frame and time series libraries ... none of which feature on this page.
- fractionalhare 6y agoYeah I'm confused. The only "data science" I can see here is the the title. How is list comprehension a data science primitive? How did this get over 4,000 stars on GitHub with a glaring lack of basic data science functionality? Is this used by actual practitioners?
- fixIt83 6y agoGitHub stars are bookmarks for me, not an indicator of usefulness. It does say it’s under heavy development. Maybe 4.3k+ GitHub users just want to make sure they get updates?
- diarrhea 6y agoRight next to Star is Watch, which would be much more suitable towards that, no?
- SamWhited 6y agoNo, Watch emails you a bunch. Stars just show up in a list so you can find it later. That being said, public bookmarks always seemed weird to me. Why not just actually bookmark it with your browser? Not that it matters.
- searchableguy 6y agoIf you are using github app, it's less friction to star it than open the page in the browser window and bookmark it.
- 6y ago
- daemonk 6y agoThere are a lot of numerical structures missing from this. Not sure if you can really advertise it as for data science without some kind of dataframe structure.
- fractionalhare 6y agoDataFrames? It doesn't even seem to have specialized array primitives like Series or NumPy. What the heck?
- micro_cam 6y agoI used to do a lot of machine learning code in go and think it has great potential as a compiled, static language with similar ease of development to python. However it is hard to get around the lack of operator overloading and (to a lesser extent at least to me) generics. I love the simplicity of the language and understand their feeling that operator overriding is too often abused but at the same time not being able to use algebraic operators for matrix and tensor libraries makes them really hard to use. The compacting garbage collector can also make it hard to pass pointers to memory to non go libraries which is key in data science. If this project could address those things I think it could have real potential
- elcritch 6y agoInteresting, I wouldn’t have thought of Go for ML. But I do share the enjoyment of static languages for Ml/data science. You might give Nim a look as it’s pretty practical for wrapping C++ code!
- matsemann 6y agoRef operator overloading. As someone not used to python but had to read a simple numpy script last week, I was stumped for a while on this line of code: X[y==1,0] Just that.I first thought, what would X[False,0] be? Since y was a vector, it obviously wouldn't be equal to one. Okey, but extracting that part, it looks like y==1 takes my vector, and replaces with an array of same size, with true or false for each element. Basically == is overridden to run a predicate over all elements. Okey, but then what does X[[True,False,False..],0] mean? Looks like numpy has overridden the [] so one can pass an array of booleans in addition to a normal index, and then it only keeps those elements corresponding to True indexes. Clever and useful when done daily I guess, but damn it was hard to understand those 9 characters as someone not well-versed in this domain.
- marcus_holmes 6y agoI never understand why operator overloading is said to make things more readable. If the meaning of an operator can change wildly with the operands then that's just confusing - you can't assume that '==' means what you think it means and you have to go find out what it means. In comparison, having an actual function name to clue me in on what something does is useful. Like, how is "X[y==1,0]" more readable in this case than something like "filterElements(arrayToFilter, arrayOfBools)"? (if I've understood what the original was trying to do, which I'm not sure I have). People seem to confuse "less typing" with "simpler", and that's not true. One of the great strengths of Go is that it rejects this and embraces true simplicity.
- JediPig 6y agonot even half baked. its a webpage with a single feature that is broken.
- chartpath 6y agoAgree with all the comments like "where are the nd arrays?" BUT, they have list comprehensions!! One of the main things I miss coming from Python.
- conradludgate 6y agoOnce generics get introduced into the language, you can write a generic map function which can take a []T and a func(T) U to return a []U. While it's not as elegant as a list comprehension, it's nicer than writing a for loop every time. Although, I can't remember what the performance impact of closures are in go, so this might not be a cheap operation.
- jhgb 6y agoWhy not something like a channel map? Give it a channel and a function and you get another channel with a goroutine running in the background.
- nemo1618 6y agoIn practice, no one will do this, unless there happens to be a function with the correct signature already available. The lambda syntax is so verbose that it's easier to just write the for loop. Another problem is that tons of Go functions return (value, error), and it's not clear how such functions should interact with a "map" function. Return all the errors in a separate slice? Stop at the first error? What if you only want to stop when the error is io.EOF? etc. I think we'll only see map/filter/reduce if the language is changed to specifically accommodate them. I've experimented with doing this myself, which people tend to view as heresy: https://twitter.com/lukechampine/status/1367279449302007809?s=19 https://twitter.com/lukechampine/status/1367279449302007809?...
- pjmlp 6y agoEasy, by using monadic operators. https://naveenkumarmuguda.medium.com/railway-oriented-programming-a-powerful-functional-programming-pattern-ab454e467f31 https://naveenkumarmuguda.medium.com/railway-oriented-progra...
- MrPowers 6y agoGo has a ton of potential in the data science space. A basic DataFrame library would go a long way. Doesn't have to be as full featured as Pandas. Just something that's maintainable and portable. I wrote a blog post a few months ago on the current Go DataFrame libraries (gota, qframe, dataframe-go): https://mungingdata.com/go/dataframes-gota-qframe/ https://mungingdata.com/go/dataframes-gota-qframe/. None of the current offerings are integrated with Arrow. An Arrow-backed Go DataFrame library that can read / write Parquet files could really jumpstart data science in Go (really data engineering in Go, which is where they should probably focus first).
- fractionalhare 6y agoI can see a lot of potential in Go for data engineering specifically, yeah. Those would probably be some very stable and performant ETLs. And the concurrency and network primitives would make it easy to develop libraries like Prefect/Airflow.
- MrPowers 6y agoYep, agreed. Go is a great language for AWS Lambda type workflows. Python isn't as great (Python Lambda Layers built on Macs don't always work). AWS Data Wrangler (https://github.com/awslabs/aws-data-wrangler https://github.com/awslabs/aws-data-wrangler) provides pre-built layers, which is a work around, but something that's as portable as Go would be the best solution.
- fractionalhare 6y agoLove awswrangler. I use that over boto whenever I have the opportunity.
- RSHEPP 6y agoWe use Go for our ETL, with some Python too. We are in the process of transitioning to Argo Workflows from a K8s CronJob/Job setup which has been pretty stable itself.
- andrewprock 6y ago
- iagovar 6y agoIf anyone is looking for an alternative to R or python, there's Julia already.
- CameronNemo 6y agoAlso Rust has a good datagrams library now, polars. Not as mature an ecosystem as Julia, but hopefully it improves in the future.
- hu3 6y agoI find Rust's borrow checker too clunky for exploratory work. It breaks my flow and imposes higher cognitive load. The slow compiler doesn't help either.
- Gibbon1 6y agoI tend to think the innately slow compiler is basically a fatal mistake. Rust will never be able to be used for large projects. It's not so obvious now because everything it's used for is tiny.
- CameronNemo 6y agoHopefully some of the work out of cranelift, gccrs, and/or rust-analyzer can be used to speed up compilation.
- Gibbon1 6y agoSeriously one hopes. I think of projects in C++ that takes half an hour to compile, in rust they would take half a day or longer.
- c-cube 6y agoWhat supports that statement exactly? If you split your big projects into crates it would recompile pretty fast. C++ can also take ages to compile (eg. compiling Firefox from scratch). Keeping your code modular to get decent compile times seems like a win win.
- kzrdude 6y agoThe env gop run shebang line is not posix-compliant; posix only requires support for a single argument in the shebang and this one has two arguments (gop run). </irrelevant unix nerd mumbling>
- nine_k 6y agoPOSIX is thirty three years old. Can we please consider certain modest improvements?
- 10000truths 6y agoThe latest revision of the POSIX standard is only 4 years old.
- pjmlp 6y agoAnd yet it is still all about writing CLI and server daemons, stuck in the early 80's timesharing computing world.
- lanstin 6y agoThat is an odd comment to put using HTTP onto a process listening to port 443 so it can be stored by way of sending certain bytes to a different process listening to a port.
- pjmlp 6y agoExcept that the application that is able to display and understand what those magical HTTP contents mean isn't part of POSIX.
- deleted 6y ago[deleted]
- m45t3r 6y agoThere is env -S that supports multiple arguments. This was always an extension available in BSD I think, and it is available now in recent versions of GNU's env.
- tpmx 6y agoSeems like there's a potential trademark risk if Google decides it wants to protect the Go trademark. https://news.ycombinator.com/item?id=20023137 https://news.ycombinator.com/item?id=20023137
- great_reversal 6y agoWhy can't you just build libraries to make Go a better language for data science? There's already Go support for a Jupyter Notebooks kernel: https://github.com/gopherdata/gophernotes https://github.com/gopherdata/gophernotes
- srer 6y agoWe could build such libraries, and people have built some. However the task at it's heart is a vast duplication of work, and while Go has a lot of things going for it, it doesn't seem enough to sway many data scientists into reinventing their wheels in Go. I don't blame them. Rewrites being difficult to justify or motivate when you already have a compelling implementation is part of the reason why we have significant amounts of FORTRAN77 code still kicking around today. It is also why for many things we opt to just write wrappers around existing C libraries to call them from other languages. It has many shortcomings, but overall I prefer the sharing of a library across languages, each with it's own bindings that can attempt to make it more idiomatic to that specific language. The Go culture/community doesn't favor this approach, the Python community embraces it.
- deleted 6y ago[deleted]
- umvi 6y agoI just barely picked up Go, and my first impression is that it's very... opinionated. It wants me to do if/else guards a certain way, you have to capitalize first letters of "exported" functions, it won't let me import `fmt` unless I use it, etc. I'm not sure I like it.
- philosopher1234 6y agoThe opinions are by design. By removing flexibility, you can increase uniformity. Instead of having 12 different styles of code, there can be 1. It removes cognitive load, so you can spend your mental energy on solving problems.
- umvi 6y agoAnd that's fine if go were the only language I ever used, but it's jarring going from unopinionated languages to a highly opinionated one. I have to have a special set of "go rules" in my mind to be sure to follow when using go which has the effect of increasing cognitive load. "oh right, go wants me to compress my if else clauses and put the brackets a certain way"
- philosopher1234 6y agoI don’t that appreciates the value enough. Yes, there is a cognitive load to learning the opinions of go (though gofmt keeps you from having to learn a lot of them, as do compiler errors) but that is true of any language. I also think the number of opinions you need to learn is far smaller, as you don’t have as big of a surface area to navigate when making design decisions about daily coding. I think they are net very positive.
- jy3 6y agoI hope you have the presence of mind to realize the amazing benefit this has for the entire Go codebase in existence.
- yashap 6y agoGo is a great language, but it seems terribly suited to data science. The popular data science languages are Python, R, Julia, and to a lesser extent Scala. They’re all extremely flexible languages, where you can easily write high level abstractions/DSLs, and they all have very strong functional programming support, because data science tends to be extremely functional. They also tend to be very concise languages. Go is at the complete opposite end of the spectrum - not flexible at all, it’s purposefully difficult and awkward to write high level abstractions/DSLs, there’s very poor functional programming support, and it’s very verbose. There are great reasons for these restrictions, they’re intentional design decisions, but they also make it a very poor fit for data science IMO.
- tmpz22 6y agoIDK if its Go's problem honestly. Data modeling is hard. Its hard for a reason. If a language like python makes it seem easy, its still hard but your perception and attitude towards it has changed because some of the busy work has been taken out of it - possibly in a way that costs you down the road. Let's be honest programming languages are the punching bags of developers.
- sabellito 6y agoNot trying to start anything, but what's functional about Python? It doesn't have/support tail recursion, a strong type system, pattern matching, immutability-by-default for lists and dictionaries. From where I'm standing, python has some features that kinda look like functional programming concepts, but overall is an OO imperative language, like Ruby and many others. My understanding for its preference from the DS community is due more for its library support in that domain.
- Nican 6y agoI think the appeal is with Jupyter [1] notebooks. Python is not about performance. Usually numpy (or other libraries) that does the heavy lifting on another language anyway. But having the Jupyter notebooks allows for intractability with the data. Make changes, and see how it affects every step after it. [1] https://jupyter.org/ https://jupyter.org/
- dm319 6y agoWhen I realised I couldn't divide a time period by a number or integer, the penny dropped that different languages excel at different things.
- icholy 6y agoAre you talking about `time.Duration`? Because you can definitely do that.
- robbyt 6y agoBut don't try to use BC dates with time.Duration, because they don't work!
- dm319 6y agoHas that changed? I couldn't before!
- TheDong 6y agoYou have been able to divide a time.Duration by an integer since before go 1.0. As you can see in the stdlib, time.Duration is just an int64 (https://golang.org/pkg/time/#Duration https://golang.org/pkg/time/#Duration). You do have to cast the integer to a time.Duration sometimes. Here's a playground showing cases where it works and cases where it require a cast: https://play.golang.org/p/6Pbqrz8ZZ3t https://play.golang.org/p/6Pbqrz8ZZ3t
- dunefox 6y agoThere are quite a few languages I would like to use before Go. Especially F# seems very interesting for DS.
- bachmeier 6y agoEven been wanting for some time to check out the F# R Type Provider: http://bluemountaincapital.github.io/FSharpRProvider/ http://bluemountaincapital.github.io/FSharpRProvider/ Unfortunately it appears to be Windows-only, and my curiosity hasn't yet reached the point that I'd boot into Windows.
- dunefox 6y agoSadly, it also doesn't work with .net core yet. Otherwise this would be a pretty convincing point for F#.
- Ambix 6y agoReally cool thing! When it will be ready for production use?
- donutloop 6y agoMany of these features should be part of the upcoming GO 2
- amelius 6y agoNobody in data science wants fragmentation. Therefore, any aspiring new platform would need to bring some serious benefits to the table. I'm not sure what they are here.
- unreliableNar8r 6y agoI wish them all the best in this but it seems like an uphill battle, and doesn't seem to have a clear use case to me. For lightweight to medium projects R and Python are so well supported it's hard to reject them as the null. If you're doing exploratory stuff and want visuals, it's the same story with Rmd and Jupyter. For more behind-the-scenes production pipeline stuff there is already Scala which has inroads with Spark. If you really want to use something new, Julia is starting to mature and has all sort of plotting and linear algebra support. To me it seems Go would aim more to compete with Scala I suppose? I suppose then it might come down to plotting. In terms of being a general-purpose DS language, I can't imagine using anything that doesn't have a clear strategy to A) get a dataset into a DataFrame or similar, B) get my collaborators a plot in a way that is quick and easy, and C) a lesser extent, some kind of notebook/reporting tool. They do say there is a lot of development going on but it seems like a space with a lot of great incumbents and a rapidly maturing up-and-comer in Julia. edit: typo
- quixoticelixer- 6y agoOh jesus christ fuck no
- InvOfSmallC 6y agoWhat's the point?