23 ms·
Haskell Problems for a New Decade
- mlthoughts2018 7y ago> “ Many of the large tech companies are investing in alternative languages such as Swift and Julia in order to build the next iterations of these libraries because of the hard limitations of CPython.” I would be interested in evidence of this. I work at a big tech company in ML and scientific computing, and have close peers and friends in similar leadership roles of this tech in other big companies, FAANG, finance, biotech, large cloud providers, etc. I am only hearing about the adoption of CPython skyrocketing and continued investment in tools like numba, llvm-lite and Cython. At none of these companies have I heard any interest in Julia, Swift or Rust development for these use cases, and have heard many, many arguments for why Python wins the trade-offs discussion against those options. In fact, two places I used to work for (one a heavy Haskell shop and one a heavy Scala shop) are moving away from those languages and to Python, for all kinds of reasons related to fundamental limitations of static typing and enforced (rather than provisioned) safety. I mean, Haskell & Scala are great languages and so are Julia & Swift. But even though in some esoteric corners people have started to disprefer Python, it’s super unrealistic to suggest there’s large movement away from it. Rather there’s a large move to it. It reminds me of the Alex Martelli quote from Python Interviews, > “Eventually we bought that little start-up and we found out how 20 developers ran circles around our hundreds of great developers. The solution was very simple! Those 20 guys were using Python. We were using C++. So that was YouTube and still is.”
- tastyminerals 7y agoyep, it is the current state of things. However, it would be better to have all this huge ecosystem written in a single performant language. People trade perceivable simplicity for things that are more beneficial long term.
- fnord123 7y ago>However, it would be better to have all this huge ecosystem written in a single performant language. It's not the language, it's the linkage model. Python glues C components together and that's why experienced professionals don't blink when going all in on Python: the performance can be easily cranked up to C level.
- jacques_chester 7y agoWithin the scope of a linked C library, this is true. But when your Python code is shuttling chunks of data between libraries it's less so. Pure Julia has the advantage that it can perform global optimisations that Python-wrapping-C can't.
- mlthoughts2018 7y agoI disagree. In practice the things I can do between Cython, llvmlite and numba are very low effort, wide reach optimizations. Practically, I have never found a case when a feature of Julia would have made this much easier for me. As a result, the tradeoff of Julia continues to be 100% focused on the switching costs, infrastructure maturity, drop in equivalents of well worn Python libraries, etc.
- socialdemocrat 7y agoThe problem is that you need C/C++ expertise to build packages, which creates a barriers between users of packages and makers of packages. This is a big selling point for Julia. One is seeing much faster package development on Julia than on Python because you use the same language on both ends. Package users are often able to contribute to the packages they use. Also the Python-Linkage has major limitations. If you create anything that requires the user to pass a function defined in Python to the C/C++ layer you take a performance hit. Think about solvers e.g. For machine learning that is a big deal. In Julia you can write your own scientific models and do automatic differentiation and training on them. If you do the same in Python you get a huge performance penalty. C/C++ cannot really help you.
- 7y ago
- socialdemocrat 7y agoComplexity matters. Python originally had success because it was easy to use. It was one of the first languages I learned after C++ to just get stuff done. People are not leaving Python for Haskell, Rust, Swift and Scala because those languages are too complex to deal with. However what we see time and again is that languages that are easy to use and which offer real advantages gain adoption. Look at the rising popularity of Go as a good example. This is why I belive Julia has a real chance against Python. It hits all the right checkboxes: It is easy to learn, simple tools, while also being highly productive and giving great performance. Sure Python is not suddenly going to get knocked off the crown. But the fact is that Julia has far more growth potential. Packages are built faster as there is no need for C/C++. Packages are combined way easier due to multiple dispatch and lack of C/C++ dependency. This is a hard one to explain in a short text but Julia has a unique ability in how packages can easily combine. Hence with a couple of Julia packages you get squeeze out the same functionality as a dozen Python packages.
- short_sells_poo 7y agoI am a researcher at an algo trading shop. We are moving our core libraries from rust+python to Julia. It really is amazingly powerful and has great interop with python (eg seamless zero copy array sharing). In the past one had to muck about with cython/c/numpy api to speed things up, now one can just write the functionality in Julia and make it available in both ecosystems. Python will likely remain the premier language to do research in, but more and more work can move to Julia which is much better at anything not immediately vectorizable.
- socialdemocrat 7y agoGreat news for me! I actually do Julia training. My colleagues do Python training. But we don't seem to be quit at the critical mass yet to gain a lot of training requests for Julia. But I am an optimist. There seems to be a shift going on. 2020 could be a breakout year for Julia. A lot of my job the way I see it at the moment is to try to learn about people's experience switching to Julia to help explain to people why it would benefit them. How did you guy actually end up switching to Julia? Was it a careful analysis ordered from the top, or was there more like some Julia evangelists who bugged people until they tried it an realized it was actually quite useful? I used to work in the oil industry, and tried to convince people to try Julia. It would have been a huge advantage over our Python interface in terms of performance but it was a very hard sell. People are conservative. They are very reluctant to try new things. So I am always curious how other people pull off making a change happen.
- logicchains 7y agoI work in HFT and now use Julia for all my research, and a couple of my colleagues now do too. Personally I'd rather retire and farm goats than go back to having to write Python professionally: there's just soo much that can go wrong that doesn't happen in a typed language, so much unnecessary stuff you have to keep in your head when coding. It also seems incredibly counterproductive to use a language that's 100x slower than necessary just because it's the only language some people know; the difference in research speed between having to wait one second for a result and ten minutes is massive. Of course, HFT is a somewhat different usecase than pure ML, as we work with data in a format that's rarely seen elsewhere (high-frequency orderbook tick data). Python's probably less painful for working with data for which somebody else has already written a C/C++ library with a nice API, as then you don't need to write your own C/C++ library and interface with it. My choice is either: write Python, and the research will take 100x longer, write Python and C++, and the development will take 2-4 times longer, or write Julia, and get similar performance to C++ with even faster development time than Python.
- mlthoughts2018 7y agoOne of the close colleagues I was alluding to also works in HFT and they do in fact use Cython libraries built in house for extremely low latency applications, order book processing, etc. Their main alternatives are coding in C++ directly or wholesale switch to Rust, but they prefer Cython. I know they evaluated Julia and found it entirely intractable to use for their production systems.
- logicchains 7y agoI'm curious why they found Julia intractable. In my experience it's much quicker to write than Rust, C++ or Cython. It's also much more expressive than Cython. Is it because they tried embedding it in C++? That can be painful, because it needs its own thread, and can only have one per process, but it's certainly doable.
- mlthoughts2018 7y agoI’m not sure what you mean by saying Julia is more expressive than Cython, given that Cython is as expressive as C. In this shop’s particular situation, it’s mostly the switching costs to Julia that cause it to lose the debate. The firm has lots of systems software, data fetchers, offline analytics jobs, research code, etc. With Python & Cython, they easily write all of it in one ecosystem, build shared libraries that span all these use cases, rely on shared testing frameworks, integration pipelines, packaging, virtual envs, etc. If Julia offered some kind of crazy game changer advantage that required a huge amount of effort to get in Python/Cython, they might consider breaking off some subsystem that has to have new environment management, new tooling, etc., and is not sharable across as many use cases. But there is no such case. They might get some sort of “5% more generic” or “5% benefit from seamless typing instead of a little rough around the edges typing in Cython”, and these differences would never justify the huge costs of switching or the missing third party packages that are heavily relied on. I always like to remind people that in any professional setting, ~95% of the software you write is for reporting and testing, and 5% at best is for the actual application. Out of that 5%, another 95% never has serious resource bottlenecks and taking care to write super careful optimized code for the 5% of the 5% can be done in nearly any language. Choose your ecosystem based on what best solves your problems in that other 99.75% of cases. This is especially true in HFT and quant finance, which is why so many of those firms use Python for everything except the 0.25% of the code where performance is insanely critical, they just use anything that super easily plugs into Python, usually C++ or Cython.
- agumonkey 7y agoPython is a fad right now, i dont think this will last. And i appreciate python dont mistake my comment. It just happened to have a simple 'ui' and some people made nice libs in it. The language is too fragile to reach more imo. i'd bet on Julia because it has stronger roots even though few care about it.
- mlthoughts2018 7y agoSeeing how it’s been the standard in scientific computing and ML for decades (plural) at this point, I don’t see how it can be a fad. It obviously might get replaced by new languages and ecosystems, but that’s a huge difference than it being a fad.
- sweeneyrod 7y ago"nice libs" are worth a lot; consider Fortran. Although quite possibly nice libs that are often just bindings to stuff in other languages have less staying power.
- seanmcdirmid 7y agoJulia is getting a lot of adoption outside of the tradition CS crowd, so data scientists who would otherwise might be using R or Matlab, or Mathematica. I’m not really sure who is adopting Swift yet, beyond iOS devs.
- socialdemocrat 7y agoDespite the positive tone, what really struck me in this article was that they struggle with getting people to work on the compiler because it is too complex. It just confirms my main criticism against advance statically typed languages: They are simply too complex. It is why I have more faith in languages like Julia, because you can achieve a tremendous amounts of power, expressiveness and performance in a relatively simple language. Dynamic typing gives you simplicity in the type system which is impossible to achieve with static type checking. Long term this has profound implications. People can keep hacking on and expanding Julia because they can wrap their head around the language. You don't need to be Einstein to grok it. I think you may also see Clojure overtake Scala for the same reason. The simplicity of Clojure may in the long run overtake Scala although Scala benefits a lot from similarity to Java at a superficial level.
- adrianN 7y agoMost compilers are very complex because generating performant code on modern hardware is hard. I don't think the complexity of the type system make much of a difference. Of course many new languages (like Julia) outsource a large part of the complexity to LLVM.
- socialdemocrat 7y agoI completely disagree. The experience with Julia points in a different direction. Julia performance has has little to do with LLVM. If performance was that easy, you could just slap LLVM in the back of any dynamic language and poff by magic you have super performance. Most dynamic languages that have reached decent performance have done so using trace-JIT compilers. These are complex and require a lot of man hours to make. PyPy is a similar approach. It is hard problem to solve. Yet Julia spending considerably less man hours than either project and with smaller and simpler code base has managed to run circles around these two projects in terms of performance. Why? Because of smart language design. The pervasive multiple-dispatch design tailored towards JIT code generation has been key. It allowed Julia to get great performance with a very simple method-JIT compiler. These are much easier to make than trace-JIT compilers. But Julia is not alone. Go is another example of a language which kicks above its weight. It is a relatively simple implementation, yet has good performance and is enjoyable for most people to work with. Except those who really hate Go of course ;-) My bottom line is: LANGUAGE DESIGN MATTERS!!! You you design a smart and simple language you can get away with simple implementations and still get good performance. Sure LLVM is an important piece of the puzzle, but it serves no more important role than C does as the backend for Haskell IMHO.
- pron 7y agoThis post highlights my biggest problem with the language: it's about the how, not about the why. For example, dependent types. They add significant complexity to the language and are intended to help with correctness. But are they an effective way to achieve it compared to alternatives? The fact that even current formal verification research is mostly looking elsewhere seems to suggest that the answer is likely negative, so why focus on them now? Or algebraic effects, about which the article says, "These effect-system libraries may help to achieve a boilerplate-free nirvana of tracking algebraic effects at much more granular levels." But is that even a worthwhile goal? Why is it nirvana to more easily do something that might not be worth doing at all? Researchers who study some technique X ask themselves how best to use X to do Y, and publish a series of papers on the subject, but practitioners don't care about the answer to that question because they don't care about that question at all. They want to know how best to do Y, period. One could say that every language, once designed, is committed to some techniques that it tries to employ in full effect. But any product made for practical use should at least strive to begin with the why rather than the how, both in its original design and throughout its evolution. I work with designers of a popular programming language, and I see them spending literally years first trying to answer the why or the what before getting to the how. So either Haskell is a language intended as a research tool for those studying some particular techniques -- which is a very worthy goal but not one that helps me achieve mine -- or it's intended to provide entertainment for practitioners who are particularly drawn to its style. Either way, it's not a language that's designed to solve a problem for practitioners (and if it is, it doesn't seem that testing whether it actually does so successfully is a priority). I think it's important to have languages for language research, but if that's what they are, don't try to present them as something else. And BTW, at ~0.06% of US market share -- less than Clojure, Elixir, Rust and Fortran, about the same as Erlang and Delphi, and slightly more than Lisp, F# and Elm -- and little or no growth (https://www.hiringlab.org/2019/11/19/todays-top-tech-skills/ https://www.hiringlab.org/2019/11/19/todays-top-tech-skills/), I wouldn't call Haskell a "popular language," as this post does, just yet.
- patrec 7y agoExactly. I find it revealing that having some moderately visible successful software does not feature on this list. Even ocaml does, IMO, better here (bits of mirageOs for non-linux docker; Xen; unison, coq, at least one successful ocaml only company) vs what, pandoc and shellCheck?
- thrower123 7y agoHaskell's main problem is that it is over-zealous, and built on axioms that make it a poor fit for real people writing real code. Everything else flows from that core issue. Functional purity and lazy evaluation are interesting, but when you can't toss a printf debug or log statement into a function without changing function signatures all the way up, it's not going to be popular. Pragmatic, sloppy languages will always be more popular, because they are more forgiving.
- cosmic_quanta 7y agoTo be fair, you can printf debug anywhere (Debug.Trace module). I'm the "print to debug" kind of guy, and I can debug in Haskell just like in Python using this. I think the issue you have stems from the fact that people can't transfer the knowledge from C/C++/Python/wtv directly. One of the standard way to build programs in Haskell (mtl-style) allows you to perform IO actions almost anywhere, provided you're willing to play with monads. But this is so different from anything most people have ever seen that I understand why it can be frustrating.
- thrower123 7y agoAny idea when this was added? I am relatively certain that it was not possible in Haskell 98, and admittedly most of my Haskell experience is a decade old. I have picked up three well-regarded Haskell books over the past decade trying to get back into the language, and don't recall any mention of this capability. That's another issue, the dearth of anything like tutorials or books on how to do something practical with Haskell, rather than the whirlwind tour of language features. I keep being told that serious software is being built in Haskell, but the knowledge of how to get from university-grade code to production-grade code is an unmarked wilderness that everyone must apparently navigate on their own.
- cosmic_quanta 7y agoThe license on the Debug.Trace module says 2001... That's all I know. I'm a big fan of the "Haskell Programming from First Principles" book. I read it on-and-off over the course of 2 years, and I think I'm an intermediate Haskell user at this point. The most production-grade Haskell software I've ever played with would be Pandoc (document conversion) and the Yesod family of webdev libraries. Playing with Yesod is a bit scarier because it uses some Template Haskell magic. On the flip side, there's a great user guide (the "Yesod book").
- coolplants 7y agoBig issue with Haskell adoption is the number of hoops you have to jump through to write efficient code in it. Haskell abstracts away the idea of a procedural VM implementing the code underneath, but you often need that exposed to finagle your code into a form that’s efficient. Haskell programming ethos is opposed to that, the language wants people to focus on the semantics of computation not the method. Yet we must be able to specify computation method if we want our programs to operate efficiently. The truth is that the Haskell compiler will never be able to automatically translate idiomatic Haskell code into something that’s competitive with C in the common case. once you start finagling your haskell code to compete with C you’ve negated the point of using Haskell to begin with.
- the_af 7y ago> The truth is that the Haskell compiler will never be able to automatically translate idiomatic Haskell code into something that’s competitive with C in the common case. Is this something you have experience with, or are you just theorizing? I know there are pitfalls to Haskell development, but to assert that "the common case" has uncompetitive performance seems wrong to me. My problem with this kind of assertions is that often -- not saying it's your case, that's why I'm asking -- they come from people who imagine how it must be (or read it in a blog), often coming from a place of never really having tried to solve something with the language, and this sadly percolates into popular wisdom ("Haskell is not good for this or that"). Something similar happened to Java's allegedly "poor" performance, way past the point Java software sometimes outperformed C in many areas.
- jerf 7y agoI think it's a common problem with almost all, if not all, languages that come to us out of the 1990s that they are fundamentally built on the presumption that "someday" compilers may save us. Well, "someday" is pretty much here, and they've all failed, as far as I'm concerned. (If you're happy with 10x-slower-than-C and substantial memory eaten to get even that fast, YMMV, but it's not what was hoped for. No sarcasm on that, BTW; sometimes that's suitable, just like sometimes 50x-slower-than-C is suitable. But it's not what was hoped for.) Haskell can do some pretty tricks with specific code if you tickle it right, but if you just write general, normal code, it's faster than Python (a low bar, Python qua Python is nearly the slowest language in anything like common use) but not a generally "fast" language. "Someday" is here; if it's not "generally C-fast" today, I see no particular reason to believe it will be in the future either, especially since GHC seems to have spent a great deal of its design budget. It would be interesting to see someone's take on Haskell, but where they write the language from the very beginning to be a high-performance language. There's probably multiple points in this design space that would be possible. Arguably Rust is nearly one of them, and getting even closer in the next few years.
- eli_gottlieb 7y agoI'm going to give the same Haskell hot-take I gave on Twitter while at NeurIPS in December: the problem with Haskell is that you can port the awesome parts of pure, compositional abstractions over to Python, far more quickly than you can port really good domain-specific libraries and applications over to Haskell. My experience was that porting some nice ADT-based machinery to Python, using an available ADT library, took roughly one weekend, whereas porting the major domain-specific package I needed from Python to Haskell would take months to years. I love Haskell. I prefer to use Haskell when I can. Haskell is a language in which I can Achieve Enlightenment. But apparently, God help me if I need to do some heavy graph processing or CUDA numerics in Haskell.
- pron 7y agoBTW, ADTs are coming to Java real soon: https://cr.openjdk.java.net/~briangoetz/amber/datum.html https://cr.openjdk.java.net/~briangoetz/amber/datum.html
- ashilfarahmand 7y agoSo, uh, I was looking through past threads (this one in particular https://news.ycombinator.com/item?id=21578769 https://news.ycombinator.com/item?id=21578769) and I came across your comment: "I might be spat at or beaten up in the street for being Jewish - that happens in Brooklyn or LA anyway these days." And I was wondering, how has it been for Jews in the US in the past few years up till now? When did things start to get noticeably worse? I know about the rising Anti-Semitism and the particular incident in NY about a month ago.
- eli_gottlieb 7y agoIf you want to talk about that, you should contact me privately. I don't want to derail a nice thread.
- ashilfarahmand 7y agoI'm totally fine doing that but I don't know how. Can we private message?
- anentropic 7y ago> If you look at the top 100 packages on Hackage, around a third of them have proper documentation showing the simple use cases for the library. This is still very very poor, and one of the biggest problems faced by newcomers to the language (speaking as one myself...)
- platz 7y agoWhich library in the top 100 that you used had documentation (or lack thereof) that wasn't sufficient to understand how to accomplish the task at hand?
- weavie 7y agoSome of the libraries need you to click into the docs for specific modules to get the full documentation for them. This confused me when I first set out. eg. Postgres Simple. This is the page you get to at first when you search for it: https://www.stackage.org/lts-14.21/package/postgresql-simple-0.6.2 https://www.stackage.org/lts-14.21/package/postgresql-simple... It looks like there is no documentation, just a list of modules. Clicking on the Database.PostgreSQL.Simple module gets you to the docs: https://www.stackage.org/haddock/lts-14.21/postgresql-simple-0.6.2/Database-PostgreSQL-Simple.html https://www.stackage.org/haddock/lts-14.21/postgresql-simple...
- whateveracct 7y agoThis is pretty standard practice. The landing page for a package is never where examples live (except for the README). Haddocks are used for examples & tutorials, and the top-level module or a documentation-only module ending in ".Tutorial" are often provided.
- weavie 7y agoSure. It just took a bit of getting used to, but could be the reason why a lot of casual developers looking at the language have the impression that the documentation is poor.
- siraben 7y agoMost undergraduates take a compiler course in which they implement C, Java or Scheme. I have yet to see a course at any university, however, in which Haskell is used as the project language. [...] The closest project I’ve seen to this is a minimal Haskell dialect called duet. Ben Lynn has been working on such a Haskell compiler[0], self-hosting and with IO too. It's a little unusual in its approach to compilation (targets combinatory logic and executes via graph reduction), but it's a great example nonetheless, from which one could build an undergraduate course. Code generation is a difficult issue, compiling a lazy functional language to assembly could easily fill up two semesters of classes alone. The VM in C is really small and the compiler can be built up incrementally[1]. [0] https://github.com/blynn/compiler https://github.com/blynn/compiler [1] https://crypto.stanford.edu/~blynn/compiler/quest.html https://crypto.stanford.edu/~blynn/compiler/quest.html
- arianvanp 7y agoWe have https://hackage.haskell.org/package/helium https://hackage.haskell.org/package/helium at our university which is being used for teaching. It's pretty close to Haskell 98 compliant
- tluyben2 7y agoI used this or similar (I forgot the name back then) when I was at UU decades ago; it was quite pleasant to work with as a teaching tool compared to ghc, esp at that time.
- vanusa 7y agoI have yet to see a course at any university, however, in which Haskell is used as the project language. Given the language's inherent complexity (and tortured history) -- is the task of writing a Haskell compiler really suitable for the undergraduate curriculum? That is, for "mere mortals". Not the kind who take tons and tons of graduate courses. In which case we'd be speaking of graduate-level courses.
- przemo_li 7y ago
- mark_l_watson 7y agoGreat forward thinking wish list! I have a strange relationship with Haskell. I love Haskell repl driven development of pure code for playing with algorithms and ideas. I use Haskell in this use case much like I use Common Lisp and Scheme languages. However, I am not very good at writing impure code and the interfaces between pure and impure code. To put it more plainly, I enjoy the language and I am productive, but I suffer sometimes figuring out other people’s Haskell code.