19 ms·
Some Insights from a Julia Developer
- tomoya 9y ago"i like julia too"
- j7ake 9y agoEverything sounds great about Julia but it's lacking sufficient critical mass to develop useful packages to make scientists and data analysts effective. At the moment the bottleneck in our scientific computing and data analysis workflow is not waiting for code to run but rather quickly inplementing, evaluating, and iterating different models on datasets.
- freyir 9y agoWhen I last looked at Julia, the language wasn't yet stable. It's not fun developing packages for a language that changes from release to release, so I don't expect their ecosystem to stand a chance until they get to v1.0. All the advantages for package developers listed in this article are moot while the language remains a moving target. It was said that v1.0 was due in early/mid 2017, but it looks like it's still a ways off. There is (or was) a window of opportunity for Julia to steal mindshare in academia from MATLAB and R, but it feels like Python is beating them to the punch. Of course, Python had a 21 year head start.
- kgwgk 9y agoI was going to add that Python has been "getting there" since 1995, but I see you have already edited you comment :-)
- sgt101 9y agoTalking to people in the community my impression was that 1.0 would appear when 1.0 appeared and that they felt that given the difficulties faced by other languages and the work required to produce a language that was fit for purpose in the modern context, a gestation time of more than a couple of years was reasonable. In terms of community I think that the economics community and private equity firms in particular seem to have picked Julia up enthusiastically. Having said that, I suspect that it will be three years before Julia becomes mainstream, and then it will be a minority choice for five or six more - if it is the big success it deserves to be!
- mkborregaard 9y agoWhat makes you say it lacks critical mass? I don't think that's accurate, but of course it takes time. Implementation time is important, yes - but julia is not just a fast-to-run-language, which is essentially the point - it is a fast-to-implement language.
- j7ake 9y agoAt least in the field of computational biology most people use R or python because of the wealth of statistical packages and biology-specific packages (biomaRt) that makes implementing and testing models on biological data much faster than in Julia. It's always this problem, if everybody is using R, it is hard to migrate to Julia. But in order to migrate to Julia you need people to stop using R and switch to Julia.
- rattray 9y agoIs there no Julia wrapper around biomaRt? Would it be challenging to create and maintain? (I don't mean to imply that _you_ should be the one to do it - I'm genuinely curious)
- ViralBShah 9y agoThere are over 1500 packages (pkg.julialang.org)in Julia now, with the ability to call Python and R, and any C and Fortran library written. Julia packages tend to be more Julian of course, and natural if you are using Julia. I say most of the basic stuff is there, but a few things remain. What kinds of things are you looking for that Julia doesn't have?
- stabbles 9y agoOne issue I have with Julia is that it is fast the second time you run your code. And developing code is mainly running things just once. In practice you might find yourself waiting on compilation a lot.
- azag0 9y agoThat all depends on factoring. If running your code once means evaluating some function hundred times, then that's no issue.
- williamstein 9y agoStrangely, the only mention of Cython is to point at that we have had less developers than Julia: "as evidenced by the over 500 committers to just the Base language, more than projects like Cython has ever had!"
- ChrisRackauckas 9y agoI don't find it strange: I wrote this to say why I like using Julia and point out what the community is missing, not as a comparison to every other JIT in existence. But if you want to know why I gave up on Cython, I'll lay it out for you. I tried it almost 2 years ago because some documents in a course had IPython notebooks which used it. So I did some standard scientific computing stuff like write some Runge-Kutta methods and yes its speed was fine (that was a pretty big part of the blog post: if you try hard and use the right tools you'll get pretty much the same performance anywhere). That's not the problem at all. The problem was extending it to be more widely useful in my own research. I wanted to make those same compiled functions also work with complex numbers to integrate spectral discretizations of a stochastic PDE (instead of the finite difference one from before). I found some SO posts like: https://stackoverflow.com/questions/30054019/complex-numbers-in-cython https://stackoverflow.com/questions/30054019/complex-numbers... https://stackoverflow.com/questions/27906862/complex-valued-calculations-using-cython https://stackoverflow.com/questions/27906862/complex-valued-... At that point it stopped looking like Python at all. I always found the SciPy syntax a little verbose since I had used a lot of MATLAB before (but I wanted to make this project not require a license to run) (this QuantEcon cheatsheet is a good demonstration of what syntax is like in my domain: https://cheatsheets.quantecon.org/ https://cheatsheets.quantecon.org/). But to make this kind of "complex or not" logic work in compiled Cython, I resorted to conditional compilation (http://cython.readthedocs.io/en/latest/src/userguide/language_basics.html#conditional-compilation http://cython.readthedocs.io/en/latest/src/userguide/languag...). These days I understand that what I created was essentially a multiple dispatch mechanism. Anyways, at around that time I started experimenting with other tools, especially Julia, because I really was getting frustrated whenever I had to "go beyond doubles" and write something that was extendable instead of a one-use script. Maybe there's some tricks I was missing, but I found it really hard in Python and MATLAB. Soon after, my PhD adviser and I started arguing about whether one of the properties in the simulation's solution was due to floating point errors. I couldn't convince him, so I wanted to write this integrator so it was fast and compiled, but allowed arbitrary precision so that way I could prove that it still existed even with very high precision. I couldn't find a page which explained how to do high precision arithmetic in Cython or Numba, so I completely gave up. Needless to say, I decided to re-write a small portion of this in Julia and it worked really well, pretty much instantly. Then I was digging around the Julia package listing and Viral pointed me to a big opportunity (https://github.com/JuliaDiffEq/ODE.jl/issues/64 https://github.com/JuliaDiffEq/ODE.jl/issues/64) and I have been developing a lot of Julia differential equation solvers ever since. Obviously YMMV, but after computing a lot without a "first-choice language" (between R, Python, C, MATLAB, Mathematica) for quite awhile, I kept with Julia because I didn't have issues when I hit less standard tasks, and I found it very easy to contribute fixes to other people's projects because it was just Julia code. This stochastic PDE integrator story is just one (significant) project that led me in this direction.
- jampekka 9y agoToo bad they somehow thought it was a good idea to make the syntax resemble MATLAB of all languages. Perhaps most of the nausea inducing warts could be worked around with some kind of transcompilation, although some semantic issues, such as one-based indexing, would remain. It'll be a sad day if Julia starts to get such popularity that high quality libraries will be Julia-only.
- azag0 9y agoOne-based indexing is not a semantic issue, it's a language-design decision you may disagree with.
- pjmlp 9y agoWhich happens to common across many languages outside C universe. Some languages, like the Algol ones, even have user defined indexing. So it is not neither 0 or 1, rather whatever the min value of the index happens to be.
- PeachPlum 9y agoJulia 0.5 introduced support for any indexing scheme you care to invent. 1-based 0-based 20-based https://docs.julialang.org/en/latest/devdocs/offset-arrays/ https://docs.julialang.org/en/latest/devdocs/offset-arrays/
- attractivechaos 9y agoThat is an ugly hack, mainly because it was not in the original language design.
- pjmlp 9y agoWhy ugly? It looks similar to what I know in the Pascal family, where indexes can be ranges or enumerations.
- 9y ago
- kevinalexbrown 9y agoThe packages-first attitude feels significant to me. A language that “users” enjoy but package developers also enjoy seems important. I hadn’t thought about language choice from a heavily package-development weighted perspective before. It seems obvious in retrospect though, which is probably a sign of something cool. A version of this would be: how can good package development be as easy as possible, and how can package use be as easy as possible? I haven’t done any serious work in Julia mainly because the python libraries are mature, good, and performant enough. I can’t speak for everyone, but for end users in science labs library support is perhaps the biggest consideration for language choice.
- pjmlp 9y agoIt is. Many of the language flamewars we do, tend to skip over the eco-system. Which is much more relevant that any language design issue. Specially given that the decision of choosing a language is a consequence of working with a specific tool, unless one is open to face some hurdles.
- mhd 9y agoIs there a good use case for Julia outside the "math" community, when your alternatives wouldn't be R or Numpy, but Ruby or C#?
- jernfrost 9y agoI use Julia for shell scripting. Remember how many ende up using ruby for this? Well Julia works even better. I don’t write big programs in Julia. It is just a kind of everyday tool. I am a C++/Swift/Objective-C developer professionally.
- fasquoika 9y agoFor the most part, no. The language and community are significantly math oriented. However, if you're someone who likes to learn languages just to broaden their horizons a little bit, Julia might be a good choice. It's one of the few languages in common use which has multimethods (the other major one being Common Lisp). Actually, in general Julia is surprisingly Lispy for a language with Algol-family syntax.
- stevelosh 9y ago> Actually, in general Julia is surprisingly Lispy for a language with Algol-family syntax. It's less surprising if you know there's still a version of the Julia REPL that works with sexpressions: https://youtu.be/dK3zRXhrFZY?t=5m58s https://youtu.be/dK3zRXhrFZY?t=5m58s
- parenthephobia 9y agoPerhaps it's a borderline case, but I've used Julia to make a "big data" database server, backed by a mmap'ed column store. Because the compiler's available at run-time, queries can be compiled into efficient kernels that can run in parallel over multi-gigabyte arrays orders of magnitude faster than MySQL or Postgres on the same hardware. Without Julia, I would have had to find some other way to compile queries into machine code at run-time, which for me would have ruled out all the languages you listed. In practical terms I would have abandoned the project.
- 88e282102ae2e5b 9y agoIt's great and all, but I can't justify switching languages for minor improvements over Python + numpy/scipy. I'd be abandoning: * My deep knowledge and experience with Python * My entire codebase * The ability to work on projects with colleagues who don't also switch * The certainty that when I leave my current job, someone will be able to pick up after me * Zero-based indexing I've started to do some work in Rust when it makes sense, since there's occasionally a compelling case and it's substantially different than Python. Incremental advances in programming languages just aren't worth it, and knowing that other people are probably coming to the same conclusion means I can't expect a serious community to ever arise around Julia.
- deleted 9y ago[deleted]
- mkborregaard 9y agoJulia supports indexing with arbitrary bases, and it generally works ecosystem-wide.
- Bromskloss 9y agoOh, it does? I've taken to saying that "zero-based indexing – that's how Julia broke my heart", because at the time, I didn't get the impression that there was anything else in sight. How do I index based on zero?
- wallnuss 9y agoIf you have a particular algorithm that is better expressed using 0-based indexing use https://github.com/JuliaArrays/OffsetArrays.jl https://github.com/JuliaArrays/OffsetArrays.jl
- Bromskloss 9y agoIt's not just something particular I need it for; it's everything.
- _Codemonkeyism 9y ago"is a good generic type-stable function for any number which has zero and + defined. That's quite unique: I can put floating point numbers in here, symbolic expressions from SymEngine.jl, ArbFloats from the user-defined ArbFloats.jl fast arbitrary precision library, etc." Is it really unique? Isn't that just a Monoid (e.g. in Haskell, Scala, ...) or am I wrong?
- StefanKarpinski 9y agoI think that unique is an overstatement there, but in the context it's a comparison to other numerical computing systems, where this kind of support for generic code is less common.
- ChrisRackauckas 9y agoYes,i it was an overstatement. But, while I do see it can be done with these Monoids, I haven't come across these kinds of numerical linear algebra libraries that allow generic code in other languages (here's an example of what I find in Haskell that requires double: https://hackage.haskell.org/package/hmatrix-0.14.0.1/docs/Numeric-LinearAlgebra-Algorithms.html https://hackage.haskell.org/package/hmatrix-0.14.0.1/docs/Nu...)
- fasquoika 9y agoWhenever I see Julia mentioned, I like to link to this blog post by Graydon Hoare (creator of Rust). https://graydon2.dreamwidth.org/189377.html https://graydon2.dreamwidth.org/189377.html
- Bromskloss 9y agoWhat would you say is the takeaway point for this discussion?
- pygy_ 9y agoQuoting the conclusion: ———— Julia, like Dylan and Lisp before it, is a Goldilocks language. Done by a bunch of Lisp hackers who seriously know what they're doing. It is trying to span the entire spectrum of its target users' needs, from numerical inner loops to glue-language scripting to dynamic code generation and reflection. And it's doing a very credible job at it. Its designers have produced a language that seems to be a strict improvement on Dylan, which was itself excellent. Julia's multimethods are type-parametric. It ships with really good multi-language FFIs, green coroutines and integrated package management. Its codegen is LLVM-MCJIT, which is as good as it gets these days. To my eyes, Julia is one of the brightest spots in the recent language-design landscape; it's working in a space that really needs good languages right now; and it's reviving a language-lineage that could really do with a renaissance. I'm excited for its future.
- fasquoika 9y agoThis is a bit broader a point than the original post of "why you should use Julia", but bear with me. Basically, most languages (and even more generally, most tools) assume that you'll have external ways of dealing with their limitations. For example, we use slow but productive languages in conjunction with fast but tedious languages. Or we have a DSL for testing. Or we have a 3rd-party package manager. The assumption that these external tools will be there is generally true, plus it makes the language designer's job easier, so it's generally a no-brainer for them to make it. However, this pushes the complexity onto the system and the user, and creates redundancy. There have been some attempts to forego this assumption and make a language that can do everything you need. This has led to the creation of Lisp, Forth, and Smalltalk, among others. (If you've ever wondered about the fanaticism generally associated with these languages, this is a major part of it.) Julia is attempting to be such a language, and in my opinion does a pretty good job of it.
- crwalker 9y agoJulia is my go-to language for numerical work. Compared to other solutions I've used like python + numpy + pandas, Matlab, Mathcad, heh even Excel, Julia is a breath of fresh air. Fast, clean, powerful.
- j7ake 9y agoI looked at your blog post on solving differential equations and it looks pretty attractive. I will install Julia and play around with this differential equation solver. At the moment R and python are awkward with differential equations, and I don't want to be married to matlab.
- domador 9y ago"...the majority of programmers are not developers." Could someone please explain the difference between the two terms? I've always used them synonymously.
- jventura 9y agoProbably he is saying developers as professional programmers..
- 3JPLW 9y agoYes, I think most folks do. He's making a distinction between programmers that are consumers of APIs vs. developers that create and maintain APIs.
- ChrisRackauckas 9y agoThanks for bringing this up. It may be more of a domain-specific use of the terms. In science and math, most of the people who are "programming" are not really that deep into programming. Most people go to a workshop on "here's how to use R to do some regressions", and they will know a small portion of the language and usually rely on some libraries to handle most of the heavy lifting. For example, the size of the community that actually develops code for analysis in bioinformatics pales in comparison to the size of the community that is doing "bench" biology and performing the analyses. There is a clear distinction because in many cases the training is very different (CS and math backgrounds, vs more science backgrounds) and the "average day" and research focus are different. You can swap out biology with pretty much any other field and see the same split and this is the "package user" vs "package developer" split in technical computing.
- Fomite 9y agoI've been pushing a notion like this when people think about (especially) languages used in scientific computing. There's something close to three types of users: 1. Developers. People interested in software development that just happens to be scientific in use. Basically anyone whose considered writing a package. 2. "The Fuzzy Middle Category". People who are pretty good at a language, know some software development concepts (testing, etc.) but whose primary focus isn't developing packages as much as it is answering questions, usually with a number of packages glued together with some interstitial code. 3. Users. "The steps I use to read a .csv file into R and then run a logistic regression are..." They are, essentially, invoking software commands, rather than programming. Now clearly, these three classes are blurry and people can move between them.
- T3RMINATED 9y agoJulia is the worst language since Go
- RivieraKid 9y agoI really like Julia overall but I'm undecided whether it's good as a general purpose language. Right now I'm working on a few-thousands-lines-of-code project and sometimes wish Julia was more like Swift: - Writing code with Nullables is cumbersome and verbose compared to Swift. - The object.method() notation is sometimes more readable, especially in more complex expression. Plus, in an IDE it works well with completion. - The ordering of types in a file matters. - Overall, Swift code looks a tiny bit more readable, cleaner. - No interfaces.
- tomsthumb 9y agoMy company decided to used Julia as a general purpose language for a number of projects around 2.5-1.5 years ago (I’ve been there 18 months). I firmly believe this choice had an opportunity cost somewhere in the 1-2M $ range for a sub 25 person company. The number of bugs in the language, which had no doubt since improved, and the lack of libraries, which leads to serious NIH (b/c really it’s just NI yet) was costly, and frankly dangerous in the context of security, and then generally with regards to reproducibility. At one point we had a tool which leaned heavily on macros and required a version of the language which no one no longer had on heir machine and had no download links on the internet anywhere. Luckily they are consistent in their version uploads and altering the download address to have that version happened to line up with what’s in their server, but running a company in something like that is tenuous at best.
- jernfrost 9y agoI get what you mean, but I kind of miss the other when working with one. For larger code bases Swift feels "safer", but Julia tends to feel more enjoyable and fast to try out things with. But yeah I REALLY wish the Julia guys can come up with a nice way of dealing with Nullable. But I actually don't miss the `object.method()` notation. I find it is so much more straightforward to compose things functionally when everything is a plain function. I feel methods just adds complexity to a language. For readability I tend to use the form `x |> f |> g |> h` if I need to make it easier to read a call like `h(g(f(x)))`. I agree completion is a bit nicer with `object.method()` but function completion works quite well in the REPL and Juno. I use `methodswith()` to locate relevant functions for a type. In some ways I think it is actually easier to deal with than in Swift, as Swift base classes have so many methods you can't find the stuff you are interested in quickly. I have kind of wanted Swift to be the solution for everything, but I see that the API design philosophy for Swift makes it difficult to turn it into a language which is as nice for data science as Julia. There is no focus on making stuff like matrix classes, multidimensional arrays, shell integration etc. Also I don't quite like that Swift ties me so much to an IDE. Writing Swift code with a plain text editor and a REPL isn't as nice as doing it with Julia.
- skybrian 9y agoI'm wondering about this part: "using this strategy Julia actually can produce static binaries like compiled C or Fortran code." Is this something that works now, something planned, or just speculation?
- mkborregaard 9y agoThe blog posts highlights this as something planned. It does sort of work already, but not completely.
- ViralBShah 9y agoIt does work. It is not user friendly just yet, and compiler work needs to be done - but yes, you can get binaries and shared libraries today. https://github.com/JuliaComputing/static-julia https://github.com/JuliaComputing/static-julia
- anon_342njlkesr 9y agoYeah, the language design of julia is brilliant (multiple dispatch, typing, llvm use, zero-cost abstractions, @code_native to see why your code is slow). This allows you to write fast code in julia, which is impossible in python (you can call into very fast C/Fortran libraries with nice bindings, though). On the other hand, I really hate the syntax. One-based array indexing (ok, minor), blocks ending with "end", and most importantly unicode support. Really, who thought that it was a good idea to allow symbol names that are unreachable on a standard US keyboard? WTF? This makes it supremely inconvenient to use libraries that happen to export symbols with crazy names. If you must, specify a name mangling scheme and let people who want to see crazy symbols use an IDE. Seriously, am I supposed to use a hex-editor when auditing julia code? I totally agree with the author that package discovery and uniformity are a big problem. Indeed, this is even moving in the wrong direction: More and more functionality is removed from Base and put into external packages. I would prefer a more "batteries included" style, with uniform quality and documentation for standard functionality (splitting into different namespaces is good, though), python-style. Second non-cosmetic problem is that the language documentation is atrocious, both from a completeness and pedagogical viewpoint. Pyhton is again the ideal to aspire to. But the point of the article stands: It is easy to write fast julia code, and the Julia (non-C/C++) parts of the Julia core form a nice tutorial for how "good" julia code looks like. If a function is not well-documented, look at the source code. In python, you quickly run into the C-wall (function is really implemented in C and the way the wrappers work is really non-uniform); in julia this is more rare, and the wrappers tend to be much easier to pierce (ok, julia does a ccall, figure out the code of the target; the wrappers are human-generated and readable and don't come out of a complex build environment). But yeah, you shouldn't use julia for systems programming until a good standard way for unmanaged code has been fixed, if ever (mixing types that are garbage-collected and types that are programmer-memory-managed, you want garbage-collected types for performance-irrelevant parts and programmer-managed types whenever precise memory alignment matters for your performance). Oh, and the nullpointer support sucks big time. My personal hope is that the python->julia interface improves. Then python users will be able to profit from fast, easily developed julia packages.
- mkborregaard 9y agoJust a response for a few comments, since you talk with confidence but could maybe do with a closer look: 1. You end with `end` in julia. 2. You can use indexing with any base, not just 1 - no performance penalty. 3. The julia repl comes with latex completions making it very easy to just type e.g. \sigma and get the sigma sign. 4. The package system is moving both directions - functionality is split into modules for interoperability, but are gathered in batteries-included metapackages again. Like typing `using DifferentialEquations` will load all 60 small DifferentialEquations packages (centrally documented), but you can easily plugin alternatives if you like.
- ACow_Adonis 9y agoSometimes, in my darker moments, I have the terrifying thought that one of the reasons that many users like R and made it popular (apart from the historical context of its now many libraries and being the main free version of statistical software), is specifically that it isn't robust and sensibly designed from a programming/analytical perspective. You can download a package, type in a preset command on a preset thing, and 95% of the time (its R, so its only ever 95% of the time), you get back an answer/number. And if an answer is what you're after, that's where the considerations stop. Was there an edge case? Is R's answer really correct? Did it wallop something in your search path? Do you have your profile set up differently to the author? Do all the dependent packages clash and silently depend on how you load them and in what version you did so? Has R coerced something in the background, or pattern matched your typo'd variable to something it shouldn't? Somewhere in your code did you hit on one of the thousands of gotchas? Who cares? It gave me an answer. I can give it to my boss or put it in a paper. I personally am all for faster, safer, stricter (while still being dynamic) languages. And i've tried to explain to many users that languages that produce errors when you do something you shouldn't are not your enemies, they're your friends. But I see many people every day aren't living that philosophy, they'd rather an answer than the right answer or a robust answer, despite what they'll tell you in plain english. With respect to Julia, someone like me might like it (well, if they got rid of the matlab syntax and just went back to the Lisp they copied and relabeled as 'Julia' :P). But perhaps many users don't want a faster, more robust language. Perhaps, many of them are comfortable with a simple but wrong answer. And for that purpose, I can't see where Julia wins out relative to R. As I said though, these are thoughts i have in my darker moments...
- x0x0 9y agoI think your fears are misplaced. R users love R because of ease of use. R does sometimes ignores ugly corner cases in favor of that ease of use (though I'm skeptical this damages the validity of the answer in anything like 5% of cases), but that's a side effect. sklearn and pandas are great, but you still simply have to be a programmer to use them, or at least much closer to a programmer than many statisticians want to be. Allowing non-programmers to do things like data <- read.csv(file='blah.csv') fit1 <- lm(outcome ~ var1 + var2, data) anova(fit1) summary(fit1) to read data, run a linear model, and get an anova and p-values is amazing, and massively widens the scope of people to whom these tools are available. This omission of most quoting, the magic inference of column names, etc etc all makes R much easier to use.
- AlexCoventry 9y agoI tried julia last year, and it was a nightmare of version skew. Has it improved in that regard at all?
- ChrisRackauckas 9y agoIf you need something stable, wait until a bit after 1.0, like you would with any Windows release. Right now it's expected that things will change and break with language updates, but the next one is the 1.0 which is the "we stop breaking things now"
- AlexCoventry 9y agoThat's what I was told about 0.4 -> 0.5 last year.
- ChrisRackauckas 9y agoBy who? Someone who isn't part of the development team? There was an issue set for v0.5 called "Arraypocolypse" that was meant to change a ton of things related to arrays, and that was known months (a year?) before v0.5 was out. Semvar is used for a reason: pre-1.0 is all breaking.
- AlexCoventry 9y agoPeople on gitter, IIRC... I'm sure there's a record of it somewhere. Half the packages I wanted to use were still on 0.4, half were on 0.5, and the combination was incompatible.
- bigJack9 9y agoJulia is really fun to work with.I really love the REPL, multiple dispatch and the way you easily introspect code upto native assembly code.
- rurban 9y ago> Julia's JiT is not like other JiTs, and it helps package development Julia's JIT is a simple plain method jit, the easy one. He doesn't describe the pro's and contra's of method jit vs tracing jit. In short, method jits explode in memory usage and forbid expensive optimizations. The advantages are of course as described easyness to work with, reproduce and debug. Most JITs start as simple method jit, and then advance to Tracing JITs. Esp. with performance orientated languages with a lot of vectorization potential. It's a great language. But the JIT will be improved sooner or later. Esp. with the memory-expensive type-optimizations.
- barche 9y agoGreat article! Coming from C++, I agree Julia feels a lot like "C++ done right", i.e. where C++ forces you to jump through hoops with the verbose template syntax, Julia does generics by default. Also, it's great not having to wonder for every function argument declaration if you should add *, &, && or any of the const variants.