6 ms·
Choosing Julia, Matlab, Python or R in economics?
- BrandonS113 4y agoAn interesting writeup. Old versus new. So Julia is the language for those doing new things, Python and R for those with specialized applications and Matlab those stuck in legacy teams.
- forgotpwd16 4y agoThe documentation, syntax, and libraries will be better with examples backing the arguments made. Also it will be nice if we had the code sources used. For environment, Python has Spyder that is similar to MATLAB desktop and RStudio.
- Sukera 4y agoThe source code is linked at the very top, under the heading "Computation Speed": https://modelsandrisk.org/appendix/speed_2022 https://modelsandrisk.org/appendix/speed_2022
- forgotpwd16 4y agoThanks. Read it twice and still somehow missed it.
- bluenose69 4y agoAn important practical consideration is the language employed within the working group, or in the sub-field. For one-off code written by an expert, this doesn't matter much, and criteria such as those mentioned in the article might be paramount. In many cases, however, code starts off as a modification of something that already exists, and forms a seed from which other new things might spring. If everyone chooses their own language, collaboration can be hard. In an academic environment, cost is a serious factor that argues against Matlab. Even in a place with a site license, somebody must be paying, and many groups have grown weary of the expense, particularly because it locks their students into a language that might not be available to them wherever they go after graduation.
- wrp 4y agoAre not certain specialized packages like Gretl and GAUSS still the lingua franca of certain fields?
- gtsnexp 4y agoA high performance, open source Julia code library for economics: https://quantecon.org/quantecon-jl/ https://quantecon.org/quantecon-jl/
- BrandonS113 4y agoThey did mention that in article. " Not surprisingly, it has been adopted in high quality projects, such as Quantitative Economics with Julia, popularised by Thomas Sargent (Perla et al., 2022)."
- javitury 4y agoAs a typescript developer who is now starting his PhD in finance, I asked myself recently which of these languages should I pick, taking into account how easy is to code in each language. To this end I tried to code a quick probability calculator based on the poisson-binomial distribution. I found that R has the best libraries. It's the easiest language to use but it's also easy to write spaghetti code. Static analysis is a joke. Writing code in Julia is slower/more tiresome. I have to annotate all structs and function input types. In languages like typescript, I would see this as an investment because tsserver will later use that information for linting and autocompletion, so the net return is clearly positive. In Julia however static analysis is very primitive and limited, you feel like having to write code twice: once for annotations and then again for the implementation, with very limited cross-checking between the two. When writing functions, output types are so hard to annotate that I just skip them completely and only annotate input types. Type errors are usually caught at runtime. The runtime messages are helpful though. Unfortunately I couldn't even find a reputable Poisson-binomial package for Python so I stopped here. From my limited experience with python in other projects, I think that static analysis using pyright is mature and worth using (nets a positive return, although inferior to tsserver). This experience varies greatly between frameworks and packages, it's only worth to annotate the code if the underlying libraries also provide type hints.
- Sukera 4y ago> I have to annotate all structs and function input types. Can you elaborate on that? Argument types in function definitions do not help the compiler infer types - you can leave them completely untyped. It's really only required in struct fields, so that the compiler can know the size of your struct and doesn't have to box everything.
- jakobnissen 4y agoIt's not even needed in struct fields. If Python, Matlab and R are viable alternatives, then surely OP does not care about performance and can choose a completely dynamic style of Julia. (I don't recommend it though, having some idea of type safety is in general a good idea)
- 4y ago
- Sukera 4y agoInteresting comparison! The criticism of julia is certainly valid, especially in terms of documentation. However, I wouldn't say that core language documentation should necessarily be geared towards any specific field. In the implementation of the julia code there is a slight performance bug though (not sure if the authors submitted the article - the username certainly suggests it): function likelihood(o, a, b, h, y2, N) local lik = 0.0 for i in 2:N @inbounds h = o+a*y2[i-1]+b*h lik += log(h)+y2[i]/h end return(lik) end Only one use of y2 is marked @inbounds - in this case, either the second use too or the whole loop could be marked @inbounds (or the loop could be modified to run for all of eachindex(y2), with the first index being explicitly skipped).
- y_lin 4y agoThank you for your comment and looking at the code. This particular @inbound is not a bug. We tried the method you mentioned but it was slower. What surprised us was how little benefit we got from the @inbound
- Sukera 4y agoThat's surprising to me, as @inbounds should at worst be a noop - it shouldn't ever result in slower code. I sadly can't benchmark right now (no suitable device for the next few days), but I'll see if I can get back to this after that.
- freemint 4y agoYou might want to look into the zero function http://www.jlhub.com/julia/manual/en/function/zero http://www.jlhub.com/julia/manual/en/function/zero having a hard-coded 0.0 is a bit of a code smell and can sometimes cause the type instabilities i mentioned. For this code it should not make a difference as log retuns floats but ... it reads as not very Julian bc of that.
- eggy 4y agoI use all languages for different things, or based on my familiarity with using them for a specific domain of problems. I really like Julia, and I think once the libraries grow, it will be the leader here. It currently has some really specific, and useful libraries for certain tasks. I like R's ecosystem, and being a Lisper, R and Julia are my favorites. Personal bias: Python bores me. I will program in C or C++ if I really need something fast and specific. The bottom line: if you are smart enough to do high-level science and engineering, then you should be able to develop a few skills with each of these in your toolkit. If I had my druthers, APL/J would be my language(s), but don't get me started down that path! Honorable mention, and overlooked, Frink [1]. It is always open on my desktop with a folder of custom programs I have written for different things over the past decade. [1] https://frinklang.org/ https://frinklang.org/
- dig1 4y agoFrink looks great! I usually use Emacs calc [1] for things like this, and since I have Emacs always open, calc can be up in a few keystrokes. But I'll add Frink to my bookmarks for all those friends that are not using Emacs but wants something more powerful than a regular desktop calculator but not too powerful like Mathematica or Maxima :D [1] https://www.gnu.org/software/emacs/manual/html_mono/calc.html https://www.gnu.org/software/emacs/manual/html_mono/calc.htm...
- eggy 4y agoFrink beats everything including Maxima when it comes to unit conversion and dealing with mixed-unit calcs. You can do money exchange calculations like: 54 USD - 25 SAR, and Frink will spit out 47.33 dollar (currency) I use the graphic example program and customize it to print out graph paper. You can set the graphics to print at true measurement! Animations, input forms, graphics, all with mixed-unit calcs. You program in Frink's lang, which is a simplified Java. Reminds me of Groovy. I just like the fact it is a desktop IDE with choices on programming or simply converting units. The examples you can download off the main page are amazing and fun. The documentation is great too. I have been using Frink for a long time now, so I just made a donation to Alan who created and maintains the project. It has saved many a headache and error with carrying different units, and I've had fun with it. It's like a cool little programming environment. EDIT: It also runs on Android, and there is a GCJ compiled exe version for Windows (experimental).
- username223 4y agoIf the article interests you, this reference is also worth your time: https://tobydriscoll.net/blog/matlab-vs.-julia-vs.-python/ https://tobydriscoll.net/blog/matlab-vs.-julia-vs.-python/
- avip 4y agoIf visualization is required Matlab with the financial toolbox is a winner before the competition began. Actually the only non perfect side of matlab is pricing. It also has the best help ever devised in history of cs. U can build a math PhD. Curriculum from matlab help pages.
- goosedragons 4y agoI think R is better at visualization than Matlab. Ggplot2 and Lattice are very good and base R plots are still very capable.
- BrandonS113 4y agowell Matlab cannot read compressed CSV files, a problem when uncompressed files are very large. And it is very slow as the article showed. But visualisation. Please, R with ggplot is miles ahead of Matlab.
- patrick451 4y agoPersonally, I can't stand the grammar of graphics approach to plotting. But leaving that aside, the last time I tried ggplot, it lacked tons of features that matlab plotting has. For instance - zoom and pan - Data cursor/data tip. - Export data under cursor to a workspace variable - Interactive editing of the plot, with code export. Zoom/pan and data tips should be table stakes for a plotting package. A package that doesn't support that is not fit for use. IMO.
- jiggunjer 4y agoI often load and process binary data with Matlab into spreadsheet format, because it supports so many file types. Then process data frames in R or python. R for non-neural models and Python for GPU stuff. Visualizations also R.
- inawarminister 4y agoI'm using a combination of R and Python. R has more niche libraries (macroeconomics, climate economics, finance and game theory etc) and more open paper codes, Python is easier for me. My supervisor uses Matlab though. Haven't heard of anyone using Julia in Economics yet.
- BrandonS113 4y agoIf you work in economics, how could you miss Nobel price winner Thomas Sargent's QuantEcon? I know of lots of econ graduate students who use Julia.
- inawarminister 4y agothanks for the heads-up, wasn't aware he added Julia now. I will take a look this semester break.
- ChrisRackauckas 4y ago> wasn't aware he added Julia now. FWIW, it wasn't a recent addition. It was added in 2015 so it's quite matured. https://github.com/QuantEcon/QuantEcon.jl/graphs/contributors https://github.com/QuantEcon/QuantEcon.jl/graphs/contributor...
- nomilk 4y agoI was an economist doing econometrics in excel when in 2014 the datasets went being a few 10,000's rows to a few 1,000,000's rows. I found R easiest to learn as a CS outsider because it was less strict about package versions and installation requirements, which made it easier for a beginner to setup and get going. I learned it by googling every little step ('how read in csv', 'how create new column in data.frame' etc) until I had a ~40 line R script that did what I was previously doing by hand in excel. It ran in a few seconds and did what took excel about 10 minutes. A few years later I wrote an open source economics library in R: https://github.com/stevecondylios/priceR#pricer- https://github.com/stevecondylios/priceR#pricer- It converts between nominal and real prices, converts between 171 currencies, and has a few regex's for pulling numeric data out of text (e.g. salaries out of job descriptions). Some specific observations regarding the article: - Comparing computation speed seems a bizarre metric to care about. 6x faster matters on things that take minutes, hours or days, but less so for operations that already run in under 1000ms. Developer experience is usually more important IME. - The article mentions R library support is superior (which may or may not be true), but it would nice to see the most useful highly-regarded libraries from each language reviewed, or at least mentioned. - The statement "(Julia) does not have any historical baggage" seems odd since a language doesn't need to be old before people find problems with it. Example from recent HN post: https://yuri.is/not-julia/ https://yuri.is/not-julia/
- y_lin 4y agoComputation speed is only one of many criteria. It can be important. We have R code that takes hours to run, and for another paper it takes 5 minutes for the R code every time we change something. Speed is certainly important but not the only criteria. I agree, it would certainly be useful to compare the most useful libraries, for economists dynlm is fantastic in R. But we only have so many words in blog piece like this, and we didn’t have space I stand by the historical baggage. If you compare R (and the others) to Julia, there are so many bizarre and inconsistent language features they carry with them. Julia doesn’t have a lot, it might have made bad design decisions as Yuri argues, but it’s not historical baggage.
- nomilk 4y ago
- dash2 4y agoGood to read. Opinionated takes add value. I wasn't completely clear when libraries were being used. For example, did the authors use base R's `read.csv()` or `readr::read_csv()` or something else to load the CSV file? Similarly, though they say R has no facilities for dependency management, that's true in the base language, but the "renv" package does a reasonable job. I'd add that for very many people, computation speed is second order. I work with datasets of about half a million rows, which is large for economics, and I'm basically OK using R on a ten-year-old laptop. Sometimes a little pipeline management is necessary. Actually, pipeline software would be a good addition to the article. In my experience, R's "drake" is pretty overengineeered, and I suspect its successor "targets" is not much better. It would be interesting to compare the solutions in other languages. Another useful addition would be the ease of expressing simple data management tasks. So far my impression is that R with the tidyverse excels. There's not much simpler than dataset |> filter(row == cond) |> mutate(a = b + c) |> group_by(group_variable) |> summarize(x = per_group_computation(a,b,c)) where the code hews very closely to the problem domain.
- jstx1 4y agoThe largest portion of the article is spent on comparing performance, and then at the end they say that performance doesn't really matter (which is probably true).
- WoodenChair 4y agoWhen I was doing an economics degree in the 2005-2009 timeframe, Stata was the rage at my university. Have programming languages like Python and R displaced Stata and SPSS in the social sciences?
- BrandonS113 4y agoMy take from the article and the discussion below is that libraries is what makes R the best. Is that true? Will python and Julia then catch up? Or is is just best to use R, warts and all?
- ElectronBadger 4y agoI'm in the neuroscience (mostly EEG analysis), so my experience may not fit into economics field. I went from MATLAB only (and Stata for statistical analysis), through R (liked it a lot) and Python, to Julia-only data processing and analysis pipeline (rarely calling R for some required functions). To me Julia is the future of scientific computing, especially for current and ex MATLAB users.
- ekianjo 4y agoIs it me or they did some actual performance benchmarks without including any reference code? Because there are many different ways to load csv files for example even in R, so we have no clue what they used there. The same goes for "calculations".
- BrandonS113 4y agoThe source code is linked at the very top, under the heading "Computation Speed": https://modelsandrisk.org/appendix/speed_2022 https://modelsandrisk.org/appendix/speed_2022
- Sin2x 4y agoTLDR: R
- tpoacher 4y agos/matlab/octave Octave is wonderful, and getting better every year.