Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Lofkin
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
61.
▲
by
Lofkin
11y ago
For interactive exploration of huge and out of core datasets, python Blaze/Dask has data manipulation with OOC parallel dataframes , and python Bokeh has interactive webGL and downsampling for visualization for said dataframe. htt
62.
▲
Easy pure python parallel workflows (OOC and in memory)
(matthewrocklin.com)
3 points
by
Lofkin
11y ago
|
0 comments
63.
▲
Python Tools for Visual Studio 2.2 Released
(blogs.technet.com)
1 points
by
Lofkin
11y ago
|
0 comments
64.
▲
by
Lofkin
11y ago
Unless you can help here, I see alot of pre coded models and older samplers, but nothing with a flexible JIT for user extensible variab;es and autodiff for newer HMC and NUTS type samplers. Exception being STAN, but that is its own C++ mode
65.
▲
by
Lofkin
11y ago
I was wondering this also
66.
▲
by
Lofkin
11y ago
R has no native bayesian library like Pymc 3 (Must use stan which is c++). Also Python is better for ad hoc and agent based modeling and for out of core data with blaze and dask. So no, R is not ahead in everything.
67.
▲
by
Lofkin
11y ago
Great comment. BTW, you can make interactive visualizations in pure python with bokeh: http://bokeh.pydata.org/en/latest/ Also with Blaze, you can use Pandas (or even Dplyr) syntax in python to query Hive, Spark a
68.
▲
by
Lofkin
11y ago
AOT compilation is in the works. Also you might have been using features that numba didn't support yet. They just added more numpy ops, array allocation and vector ops, so your code might be working now.
69.
▲
A view from the boundary – JuliaCon 2015
(ljuug.tumblr.com)
2 points
by
Lofkin
11y ago
|
0 comments
70.
▲
by
Lofkin
11y ago
Just use numba and get fortran like performance by writing in python.
71.
▲
by
Lofkin
11y ago
Try converting to a numpy array, executing the operations and converting back to pandas. More details here: http://pandas.pydata.org/pandas-docs/version/0.16.2/enhancin... Also they recently added support for
72.
▲
by
Lofkin
11y ago
Yes, python is not the future...but no, Julia and Python's Numba is just as fast as FORTRAN.
73.
▲
Look ma, no spark! Pure python Distributed Cluster data analysis with dask
(continuum.io)
1 points
by
Lofkin
11y ago
|
0 comments
74.
▲
by
Lofkin
11y ago
He missed a critical option: you can write those loops in python and JIT them to C fast LLVM with numba: https://github.com/numba/numba
75.
▲
by
Lofkin
11y ago
For huge datasets, Python has distributed and out of core data structures: http://dask.pydata.org/en/latest/ https://github.com/ContinuumIO/blaze This is pretty unique, and works better than
76.
▲
by
Lofkin
11y ago
With Pymc3, you can code a bayesian generalization of most of those packages pretty easily. For bayesian models, it is easier than calling into C++ with stan in R. For everything else, one can call arbitrary R packages with Rpy2, albeit wit
77.
▲
by
Lofkin
11y ago
You forgot python, with package: (similar in that R needs parallel packages): http://matthewrocklin.com/blog/work/2015/06/26/Complex-Graph...
78.
▲
by
Lofkin
11y ago
Awesome!
79.
▲
by
Lofkin
11y ago
It is already possible to some degree: https://github.com/JuliaLang/julia/issues/9973 More robust support starts with package precompilation which is slated to go into 0.4 https://github.com/J
80.
▲
by
Lofkin
11y ago
You clearly haven't used pythons major data science packages. Scikitlearn and statsmodels are known as being specifically easy to use with a uniform API that disparate R packages lack(though this is somewhat fixed with caret). Furth
81.
▲
by
Lofkin
11y ago
Definitely.
82.
▲
by
Lofkin
11y ago
Take a look at julia!
83.
▲
by
Lofkin
11y ago
What is missing? If its something like shiny, Julia is well on way to obviating that problem: https://shashi.github.io/Escher.jl/
84.
▲
Escher – Build beautiful interactive Web UIs in Julia
(shashi.github.io)
128 points
by
Lofkin
11y ago
|
17 comments
85.
▲
by
Lofkin
11y ago
Aside from the written in python making it more natural and extensible and amenable to messing around with the models, pymc 3 can sample directly from discrete parameters and STAN cannot. Pymc3 has more sampling routines.
86.
▲
by
Lofkin
11y ago
More good python resources: http://web.bryant.edu/~bblais/statistical-inference-for-ever... Harvard Data science class, in python: http://cs109.github.io/2014/
87.
▲
by
Lofkin
11y ago
Ultimately Julia combines the strengths of both and much more. It is the future IMO, not R or python.
88.
▲
by
Lofkin
11y ago
Python has more consistent stats, time series and programming syntax (for the packages it does have), and better bayesian inference package (pymc 3 > stan). Python is also better than R for ad hoc statistical modeling and algorithim dev
89.
▲
An Introduction to Statistics with Python
(work.thaslwanter.at)
198 points
by
Lofkin
11y ago
|
34 comments
90.
▲
by
Lofkin
11y ago
Why can't it be used for general computing?
More ›