Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
dandermotj
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
dandermotj
10y ago
The latter
32.
▲
by
dandermotj
10y ago
I studied statistics - my point was that statistics is taught in a linear manner, starting with distributions and hypothesis testing (p-values) and then move onto more advanced treatments like Bayesian stats.
33.
▲
by
dandermotj
10y ago
You're certainly not going to study Bayesian statistics without knowing (or at least having studied) what a p-value is.
34.
▲
by
dandermotj
10y ago
It doesn't assume the Effiecient Market hypothesis - empirical studies of returns support random returns without the imposing a model (non-parametric tests). That's not to say returns are actually random, but in any given time ran
35.
▲
by
dandermotj
10y ago
The standard evaluation versions of all dplyr functions are available, just add an underscore to the end: filter_, select_, ...
36.
▲
by
dandermotj
10y ago
Worth the read if only for the meme of the dog at the end.
37.
▲
by
dandermotj
10y ago
Mark Cuban is talking about it - maximal hype has been reached.
38.
▲
by
dandermotj
10y ago
You're trying to GROUP BY on a distributed data store; your code is the problem, not Spark SQL. Use CLUSTER BY - it's distributed sibling. Query languages like HiveQL and Spark SQL were designed to look like SQL, but they're
39.
▲
by
dandermotj
10y ago
This is excruciatingly accurate.
40.
▲
by
dandermotj
10y ago
Big four employee outsourced to big bank. The lack of in-house knowledge is always troubling. Also screen scraping data and automating processes by having a program take control of the mouse/keyboard and click/type is a big seller
41.
▲
by
dandermotj
10y ago
The open source library simply helps a programmer put together HTTP requests. If an API crashes because it receives an unexpected HTTP request its certainly not the requesters fault.
42.
▲
by
dandermotj
10y ago
It is about generative models - models that let us draw samples from complicated distributions.
43.
▲
by
dandermotj
10y ago
I'm really looking forward to seeing the scientific community adopt docker as a way to distribute reproducible research and coursework. MIT 6.S094 has a Dockerfile[^1] that contains all the software required for taking part in the clas
44.
▲
by
dandermotj
10y ago
Use Docker on Ubuntu on my own machine, but forced to use a Docker on Windows at work. Using Docker on Linux is an absolute joy. I totally advocate it. Dockerhub is my first port of call when I'm installing something new on my machine
45.
▲
by
dandermotj
10y ago
The references made in the article are probably the best place to start.
46.
▲
by
dandermotj
10y ago
There's no such thing as reliably predicting a crash - that's why it's called a crash. There's bulls and bears but no one side is 'right' in aggregate.
47.
▲
by
dandermotj
10y ago
It's not a replacement for TeX. R originally had Sweave for weaving R and text together. Then we got Rmarkdown, which introduced the use of markdown for authoring documents via pandoc, but the output document only had simple features l
48.
▲
by
dandermotj
10y ago
If my understanding is correct, the perturbations are inherent in the model, not the data. It's a vulnerability in the high dimensional decision boundary of n nets.
49.
▲
by
dandermotj
10y ago
It depends on what the variance of the distribution is and how skeweded it is.
50.
▲
by
dandermotj
10y ago
For those who are interested in the actual implementation of these models check out simmer in R, SimPy in Python or SimJulia in Julia. Simmer was just released recently and has excellent examples/tutorials [0] for an introduction. [0]
51.
▲
by
dandermotj
10y ago
A contradiction - data science is the place where statistics and computer science meet but this book definitely isn't characterised as a book about computer science and graphs. From the TOCs, it is heavily based in statistics.
52.
▲
by
dandermotj
10y ago
Table of Contents: * High Dimensional Space * Best Fit Subspace & SVD * Random Walks & Markov Chains * Machine Learning * Massive Data: Streaming, Sketching, Sampling * Clustering * Topic Models, Hidden Markov Process, Graphical Mod
53.
▲
by
dandermotj
10y ago
Does anyone know if this is available as a hard copy?
54.
▲
by
dandermotj
10y ago
https://github.com/rho-devel/rho
55.
▲
by
dandermotj
10y ago
You're not wrong and absolutely totally wrong at the same time. R is the furthest thing from C you could find in paradigm, syntax and performance, but yes much of the underlying code is C or Fortran. But really you're missing the
56.
▲
by
dandermotj
10y ago
This is awesome - been waiting for it for ages! But bindings to the python API and not C++?
57.
▲
by
dandermotj
10y ago
I saw the vctrs package repo the other day on your GitHub. What's your plan with that? I guess your covering all of R's base data types (dplyr:data frames, purrr:lists, forcats:factors, vctrs:vectors)? Also do you plan on developi
58.
▲
by
dandermotj
10y ago
I read the opening two pages but I'm not sure what this really is. A higher math text book? It looks great though!...
59.
▲
by
dandermotj
10y ago
It looks exactly like a population heatmap.
60.
▲
by
dandermotj
10y ago
Or Rmarkdown
More ›