Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
louden
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
louden
11y ago
This is an interesting read of ML in Rust, but it jumps straight into the modeling part. One question I have is how does one get a good idea of what the data is (e.g. data visualization and simple summary statistics) in Rust? Or is this s
32.
▲
by
louden
11y ago
R is a language with a lot of gotcha's. I usually get burned by characters being converted to factors in read.csv() and converting factors to numeric (it works, but not how you intend). The R Inferno ( http://www.burns-stat
33.
▲
by
louden
11y ago
I am a data scientist. I am working on two side projects. The first project is a cloud based power/sample size calculator as a competitor to PASS. The second project is a statistical distribution visualization website with API to gene
34.
▲
by
louden
11y ago
I use a 17" laptop personally because I love the number pad on the keyboard and the bigger screen (and I do take it around with me). I would get the more powerful laptop if possible due to the text analysis and neural network stuff you
35.
▲
by
louden
11y ago
Within the same repo, have two branches: database and Database and alternate between them.
36.
▲
by
louden
11y ago
I'm working on a browser based set of study design calculators (samples size and randomization lists). I'm using flask and scipy. It is a lot of fun and keeps my theory sharp (I'm a statistician by trade).
37.
▲
by
louden
11y ago
Tree based methods, such as Random Forests, provide some nice properties at the expense of interpretability. Trees by themselves tend to be fairly variable.
38.
▲
by
louden
11y ago
Not everybody needs to know how to do math beyond arithmetic (though the opportunity should be provided). The focus should be making sure everybody can understand the math, and especially statistics, they are presented with everyday.
39.
▲
by
louden
11y ago
Even with those other methods in place, the leaderboard would still favor methods that overfit. That is why the final score is determined on a dataset (the validation set) that is not used until the model is locked down. The public leader
40.
▲
by
louden
11y ago
To be fair to Stripe, this will be an issue regardless of the provider you use. They share the same back-end partners. As long as a business appears to be high risk, the risk of shut down by payment provider will exist.
41.
▲
by
louden
11y ago
This has been my experience as well. To combat this, I always start with a written analysis plan that forces me to think out all of my analyses.
42.
▲
by
louden
11y ago
If you have time, in the long term I would recommend looking at the theory behind the methods. It will give you a lot of insight on why a person is using a particular method and when it is inappropriate to use a certain method.
43.
▲
by
louden
11y ago
I second using kaggle. They have a lot of 'competitions' that are really just learning exercises with a lot of information about the techniques used directly linked to the competition.
44.
▲
by
louden
11y ago
This article illustrates the problem with over-fitting a model even when some data is withheld for testing. This is a trap that one can fall into when using training and testing sets.
45.
▲
by
louden
12y ago
It's a trade off. Clopper-Pearson can be overly conservative in many instances. I tend to use Jeffreys interval personally, which is a Bayesian method. Brown et al. give a formal treatment of the subject with simulations showing actu
46.
▲
by
louden
12y ago
R does support functional programming. In fact, it was based on Scheme (see discussion for comparison: https://stat.ethz.ch/pipermail/r-help/2008-December/181982.h... ).