Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
christopheraden
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
61.
▲
by
christopheraden
13y ago
Btilly outlined the steps of how you'd go about showing mathematically that a binomial converges to the normal distribution, but the reason you don't see this approach, even after two courses in statistics, is that it takes awhile for peopl
62.
▲
by
christopheraden
13y ago
I've surfed around on this article several times while trying to find a distribution I needed in order to make my rejection samplers work better for distributions that have no name or are obscure (posteriors in Bayesian inference often gene
63.
▲
by
christopheraden
13y ago
I'm worried your explanation may be confusing to people in some places. The bell curve and Gaussian curve (and normal curve) are almost always synonymous. Bell Curve describes its shape, while the Gaussian curve describes the guy who derive
64.
▲
by
christopheraden
13y ago
Consider that the source is High Scalability, a blog that focuses on how orthodox methods are inadequate for 1-10M concurrent connections and how different methods are being employed to reach these lofty goals, I think the advice is pretty
65.
▲
by
christopheraden
13y ago
I like that you ask for mindfck movies. They definitely hold a special place in my heart. A lot of the entries here have been very sci-fi. I tend to like more historical or present-day dramas with great acting that expose some sort of undes
66.
▲
by
christopheraden
13y ago
While I might generally agree that the interview process is asinine and a rat race, I actually use some of those formal methods and algorithms in my day-to-day work--and I'm not even a computer scientist, nor do I have a CS degree. I'll bet
67.
▲
by
christopheraden
13y ago
From perusing the main page, this looks exactly like what I want. Thanks for the awesome suggestion.
68.
▲
Ask HN: What are good resources for modern web development?
2 points
by
christopheraden
13y ago
|
2 comments
69.
▲
by
christopheraden
13y ago
You will find the intro books don't talk much about parallel computing. Most of the general data sets in intro books will be no more than 30 observations. They are trying to teach classical methods moreso than useful computational technique
70.
▲
by
christopheraden
13y ago
Yankoff, you might want to be more specific. Intro statistics in general, or for computer scientists, or scientists, or looking to learn R at the same time? I liked Freedman, Pisani, and Purves [1], and have TA'ed using McClave, Sincich, an
71.
▲
by
christopheraden
13y ago
Seems more like a brain fart than an error in understanding, especially because he didn't highlight why it was important to have unbiasedness in the first place. Having to explain the unbiased estimator for standard deviation would probably
72.
▲
by
christopheraden
13y ago
Hey Evan, from one statistics guy to another, thanks for fighting the good fight :). The formulas might benefit from examples, especially with some of the more complicated cases (KS test and onwards). The important part of statistics comes
73.
▲
by
christopheraden
13y ago
On #7: ACID is a concept that matters just as much now as it used to, but I think a lot of people that thought ACID was important realize that it might not be so. I wouldn't say a RDBMS is uniformly better than text files, but in some appli
74.
▲
by
christopheraden
13y ago
Where's openbox? #! users ought to be outraged!
75.
▲
by
christopheraden
13y ago
I work in corporate where there is no admin access to my work machine and we're expected to use Notepad for small tasks and Word for big ones. The windows version of Gvim gives me all the familiar keybindings I have on my Linux box at home,
76.
▲
by
christopheraden
13y ago
Power analysis and CI's should be elementary, but I would assert that they are actually not commonplace. Most people have a very surface-level understanding of the latter, and little understanding of the former. In my opinion, A/B Testing h
77.
▲
by
christopheraden
13y ago
"Use a dedicated statistical package from the '80s" Is your Wizard app not also a dedicated statistical package? Also, I'm being pedantic here, but how many "dedicated statistical packages" are actually from the 80s? The only ones that come
78.
▲
by
christopheraden
13y ago
Like all of A/B testing, it's applying a _very_ old statistical method (Chi-Square was one of the first modern statistical techniques--by that I mean it's 113 years old) to an area where statistics has not commonly been used. This makes it
79.
▲
by
christopheraden
13y ago
I meant checking assumptions not just to see whether the use of the big data moved a business metric, but also that the model makes sense from a statistical perspective. A lot of statistics in business does not bother to check modeling assu
80.
▲
by
christopheraden
13y ago
Statistics involves checking modeling assumptions. A lot of what I've seen with the big data people is the repetition of algorithms to the exclusion of understanding and checking modeling assumptions. While it's nice that the big data craze
81.
▲
by
christopheraden
13y ago
The fact that so many people are calling things "big data" when the data is not high volume (the most popular definition I've seen is the 5 V's definition--big seems to be a misnomer in this case, as only volume could really be called a mea
82.
▲
by
christopheraden
13y ago
I am grateful to finally see this in an article. The "big data" craze is being pushed in areas where it really doesn't make sense. We've been bit by the Big Data bug where I'm at, but it's not coming from the statisticians. It's usually the
83.
▲
by
christopheraden
13y ago
I met Chad Whitacre at the Gittip booth at PyCon (the heart-punching coin machine was pretty cool!). I found him to be very earnest and interested in the perception of his company. I'm glad he felt strongly enough about the openness to refu
84.
▲
by
christopheraden
13y ago
Disregarding the morality and ethics that are often center-stage, they are faced with really interesting technical hurdles. It's an old article, but I was fascinated to hear about some of the challenges they face as well: http://highscalab
85.
▲
by
christopheraden
13y ago
Most of the work I do involves recent data--cycles are six months at most and 3 months on average. 3 months of data, sifting by a pretty strict filter, it's not unsurprising that hundreds of terabytes of claims over years and years gets fil
86.
▲
by
christopheraden
13y ago
The types of queries we run against it don't require real-time results, and we do a pretty heavy amount of subsetting. By the time it reaches the point where we do numerical summaries and statistics, the largest set I've worked with here wa
87.
▲
by
christopheraden
13y ago
I've had a similar question before. I have heard of as few as 60,000 observations was considered "big data" [1], yet at my company, we generate about 60 million pharmacy claims every 3 months, and no one here calls it big data. In terms of
88.
▲
by
christopheraden
13y ago
Thanks for pointing out the distinction. You're right that this is what I was going for, but the articles I've seen thus far about the inaccuracies in Excel deal mainly with the poor capabilities of the optimization routines, RNG, and distr
89.
▲
by
christopheraden
13y ago
We got around the inaccuracy problems where I'm at by doing all the sensitive calculations in a more capable program (SAS where I'm at now--R when I consult), then exporting the results to Excel. It avoids most of the problems faced by the
90.
▲
by
christopheraden
13y ago
They are just now moving people off MSO 2003 at ours as well, following complaints from several people that they can't open the newer xlsx, docx, and accdb file extensions. IT and Clinical Analytics are the first wave to move to "brand new"
More ›