Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
tansey
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
by
tansey
11y ago
Thanks for the write-up Chris. Now I understand why I couldn't follow the path of logic you were laying out in our original discussion in PG's article's comments. The main problem I was having is that you are assuming our obs
32.
▲
by
tansey
11y ago
Okay, I'll have to wait for your full write-up, because I am not seeing the path of thought here.
33.
▲
by
tansey
11y ago
Sorry, I'm confused here. A p-value makes an implicit assumption that your null hypothesis is a known N(0,1). That may be throwing me off a bit. I get the point of you want to look at the likelihood function which is just one minus the
34.
▲
by
tansey
11y ago
Does that assume both samples are identically distributed and the only difference is the cutoff? If it does, then couldn't we just continue to do a difference of means test and still be consistent? If it doesn't, how do you handle
35.
▲
by
tansey
11y ago
So the PG estimator is clearly problematic. I agree that the yummfajitas (YM) estimator looks to be consistent. In this case though, we're dealing with (small) finite sample sizes, so we need to come up with some sort of test statistic
36.
▲
by
tansey
11y ago
So true. PG's articles are generally filled with good intuitive insight. Unfortunately, statistics can be very tricky to turn into folksy wisdom. Rules of thumb like "you need 30 samples before you can say anything" that are
37.
▲
by
tansey
11y ago
> Other than being more complex, the biggest downside is that all of these methods have some new parameter(s) to tune. They're so fast to run though, that just doing warm-starts and a huge solution path (or grid in the case of addit
38.
▲
by
tansey
11y ago
Strangely enough, the author only mentions total variation denoising in passing as a feature of SmartBlur. I would say this method is one of the most common, especially when your image has sharp transitions and lots of solid regions of colo
39.
▲
by
tansey
11y ago
Actual paper: http://www.pnas.org/content/112/13/E1569.full The idea (from 5 minutes of skimming plus watching the little 3min tutorial) seems to be that any multi-dimensional time series implicitly contains
40.
▲
by
tansey
11y ago
Perhaps more precisely, they're "statistical engineering" jobs. A machine learning PhD can derive an algorithm and provide you with a reassuring bound or guarantee regarding performance in terms of runtime, convergence, etc.
41.
▲
by
tansey
11y ago
The AIMA book is sort of a Good Old-Fashioned AI (GOFAI) book that focuses a lot on agents and planning. The jobs this article is talking about are really machine learning ones-- taking large volumes of data and extracting knowledge, so as
42.
▲
by
tansey
12y ago
Is the raw data available anywhere?
43.
▲
by
tansey
12y ago
> 85% of the US population has internet access Right. We're talking about the people more likely to be in the 15%. >...and (as Sam Altman later says[1]) libraries are available for the less fortunate. Compare that to biology.
44.
▲
by
tansey
12y ago
Yep, that's a more precise version of what I was saying w.r.t. estimating sample size. I think the tool makes some assumption about variance, but the other 3 are things you supply. Note that I wasn't saying anything about the A&#x
45.
▲
by
tansey
12y ago
> Shouldn't it be: "Your experiment requires 5,000 visitors and after that we'll check if the result was significant enough to not be merely due to random chance"? That's basically what is happening with the to
46.
▲
by
tansey
12y ago
I had the same thought. The bootstrap is a really simple and easy-to-implement technique that we've had for decades. There is even some recent work on a "Big Data" (distributed) version of the bootstrap from Michael Jordan&#x
47.
▲
by
tansey
12y ago
Indeed! I have a particularly relevant horror story. For one of my graduate classes, I built a game where two people (a liar and a truth teller) would try to convince a third person (the judge) they had experienced something. The goal was t
48.
▲
by
tansey
12y ago
The gensim [0] package has a nice implementation of online LDA that can handle massive streaming datasets. If you want to avoid specifying the number of topics, you can use HDP-LDA. David Blei (inventor of LDA) has a reference implementatio
49.
▲
by
tansey
12y ago
Looks like they're moving around a 3D cube on a tablet screen.
50.
▲
by
tansey
13y ago
Except the author adjusted the timelines to match them up. The skirt widths are shifted by 21 years.
51.
▲
by
tansey
13y ago
Good luck convincing the rest of the country to give your region 10 new senators.
52.
▲
by
tansey
13y ago
Same here. I was thinking after my edit that I don't know of such an algorithm either, but I don't think in principle that there is anything stopping someone from coming up with one. So... who wants a NIPS paper? Noel and I are ha
53.
▲
by
tansey
13y ago
I'm not sure why the author picks Myth #5 as the big pro-frequentist point. The "online learning" problem he's describing is a simple multi-armed bandit problem. The magical frequentist algorithm that he's touting i
54.
▲
by
tansey
13y ago
But this data is not de-identified. They explicitly list prices for data with confidential, patient-identifying information.
55.
▲
by
tansey
13y ago
Sorry, it's for UT Austin PhD students only.
56.
▲
by
tansey
13y ago
I actually read this paper a couple weeks ago as part of a deep learning reading group that I co-run. While several of these authors are household names in the RL community, this paper was not actually that impressive to me. The only real &
57.
▲
by
tansey
13y ago
Should that matter? Doctors have a hippocratic oath to uphold.
58.
▲
by
tansey
13y ago
Note that this show is terrible. They are trying to make Silicon Valley and startups look like Hollywood and scriptwriting. The first episode has the main characters (startup co-founders) hanging out at a trendy bar and programming there. O
59.
▲
by
tansey
13y ago
> ...for those kids it can be a living hell... I think you need to check your privilege here. You grew up in a household with successful parents that wanted their children to succeed. If you didn't get into Berkeley, you'd ha
60.
▲
by
tansey
13y ago
If you really just want one email for everything, go with team@mydomain.com. All the other choices imply a purpose for the email (e.g., first contact or help with a problem).
More ›