Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
datastoat
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
by
datastoat
5y ago
I think the article is confusing three things. In Bayesian terminology, suppose we've found a posterior distribution for effect size. Then there are three things we might consider: (a) how tightly concentrated the posterior distributio
32.
▲
by
datastoat
5y ago
Ideally, there's some field in our data that we can use -- e.g. train on data from cities in several countries, test on data from cities in a different country; train on one corpus, test on another; train on climate data at one level o
33.
▲
by
datastoat
5y ago
Big enough data that you can afford not to use some of it for training! Different disciplines hit this threshold at different times -- language and speech much earlier, as you say; clinical trials not there yet. Maybe we could talk about tw
34.
▲
by
datastoat
5y ago
Validation on holdout sets. When I was a student in the 1990s, I was taught about hypothesis testing (and all the hassle of p-fishing etc.), and about Bayesian inference (which is lovely, until you have to invent priors over the model space
35.
▲
by
datastoat
5y ago
When I scrape, I attach a header that says roughly “By responding to this request, the provider allows me to use the response for my own personal use; and accepts that this overrides any terms stated on the web page.” Cut and dried? A lawye
36.
▲
by
datastoat
5y ago
Relatedly, a lot of recipe books just have a dry statement of what the ingredients are and how to combine them. They should be more like CS and science papers, and explain the narrative of why the recipe is exciting, and where it comes from
37.
▲
by
datastoat
5y ago
> It is the dataset used for training that is often at fault, not the source code of the model This remark is absolutely spot on. This is why I don't like the language around 'algorithmic bias' and 'algorithmic accoun
38.
▲
by
datastoat
5y ago
I've not heard of nonlinear causal reasoning. It sounds interesting. Could you suggest some reading? (A quick search throws up uninteresting hits about nonlinear models.)
39.
▲
by
datastoat
5y ago
The article seemed to me to go back and forth between profound and tendentious. To pin down the ideas precisely, or perhaps to trivialize them, here's a simple technical + mathematical example: TCP, the congestion control protocol for
40.
▲
by
datastoat
5y ago
What are the facts that changed, in this lab leak story? As far as I'm aware, our knowledge of the real-world facts haven't changed -- we're still in a position of great uncertainty. All that's changed is how various com
41.
▲
by
datastoat
5y ago
How about making the tax increase each month the unit is unoccupied, but also making it proportional to the asking rent? No one has any use for it, high asking rent => high tax. No one has any use for it, very low asking rent => low t
42.
▲
by
datastoat
5y ago
Here's an example: https://youtu.be/eHwy-neG_W8?t=791 It's from an introductory course on Algorithms; the dataset is the students' programming assignments, rated for similarity by an off-the-shelf similarity
43.
▲
by
datastoat
5y ago
> none of which seem more intuitive than writing "mu" Write for the convenience of the reader, not of the writer!
44.
▲
by
datastoat
5y ago
For me: alt+enter (switches to Greek keyboard layout), m, alt+enter (back to usual layout). I use Greek symbols in Python and in Javascript. I do it almost without thinking now, as lots of my coding is tightly linked to maths.
45.
▲
by
datastoat
5y ago
That forum post advises "new drugs aimed at preventing Alzheimer’s should probably target surrogate markers rather than trying to fix the end-stage clinical problems". Here's One Weird Trick for targeting Alzheimer's, to
46.
▲
by
datastoat
5y ago
I completely agree about the steep learning curve and the feeling of dark magic -- how many times have I had to relearn what deparse(substitute(x)) means -- but oh the satisfaction of broadening my programming horizons. For me it didn'
47.
▲
by
datastoat
5y ago
R certainly expanded my programming views! Haskell did too, but the lessons of Haskell didn't stick the way that R's lessons did. Here are some of the things I learnt from R (though they can be found in other languages of course).
48.
▲
by
datastoat
5y ago
The article explains why it's simplistic to think of legal argument as pure maths-style logical deduction. This goes back at least to Oliver Wendell Holmes (US supreme court justice from 1902), who made fun of those who treated a disse
49.
▲
by
datastoat
5y ago
The probability that at least one tracker is compromised actually grows sub-linearly in the number of trackers. Let n be the number of trackers, let p be the probability that an individual tracker is compromised, and assume compromises are
50.
▲
by
datastoat
6y ago
> But not Greek plurals, so it's octopuses, not octopodes. Though Australia is in the antipodes (to the UK), not the antipuses.
51.
▲
by
datastoat
6y ago
To be precise, it's impossible to measure one-way latency between a pair of nodes without external information. But if you have a mesh of nodes, the story is different [1]. If you have N nodes and hence N unknown clock offsets, and i
52.
▲
by
datastoat
6y ago
As the article explains, latency and clock sync go hand in hand. Here's a blog post [1] that goes further into clock sync, contrasting NTP and hardware-based systems. The company behind the blog post says that their solution is availab
53.
▲
by
datastoat
6y ago
While we're waiting for the paper on the Platonic link ... here's a book about how machine learning is all just a rerun of Oliver Wendell Holmes's theory of epistemology in the law, from the 1890s. https://link.spr
54.
▲
by
datastoat
6y ago
The article starts with "if it's not reproducible, it's not science, right?" And then it goes on to talk about a limited idea of reproducibility, i.e. plain duplication and recipe-style rerunning an analysis. I think the
55.
▲
by
datastoat
6y ago
Probability notation used in ML and engineering has this problem, of overloading p(). Probability notation as used by probabilists in maths departments is completely different: it’s more explicit, and sometimes more clunky. There’s a hybrid
56.
▲
by
datastoat
6y ago
He bricked the marble!
57.
▲
by
datastoat
6y ago
The outputs of the model _were_ being treated as predictions. The Ferguson paper from 16 March used the language of prediction: "In the (unlikely) absence of any control measures [...] given an estimated R0 of 2.4, we predict 81% of t
58.
▲
by
datastoat
6y ago
I would pick a value of R that shows itself to have good predictive accuracy. The way to test predictive models is always to look for their predictive accuracy on holdout data. Machine learning has this ingrained. Classic statistics does th
59.
▲
by
datastoat
7y ago
It's good you put "contained" in quotes. Japan, Singapore, and Hong Kong are all (according to this plot [0]) showing exponential growth in the number of cases. It's a slower exponential than elsewhere, but it's sti
60.
▲
by
datastoat
7y ago
If Facebook lawyers understood the implications of LeCun's argument, they wouldn't be happy! There are two types of explanations here: (1) why did the data come to be as it is, (2) why did my ML make the prediction it did. Science
More ›