Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
simonster
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
by
simonster
4y ago
It seems inevitable that LLMs will eventually work for search. They just don't work yet.
32.
▲
by
simonster
4y ago
The bad versions should be developed, but they should not be marketed as a replacement for traditional education.
33.
▲
by
simonster
4y ago
In modern society, we prioritize the ease of bringing new products to market over regulation and testing. The only industry we prospectively regulate for safety is the pharmaceutical industry. Otherwise, regulation is retrospective and slow
34.
▲
by
simonster
4y ago
A brief look at the first few paragraphs of Vannevar Bush's Wikipedia article ( https://en.wikipedia.org/wiki/Vannevar_Bush ) would clearly establish that the answer to your question is yes.
35.
▲
Paul Graham predicts food shortages
(twitter.com)
25 points
by
simonster
5y ago
|
48 comments
36.
▲
by
simonster
7y ago
The DOT Full Fare Advertising Rule requires that airfare prices include all taxes and fees. Perhaps the cable industry needs a similar rule.
37.
▲
by
simonster
7y ago
The universal approximation theorem guarantees that a finite-width neural network that approximates the function to within some epsilon exists. But, regardless of the approximation method, there is no way to certify that a given approximati
38.
▲
by
simonster
7y ago
By this logic, (naive) matrix multiplication is not O(n^3) because it is a function of the precision required. The size of the neural network required to approximate a given function to within some epsilon does not change with the dataset s
39.
▲
by
simonster
7y ago
Yes, neural nets are successful is in large part because they are asymptotically more efficient than other models. Training time is O(n) with O(1) memory, and prediction time is O(1) with O(1) memory. Compare to e.g. kernel methods, which h
40.
▲
by
simonster
8y ago
Schmidhuber's take on the Wright brothers: https://www.nature.com/articles/421689c https://www.nature.com/articles/d41586-019-00491-5
41.
▲
by
simonster
8y ago
There's at least one existing paper about this idea ( https://arxiv.org/abs/1511.05641 ). Also, it is possible to initialize a convolutional layer so that it passes through its input, but initializing the weights pr
42.
▲
by
simonster
8y ago
Yes. They are a significant enough expense that they are a separate line item on SEC filings (search for "European Commission fines" in https://www.sec.gov/Archives/edgar/data/1652044/0001652044
43.
▲
by
simonster
8y ago
I don't think machine learning suffers from the same kind of p-value-driven replication crisis as other fields. It is true that people don't generally perform proper statistics to compare machine learning models [1], but ML resear
44.
▲
by
simonster
8y ago
#3 is actually wrong. The results of Recht et al. do not show that people are performing validation on the test set. If this were true, one would expect a poor correlation between accuracy on the original CIFAR-10 test set and the new test
45.
▲
by
simonster
8y ago
> What if no-one had a mind's eye and Aphantasia is simply the lack of a delusion of a mind's eye. The article proposes and falsifies a different hypothesis (that aphantasic subjects actually have a mind's eye but have the
46.
▲
by
simonster
8y ago
These sound like hypnagogic hallucinations. I think they're pretty common, although the precise experience varies by individual.
47.
▲
by
simonster
8y ago
It is possible that the problems are related—-it may be that, to achieve human-like generalization, neural nets need to learn in a human-like environment, instead of from a folder full of images. But time will tell.
48.
▲
by
simonster
8y ago
I don’t think this is a novel idea, but it is still a great topic for a PhD. While the results in this paper look impressive, my suspicion is that the system doesn’t generalize particularly well. (I suspect this from experience with similar
49.
▲
by
simonster
8y ago
Yep, the minima are in the same locations. However, if the problem has multiple minima, then the parametrization can affect which minimum gradient descent actually reaches.
50.
▲
by
simonster
8y ago
tl;dr ordinary gradient descent is sensitive to the parametrization of the problem but natural gradient is not. This is an important fact (and one that is fairly well-known within the ML community), but it is not totally clear to me why it
51.
▲
by
simonster
8y ago
It's a fair point that the Universal Approximation Theorem does not guarantee that the weights can be learned. OTOH, the physical laws that the article states a neural network cannot discover are computable functions.
52.
▲
by
simonster
8y ago
There are a couple of factual errors here. First, the difference between backprop and evolution is smaller than the author indicates. The error signal used in modern backprop training is stochastic because it is computed on a minibatch (whi
53.
▲
by
simonster
8y ago
Given that TSMC just started high volume production of 7nm chips ( https://www.anandtech.com/show/12677/tsmc-kicks-off-volume-p... ), this sounds bad for Intel.
54.
▲
by
simonster
9y ago
It's not totally clear to me whether there is a real story in ML/AI or neuroscience. The authors used logistic regression to try to determine whether a subject will remember a word or not, which the classifier did better than chan
55.
▲
by
simonster
9y ago
I'm somewhat amused that, although this article suggests that Buso's work was vital to the Nature paper, as does the first sentence of the paper itself, he is author 7 out of 21: https://www.nature.com/articles
56.
▲
by
simonster
9y ago
Let's assume that evaluating a program takes one microsecond and one atom, and that we can parallelize the search across every atom in the observable universe. If the AGI program is 500 bits long, it will take about 10^57 years to find
57.
▲
by
simonster
9y ago
It's a more accurate neuron model of some pretty weird neurons. In nearly all organisms, the vast majority of neurons fire all-or-nothing action potentials ("spikes"). C. elegans neurons do not.
58.
▲
by
simonster
9y ago
Beyond the initial stages of the network, current SOTA CNNs use strided convolution in addition to (Inception, NASNet) or instead of (ResNet, DenseNet) max pooling. But my impression is that this has more to do with computational efficiency
59.
▲
by
simonster
9y ago
As a neuroscientist who just started doing ML research, I would call this paper cargo cult programming. If you cobble together a hodgepodge of ideas from neuroscience and build a network to accomplish some trivial task with no baseline to c
60.
▲
by
simonster
9y ago
If you spent 7 years of your life on a project only to haven it stolen out from underneath you by an unscrupulous "mentor", you're saying wouldn't get emotional about it? As far as I can tell, her story reveals excepti
More ›