Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
yldedly
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
by
yldedly
4y ago
The notions that are crucial for distinguishing between intelligence and what large NNs are doing, are generalization and abstraction. I'm impressed with DALL-E's ability to connect words to images and exploit the compositionality
32.
▲
by
yldedly
4y ago
>It is not true that a piecewise-linear model trained on a set of data points will produce only outputs which are a convex combination of outputs that appear in the training set. No, and I didn't claim that. I said that, outside th
33.
▲
by
yldedly
5y ago
It's impossible for a piecewise linear function to be anything other than linear outside the training sample. They are by their definition unable to do anything but interpolate.
34.
▲
by
yldedly
5y ago
I'm not saying that's what they did. I'm saying it's functionally equivalent.
35.
▲
by
yldedly
5y ago
Take 100k joke and explanation tuples as training. Pattern match input joke to training jokes. Edit amalgamation of most structuraly similar training jokes to match content of input joke. Edit corresponding training explanation with same s
36.
▲
by
yldedly
5y ago
>I don't know why you think language models are fundamentally unable to deduce the knowledge of the points you mention. Because the knowledge is not there in the text, the models are not able to represent it, and as seen in the demo
37.
▲
by
yldedly
5y ago
>There is much structural regularity in a large text corpus that is descriptive of relationships in the world. Sure, there is a lot. But let's say we want to learn what apples are. So we look at occurrences of "apple" in t
38.
▲
by
yldedly
5y ago
Understanding of what? What the joke is about? Then no, it has no idea what any of it means. The syntactic structure of jokes? Sure. Feed it 10 thousand jokes that are based on a word found in two otherwise disjoint clusters (pod of whales,
39.
▲
by
yldedly
5y ago
It's a pretty big deal, and there's a big difference between a Markov chain and a deep language model - the Markov chain will quickly converge, while the language model can scale with the data. But the way these models are talked
40.
▲
by
yldedly
5y ago
>At any rate, I don't care if it only works one in ten times >you are asking me to disbelieve plain evidence
41.
▲
by
yldedly
5y ago
It doesn't. It's pattern matching, and you're seeing cherry picked examples. The pattern matching is enough to give the illusion of understanding. There's plenty of articles where more thorough testing reveals the differ
42.
▲
by
yldedly
5y ago
Sounds like no amount of math will convince you otherwise.
43.
▲
by
yldedly
5y ago
https://medium.com/analytics-vidhya/you-dont-understand-neur...
44.
▲
by
yldedly
5y ago
The point is that not only is it impossible to infer the structure of the world from text, deep learning is incapable of learning about or even representing the world. The reason language makes sense to us is that it triggers the right repr
45.
▲
by
yldedly
5y ago
No, no matter how many piecewise linear functions you compose, the result is still piecewise linear.
46.
▲
by
yldedly
5y ago
Really? Better inform all the researchers working on this that they're wasting their time then: https://arxiv.org/abs/2001.05016 More fundamentally, any finite neural net is either constant or linear outside the t
47.
▲
by
yldedly
5y ago
The structure of having X apples in Y buckets is the same as the structure in the expression "X * Y", as long as the expression exists in a context that can parse it using the rules of arithmetic, such as a human, or a calculator.
48.
▲
by
yldedly
5y ago
I got a lot out of this course https://people.csail.mit.edu/asolar/SynthesisCourse/ from one of the authors of OP.
49.
▲
by
yldedly
5y ago
I dabble in it. FlashFill in Excel uses program synthesis, but yeah, not many applications exist yet. Have you seen DreamCoder?
50.
▲
by
yldedly
5y ago
But you make the assumption that the data can be generated by your model, and your variance estimate only holds asymptotically.
51.
▲
by
yldedly
5y ago
Yes, but that's true of all statistics. You have to make some assumptions to get off the ground. If you estimate parameter variance the frequentist way, you also make assumptions about the parameter distribution.
52.
▲
by
yldedly
5y ago
Of course it does. You can put hyperpriors on the priors, and hyper hyperpriors on the hyperpriors, but the regress has to stop somewhere. What is your point?
53.
▲
by
yldedly
5y ago
To get the prediction variance in a Bayesian treatment, you integrate over the posterior of the parameters - surely computing or approximating the posterior counts as considering parameter variance?
54.
▲
by
yldedly
5y ago
From personal experience and what I've read, there's a clear dose response relationship. If you meditate enough, especially if do a lot of insight (vipassana) meditation without doing much concentration (samatha) or Metta meditati
55.
▲
by
yldedly
5y ago
Right, I see. That's not really possible imo. For things like mlops, sure. But model development, selection, evaluation? From what I've seen, it's exactly when engineers reach for standard tools without giving thought to how
56.
▲
by
yldedly
5y ago
In my experience, both are true. I'm more on the ML side, and I can tell I don't have the kind of routine and habits that good software engineers have, though I'm learning. But on the other hand, and I've seen this from
57.
▲
by
yldedly
5y ago
I started a PhD believing that scientists still held on to the habit of truth. The gradual realization that modern science rewards bullshit artists, and that great research happens in spite, not because of today's science culture, is o
58.
▲
by
yldedly
5y ago
>There's no guarantee that the phenomenon we're interested in will be expressible in an F=ma manner. Complexity exists. That's a fair point. But having a good theory doesn't necessarily mean that it's expressible
59.
▲
by
yldedly
5y ago
Karl Popper just turned in his grave.
60.
▲
by
yldedly
5y ago
Can't find a good general intro, but these articles take this viewpoint: https://science.mit.edu/life-away-from-equilibrium/ https://www.pnas.org/content/114/3/423 Edit: Here's
More ›