Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
maxwells-daemon
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
maxwells-daemon
5y ago
Also went to Caltech, and my progression was the complete opposite. Started out going to every lecture and trying to write everything down, but things just moved way too fast -- I'd write down every word and figure without ever interna
32.
▲
by
maxwells-daemon
5y ago
Fond memories of running into my math professor walking around campus at 3am...
33.
▲
by
maxwells-daemon
5y ago
Look at the "math test" video. Given the question: "Jane has 9 balloons. 6 are green and the rest are blue. How many balloons are blue?" The model outputs: "jane_balloons = 9; green_balloons = 6; blue_balloons = jan
34.
▲
by
maxwells-daemon
5y ago
Autoregressive transformers take a while to generate text, since you need to run the whole model once for every word in the output.
35.
▲
by
maxwells-daemon
5y ago
As an add-on to this: I'd encourage anyone interested in this debate to read Rich Sutton's "The Bitter Lesson" ( http://www.incompleteideas.net/IncIdeas/BitterLesson.html ). At every point in time, th
36.
▲
by
maxwells-daemon
5y ago
I would argue humans ingest a lot more than 159GB before they can write code. Most of it isn't Python, and humans currently transfer knowledge a lot more efficiently than NNs, but I suspect that'll change as incorporating more var
37.
▲
by
maxwells-daemon
5y ago
The "language models don't really understand anything" corner is getting smaller and smaller. In the last few months we've seen pretty definitive evidence that transformers can recombine concepts ([1], [2]) and do simple
38.
▲
by
maxwells-daemon
5y ago
I'd rate myself as "above-average receptive" to ML-based tooling, but after trying two "AI autocomplete" tools (Kite and TabNine) I've decided it's not for me. The suggestions were usually good, but I foun
39.
▲
by
maxwells-daemon
5y ago
Most of the responses here seem to imply that the author doesn't understand that physics can be complicated (in the sense of being hard to learn or having big equations). He studies theoretical physics at MIT [1], so I expect he does.
40.
▲
by
maxwells-daemon
5y ago
I think it depends! If you want to zoom out and take the "systems view" using standard components, then you probably don't need much math. If you want to develop new architectures or algorithms, then you definitely will. The
41.
▲
by
maxwells-daemon
5y ago
I often want to run small "experiment code" many times with different inputs, but don't like wrapping everything in a big framework to do it. So I wrote a little tool that calls any command-line program multiple times with di
42.
▲
by
maxwells-daemon
5y ago
Not Fast.ai, but I self-studied ML during undergrad (mostly from books) and am currently working as an ML research scientist. That being said, I'm also thinking about starting an ML PhD because it does honestly open more doors to top r
43.
▲
by
maxwells-daemon
5y ago
It's interesting that the "probability to distance" function (-log(prob(edge))) wound up being the information-theoretic entropy [1]. I wonder if there's anything deeper there? [1]: https://en.wikipedia.org&#x
44.
▲
by
maxwells-daemon
6y ago
My research (automated theorem proving with RL) sits partway between "good old-fashioned AI" and modern deep learning, and GEB struck me as amazingly prescient, with lots of lessons for modern AI research. There's a growing s
45.
▲
by
maxwells-daemon
6y ago
A couple on and off, but most recently a GPU-accelerated differentiable fluid simulator: https://github.com/maxwells-daemons/deltaflow
46.
▲
by
maxwells-daemon
6y ago
This seems true about smallish groups, like the example of a Jew in czarist Russia, but when the group is big enough to divide into multiple public-opinion subgroups, I'm not so sure. I think the feminism argument doesn't work her
47.
▲
by
maxwells-daemon
6y ago
I'm pretty worried about designing new datasets specifically around observing things we've already seen in real data. Sure, we can observe spatial priors and double descent with this dataset by design, but when we see a new intere
48.
▲
by
maxwells-daemon
6y ago
I graduated from Caltech last year, and the exams are all still take-home. You're expected to time yourself and not consult course materials, and basically nobody cheats. It's a rare sort of equilibrium where nobody's cheatin
49.
▲
by
maxwells-daemon
6y ago
One cool thing about Lean is that the next version is trying to become something like a general-purpose programming language, with the idea of "using proofs to write more efficient code." For example, you might prove that your cod
50.
▲
by
maxwells-daemon
6y ago
I don't think we can say for sure that early stopping is the main reason deep networks generalize. Double descent [1] shows that models continue to improve even once they've "interpolated" the training data (fit every po
51.
▲
Show HN: Argsearch – A composable tool for sweeping over command-line arguments
(github.com)
6 points
by
maxwells-daemon
6y ago
|
0 comments