Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
fnbr
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
fnbr
3y ago
Ah, well you could use a standard value network, but it’d end really slow, so you probably want to train a smaller one and rely on the implicit ensembling that MCTS does to make it better. In my experience, PUCT does a lot better than UCT,
32.
▲
by
fnbr
3y ago
It’s very difficult to implement, and requires training the network to use it. I worked at DeepMind on projects that used MCTS. Even with access to the AlphaZero source code, it was very difficult to write an other implementation that got t
33.
▲
by
fnbr
3y ago
You need performance in the high level libraries to match, on a flops/$ basis. That’s “it”. That’s easier said than done, though. Even google’s TPUs still struggle to match H100s at flops/$, and they’re really annoying to use unle
34.
▲
by
fnbr
3y ago
I see the opposite fwiw. Numbers significantly lower
35.
▲
by
fnbr
3y ago
These numbers are wildly inaccurate. I spotchecked multiple places that I know the salaries and they're wrong.
36.
▲
by
fnbr
3y ago
Less expensive. Big discounts exist.
37.
▲
by
fnbr
3y ago
A number of large AI companies use it to train their large models; Midjourney, Stability, Anthropic, DeepMind, among others.
38.
▲
by
fnbr
3y ago
The main benefit in my experience is that it’s much easier to do distributed computations in JAX. It has a much nicer API. For single device computing there’s no advantage either way.
39.
▲
by
fnbr
3y ago
Not sure about VTI, but if you invested in the S&P 500 and reinvested dividends, you would have got 130%: https://dqydj.com/sp-500-return-calculator/
40.
▲
by
fnbr
3y ago
You’re welcome :)
41.
▲
by
fnbr
3y ago
It’s definitely a very cool moment for me :)
42.
▲
by
fnbr
3y ago
Yeah, you’re totally right. I actually wrote a follow up essay about that: https://finbarr.ca/llms-not-trained-enough/ I think the conversations were partly (largely?) a snapshot in time. I was talking to people in Feb
43.
▲
by
fnbr
3y ago
This is the original page, I’d link to this: http://www.incompleteideas.net/IncIdeas/BitterLesson.html u/dang, swap links if you see this?
44.
▲
by
fnbr
3y ago
I went from being a lifelong fan, especially of the EU, to not caring at all. I’m now a Trekkie.
45.
▲
by
fnbr
3y ago
I was about to ask the same thing. Why does this matter? It's like when the Canadian Parliament subpoenaed Mark Zuckerberg- he just ignored it [1]. [1]: https://www.cnn.com/2019/05/27/tech/zuckerberg
46.
▲
by
fnbr
3y ago
man, I want a tiny Japanese pickup truck
47.
▲
by
fnbr
3y ago
Yeah, but Ilya left. Doesn't that prove your point?
48.
▲
by
fnbr
3y ago
Sous vide translates literally to “under vacuum” as it involves cooking food in plastic bags that are vacuum sealed and cooked in water baths at precise temperatures. The technique, however, is much more useful for the temperature control e
49.
▲
by
fnbr
3y ago
I think this is an excellent resource. I used to work at DeepMind, and Hado & Matteo are two of the best RL researchers there.
50.
▲
by
fnbr
3y ago
I was at google when he led the swift for tensorflow team and it was really unimpressive. Hopefully this does better
51.
▲
by
fnbr
3y ago
I’m doing my part. I got a new work M1 MBP. Then, I was laid off, so I got a new personal M2 air. Then, I got a new job, which gave me a new M2 MBP. North of $10k in laptops. Oh, and a new iPhone.
52.
▲
by
fnbr
4y ago
Biggest advantage is the high bandwidth interconnect. It’s very painful training large models on multiple GPUs. TPUs help a lot with that, although the software side of things is pretty brutal.
53.
▲
by
fnbr
4y ago
Yes! This is something that is done. The problem is that a) it’s tough to find a sane denominator as the likelihood of the entire sequence can be quite small, even though it’s the best answer and b) the answer isn’t grounded in anything, so
54.
▲
by
fnbr
4y ago
It is not common for new projects. The vast majority of new projects use PyTorch, with some using tensorflow and some using JAX.
55.
▲
by
fnbr
4y ago
Do the baskets really make that much of a difference compared to stock ones?
56.
▲
by
fnbr
4y ago
I love this (even though I have ~25% of my net worth in $GOOG). Google got fat and lazy. They focused inwards, on promo driven development and executive infighting. This was awful for the tech ecosystem as Google was held out for a long tim
57.
▲
by
fnbr
4y ago
When I was laid off, and all my access to internal systems was disabled, I had no way to get in touch with my former colleagues, who didn’t have my contact info. Having a personal website enabled them to get in touch with me, which led to a
58.
▲
by
fnbr
4y ago
I’m a naive techie. What, exactly, did Mitchell do wrong here? What could he have done better?
59.
▲
by
fnbr
4y ago
FAANGs aren’t anywhere near the high end of the ML job market currently.
60.
▲
by
fnbr
4y ago
ML engineers work on products, AI research engineers work for AI research labs. Skillset is almost identical, biggest difference is having research experience.
More ›