Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
WanderPanda
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
WanderPanda
2y ago
Yes, it seems like capability is logarithmic wrt compute but utility (in different applications) is exponential (or rather s-shaped) with capability again
32.
▲
by
WanderPanda
2y ago
It needs to be intact, the pdf is rendered by the arxiv backend based on the source
33.
▲
CUDA 12.8 released; Introducing Blackwell support
(docs.nvidia.com)
1 points
by
WanderPanda
2y ago
|
0 comments
34.
▲
by
WanderPanda
2y ago
I mean it is humanity’s LAST exam. Humanity’s first exam would probably be something about communication? Or about building and predicting effects of certain tools?
35.
▲
by
WanderPanda
2y ago
It’s sad that we have to assume this is the agenda for all offerings that don’t allow paying current price x years in advance (possibly including adjustion for expected inflation). Leaves a very sour taste, especially in the case of hardwar
36.
▲
by
WanderPanda
2y ago
Or is it just white collar workers experiencing what blue collar workers have been experiencing for decades?
37.
▲
by
WanderPanda
2y ago
Crazy conglomerate discount on Alphabet if you can see TPUs as the only Nvidia competitor for training. Breaking up Alphabet seems more profitable than ever
38.
▲
by
WanderPanda
2y ago
No you don‘t need much bandwidth between cards for inference
39.
▲
by
WanderPanda
2y ago
Damn good to know! Have been gaslit by the ever-changing docker install instructions. Of course it would be a lagging version, but I think the docker feature-set has converged years ago, why would I care any more about the docker version th
40.
▲
by
WanderPanda
2y ago
Makes me wonder why docker still didn't make it to the ubuntu/debian repositories. Would be such an easy net benefit
41.
▲
RLtools: The Fastest Deep Reinforcement Learning Library
(github.com)
1 points
by
WanderPanda
2y ago
|
0 comments
42.
▲
How is this Website so fast? (McMaster-Carr) [video]
(youtube.com)
3 points
by
WanderPanda
2y ago
|
1 comments
43.
▲
by
WanderPanda
2y ago
This sounds like the (debunked) labor theory of value :p
44.
▲
by
WanderPanda
2y ago
I wonder why that is? because they are trained with dropout?
45.
▲
by
WanderPanda
2y ago
+ being early on capacitive touchscreens and multi-touch
46.
▲
by
WanderPanda
2y ago
You can just alt + click the audio options in the control center to select the input source
47.
▲
by
WanderPanda
2y ago
JAX still has the "This is a research project, not an official Google product. Expect bugs and sharp edges. Please help by trying it out, reporting bugs, and letting us know what you think!" disclaimer in its readme. This is quite
48.
▲
by
WanderPanda
2y ago
How will we even measure this? Benchmarks are gamed/trained on and there is no way that there is much signal in the chatbot arena for these types of queries? I think in just a few month the average user will not be able to tell the dif
49.
▲
by
WanderPanda
2y ago
I'd be shocked if we don't see diminishing returns in the inference compute scaling laws. We already didn't deserve how clean and predictive the pre-training scaling laws were, no way the universe grants us another boon of th
50.
▲
by
WanderPanda
2y ago
If the writing & arts vs. doing laundry & cleaning dishes is any indication, it does not look rosy. All the fun and rewarding parts (low hanging fruits / quick wins) of coding might be automated. What remains are probably thing
51.
▲
by
WanderPanda
2y ago
The St. Petersburg paradox is where hypers and doomers meet apparently. Pricing the future infinitely good and infinitely bad to come to the wildest conclusions
52.
▲
by
WanderPanda
2y ago
Doesn‘t even have to be about money. Would be perfectly rational to comply to minimize uncertainty, public outcry, getting into the cross-hair of other countries‘ regulators and distractions from operating the core product/service. In
53.
▲
by
WanderPanda
2y ago
This is one of these technologies that is indistinguishable from magic.
54.
▲
by
WanderPanda
2y ago
Agreed, on another note I also struggle to see how no one could create a (better) CUDA implementation for e.g. $1B engineering budget
55.
▲
by
WanderPanda
2y ago
I washed my Airpods in a washing machine once and they continued working normally for 6 months until the noise cancelling started to fail on one side.
56.
▲
by
WanderPanda
2y ago
I found it fascinating how far we can go with even just < 10 kilobytes per second [1] [1] https://bellard.org/tsac/
57.
▲
by
WanderPanda
2y ago
I would have expected fine-tuning to be good at imparting knowledge. Pre-training is often done for a single epoch only and models soak up the knowledge like crazy without multiple passes so why would fine-tuning be any different?
58.
▲
by
WanderPanda
2y ago
Compiler folks: Is there any chance compilers will be able to find optimizations like FlashAttention on their own? Seems like TVM and tinygrad are working in that direction but I find it hard to believe that that would be feasible
59.
▲
by
WanderPanda
2y ago
I enjoyed the docu recently and I‘m still in awe how „software-defined“ the voyager system was/is for it’s time
60.
▲
by
WanderPanda
2y ago
Also that level of hustle is not (yet) common in Germany I believe
More ›