Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
musebox35
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
61.
▲
by
musebox35
1y ago
Honestly, if you have any actual interest in LLMs or other generative ai variants, just go after a concrete goal post that you yourself set with measurable metrics to gauge your progress. Then the predicted timeline from podcasts and blog p
62.
▲
by
musebox35
1y ago
AFAIK, certain abilities such as understanding arithmetic manifest at discrete scale points even though there is a continuous build up of potential. There is also the more remote possibility of a discrete scale that AI takes over its own tr
63.
▲
by
musebox35
1y ago
While I agree with the general sentiment on throwaway compute infra, the generated know-how with large scale experiments is not thrown away. I think a lot hinges on the scaling laws and whether you will hit the jackpot at a certain scale be
64.
▲
by
musebox35
1y ago
I think conceptually diataxis is brilliant. However, it is not trivial to implement. Every project needs a varying ratio of each component and stacking all forms in a single website in the same format is very ineffective. The ratio also evo
65.
▲
by
musebox35
1y ago
Brilliant, I have always felt that one of the major problems with machine learning, consequently LLMs, is the boring average based loss functions that under-represent the unique and the rare. It seems our collective civilization is using a
66.
▲
by
musebox35
1y ago
That is a fine point. However I am not sure if replacing the gpus themselves will be the bottleneck investment for datacenter costs. After all you have so much more infrastructure in a datacenter (cooling and networking). Plus custom chips
67.
▲
by
musebox35
1y ago
I found the book from David Mackay on Information Theory, Inference, and Learning Algorithms to be well written and easy to follow. Plus it is freely available from his website: https://www.inference.org.uk/itprnn/book.
68.
▲
by
musebox35
1y ago
Context is also a bottleneck in many human to human interactions as well so this is not surprising. Especially juniors often start by talking about their problems without providing adequate context about what they’re trying to accomplish or
69.
▲
by
musebox35
1y ago
"The prompt could be perfect, but there's no way to guarantee that the LLM will turn it into a reasonable implementation." I think it is worse than that. The prompt, written in natural language, is by its very nature vague an
70.
▲
by
musebox35
1y ago
With the rise of LLM training, Nvidia’s main revenue stream switched to datacenter gpus (>10x gaming revenue). I wonder whether this have affected the quality of these consumer cards, including both their design and product processes: h
71.
▲
QuACK: A Quirky Assortment of Cute Kernels
(github.com)
1 points
by
musebox35
1y ago
|
1 comments
72.
▲
by
musebox35
1y ago
CuTe DSL examples from the MIT Dao-AILab for writing high performance cuda kernels using Python (see https://github.com/NVIDIA/cutlass for more background info).
73.
▲
by
musebox35
1y ago
From the acknowledgment at the end, I guess the author has access to TPUs through https://sites.research.google/trc/about/ This is not the only way though. TPUs are available to companies operating on GCP as an al
74.
▲
by
musebox35
1y ago
I think https://jax-ml.github.io/scaling-book/ is one of the best references to go through. It details how single device and distributed computations map to TPU hardware features. The emphasis is on mapping the transfo
75.
▲
by
musebox35
1y ago
I totally agree that the resulting kernel will be rarely useful. I just wanted to highlight that it is a commonly used educational exercise to showcase how to optimize for memory throughput. If the post showed how to fuse a transpose + rmsn
76.
▲
by
musebox35
1y ago
Matrix transpose is a canonical example of a memory bound operation and often used to showcase optimization in a particular programming language or library. See for example the cutlass matrix transpose tutorial from Jay Shah of flash attent
77.
▲
by
musebox35
2y ago
I suggest having a look at https://m.youtube.com/@GPUMODE They have excellent resources to get you started with Cuda/Triton on top of torch. It also has a good community around it so you get to listen to some amazing p
78.
▲
by
musebox35
2y ago
Considering recent developments in GPU hardware (Tensor Cores for GEMM), another hardware accelerated algorithm is ray tracing for photo-realistic rendering. As far as I understand the Ray Tracing Cores provide an efficient hardware impleme
79.
▲
by
musebox35
2y ago
Fun to read. The style of the article reminded me of https://scholar.harvard.edu/files/mickens/files/thenightwatc...
80.
▲
by
musebox35
3y ago
It seems to me that technological developments and empirical scientific breakthroughs come in cycles, technology making it cheaper to experiment, science reducing cost of new technical developments. I would be happy to hear about pointers t
81.
▲
by
musebox35
3y ago
For a robotics oriented, part theory part hands-on learning material on Kalman Filtering, I would suggest The Robot Mapping course from Cyrill Stachniss: http://ais.informatik.uni-freiburg.de/teaching/ws22/mapping&
82.
▲
by
musebox35
3y ago
This seems like a specific application of the "inverted thinking/inversion" mentality. I once came across a post on this but can not find it now. I think it was about this article: https://fs.blog/inversion&#x