Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
brrrrrm
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
61.
▲
by
brrrrrm
2y ago
WebGPU cannot even come close unfortunately since they don't have support for hardware specific memory or warp-level primitives (like TMA or tensorcores). it's not like it gets 80% of perf, it gets < 30% of the peak perf for a
62.
▲
by
brrrrrm
2y ago
I don't think this is absurd at all, I'm in the exact same boat. In fact, I suspect most people have far more sophisticated relationships with digital companies these days than ever before. Grievances like cancellation pain are a
63.
▲
by
brrrrrm
2y ago
fake it. add some latency to the first token and then "stream" at the rate you received tokens even though the entire thing (or some sizable chunk) has been generated. that'll give you the buffer you need to seem fast while
64.
▲
by
brrrrrm
2y ago
an oracle filtering a random generator will always produce oracle-grade results. is the computer actually getting better or is Lenat just acting as an oracle and the writer is running with it?
65.
▲
by
brrrrrm
2y ago
that silly softmax1 blog post is not worth the read. no one uses it in practice if you think about it, the "escape hatch" is the design of the entire transformer dictionary. if Key/Query attention misaligns with Value's
66.
▲
by
brrrrrm
2y ago
yea. everything. try it out (point it at a website or something)
67.
▲
by
brrrrrm
2y ago
it's not about the API. its about the documentation + ecosystem. TF's doesn't seem very good. I just tried to figure out how to learn a linear mapping with TF and went through this: 1. googled "linear layer in tensorfl
68.
▲
by
brrrrrm
2y ago
You almost always have to for good perf on non-trivial operations
69.
▲
by
brrrrrm
2y ago
depends on your phone, but try a couple of these variants with ollama https://ollama.com/library/llama3.2/tags e.g. `ollama run llama3.2:1b-instruct-q4_0`
70.
▲
by
brrrrrm
2y ago
Any citations for that? Id have thought the goofy width of spectacles was related to the screen projection
71.
▲
by
brrrrrm
2y ago
should really be titled streaming output, as full duplex streaming isn't mentioned at all. that'd be necessary for things low latency things like speech etc.
72.
▲
by
brrrrrm
2y ago
goes to show formal language isn't a necessary component of high communication. might even be antagonistic
73.
▲
Chicken Tax
(en.wikipedia.org)
2 points
by
brrrrrm
2y ago
|
0 comments
74.
▲
by
brrrrrm
2y ago
It kinda has already with fixed matrix multiplication units. But beyond that, no chance. Bitcoin is an unchanging hash algo, not a developing software
75.
▲
by
brrrrrm
2y ago
If you're hitting a memory wall it means you're not scaling. This stuff really doesn't apply to scaled up inference but rather local small batch execution
76.
▲
by
brrrrrm
2y ago
perhaps the title should be "Llama2-70B worse than humans ..."
77.
▲
by
brrrrrm
2y ago
> actually conducting research 5 people evaluating 45 responses? This doesn't take a year to do. The issue is that this study is was poorly funded and slow - model development has far outpaced the results and there's likely l
78.
▲
by
brrrrrm
2y ago
> The most promising model, Meta’s open source model Llama2-70B This is an old model that was not dominant even when released. This study must be fairly old or I question the qualification of the group running it.
79.
▲
by
brrrrrm
2y ago
It’s all about the kernels tho. The language doesn’t matter much. For the things that matter, everything is a dispatch to some cuda graph I’m not really a fan of this convergence but the old school imperative CPU way of thinking about thin
80.
▲
by
brrrrrm
2y ago
not really. embeddings are embeddings. check out llava
81.
▲
Up to 1.9X Higher Llama 3.1 Performance with Medusa
(developer.nvidia.com)
3 points
by
brrrrrm
2y ago
|
0 comments
82.
▲
by
brrrrrm
2y ago
LLMs can deal with more than text. Impressive today is nothing tomorrow
83.
▲
by
brrrrrm
2y ago
we've like barely trained these things? the entirety of common crawl is 424 terabytes. that's merely 6 days of 8K raw video.
84.
▲
by
brrrrrm
2y ago
So, noticing that linearized models have tiny KV caches ahem i mean state spaces, this approach tries to increase their size along the embedding dimension. Increasing this enormously by applying a different softmax (which is compatible w
85.
▲
by
brrrrrm
2y ago
how does this approach differ from Nvidia's 2019 writeup on using trees to improve allreduce operations? https://developer.nvidia.com/blog/massively-scale-deep-learn...
86.
▲
by
brrrrrm
2y ago
this is true of even just matrix multiplication (A*B) of which attention has two
87.
▲
by
brrrrrm
2y ago
For most LLM workloads today (short text chats), hundreds or a couple thousand tokens suffice. attention mechanisms don’t dominate (< 30% compute). But as the modalities inevitably grow, work in attention approximation/compression
88.
▲
by
brrrrrm
2y ago
amazingly, the browser has pretty much solved all of this. fully compatible EMCAscript implementations on every single device with hardware access (such as the camera, as is needed in this post) I don't buy your "security through
89.
▲
by
brrrrrm
2y ago
Empowered with chatGPT (now claude), my partner as a designer pulls off extremely workable implementations of all different ideas without much need for sketching in a non-functional design tool. When she does draft in design tools it’s usua
90.
▲
by
brrrrrm
2y ago
check out the paper. it's pretty comprehensive https://ai.meta.com/research/publications/the-llama-3-herd-o...
More ›