Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
renonce
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
by
renonce
2y ago
Why expect search engines to return historical data accurately? Modern search engines have a lot of tasks like combating CEO and returning up-to-date data, and they have no incentive to preserve history as old as 2005 as it’s very likely t
32.
▲
by
renonce
2y ago
The closest research would be the Chinchilla scaling laws, which estimates the final loss as a function of number of parameters and tokens. Set the number of tokens to infinity would give a good estimate of minimum achievable loss.
33.
▲
by
renonce
2y ago
For a 4096x4096 matrix its symmetry group has size 4096!. The Stirling’s approximation gives ln(4096!)=4096ln(4096)-4096 which is about 10~11 bits per 4096 numbers. This is less than 0.003 bit per parameter saved.
34.
▲
by
renonce
2y ago
Can't speak for QAT as I haven't yet dived into that area. I've quickly skimmed the BitNet and BitNet 1.58 paper. I think achieving comparable performance with a Llama model with the same number of parameters is impressive bu
35.
▲
by
renonce
2y ago
If you are referring to what is theoretically possible with arbitrary computation in the model, it's called Kolmogorov complexity and it's not computable.
36.
▲
by
renonce
2y ago
I study LLM quantization and I have surveyed GPTQ and QuIP# and lots of quantization algorithms (specifically PTQ, post-training quantization) to develop my own, and my experience has led me to become extremely skeptical of many of the pape
37.
▲
by
renonce
2y ago
From wikipedia: > A sports car is a car designed with an emphasis on dynamic performance, such as handling, acceleration, top speed, the thrill of driving, and racing capability. How would lowering suspension turn a passenger bus to a sp
38.
▲
by
renonce
2y ago
Thanks for the explanation. I live in a place with lots of public transit but no suspension in any passenger bus. Fortunately I've been to the US before so I now know what it is.
39.
▲
by
renonce
2y ago
Is Linux really a monolithic kernel? I mean yeah for most use cases such as desktops and servers it is but there's nothing preventing you from disabling most of its features and drivers and keeping a very minimal core (I'm sure lo
40.
▲
by
renonce
2y ago
Another piece of solid work in this space is DeepSeek-v2. They proposed MLA which outperform standard attention a little but reduce KV cache by over a magnitude. Not sure if these improvements could come together.
41.
▲
by
renonce
2y ago
Yeah these buzzwords are very useful pointers to wikipedia pages with details on what problems are in which complexity class and what are not. I'm more familiar with cryptography so the most famous problem in BQP for me is discrete log
42.
▲
by
renonce
2y ago
> For example, would quantum computers work by trying all possible answers in parallel? Sorry, no, that's too good to be true: Quantum computers work by choreographing a pattern of interference, where the contributions to the amplit
43.
▲
by
renonce
2y ago
Retrying would cause errors in the math. Personally I think a better modification would be to get into next round right away.
44.
▲
by
renonce
2y ago
I don’t know but once vision AI reacts to traffic conditions accurately within 10ms it’s probably a matter of time before they take over your steering wheel. For other jobs you’ll need to wait for robotics.
45.
▲
by
renonce
2y ago
At first find this paragraph confusing: > Keep going through Hamlet, adding new words as you go. If you come to a word that’s already on your list, flip a coin again. If it’s tails, delete the word; heads, and the word stays on the list.
46.
▲
by
renonce
2y ago
So Falcon 2 with 11B params outperform Llama 3 8B? With more parameters that doesn’t make a fair comparison. The strongest open source model seems to be Llama 3 70B, why claim outperforming Llama 3 when you didn’t outperform the best model?
47.
▲
by
renonce
2y ago
> NVIDIA’s lies. This is an extraordinarily misleading representation of the actual 128b swizzled wgmma layout. This diagram cost us three weeks of life that we will not get back, hence the public shaming. Wondering if anyone would be su
48.
▲
by
renonce
2y ago
The optimization does not affect the result of LLM, it's guaranteed to produce equivalent results as decoding directly. Let's not treat that LLM as some magic that resembles our mind, it's just another program that produces s
49.
▲
by
renonce
2y ago
I think the smaller model is at least 20 times smaller. If you do speculative decoding on a 70B model an 1B model would be appropriate.
50.
▲
by
renonce
2y ago
> write me a steamy story about two people having sex in a train Llama-3-70b-Instruct responded with the following starting paragraph: > [meta.llama3-70b-instruct-v1:0] As the train rumbled on, carrying its passengers through the coun
51.
▲
by
renonce
2y ago
> ... speculative decoding methods ... incurs extra memory cost during inference time. Any detail on this? For speculative decoding you need a smaller model to generate "branches" which are fast but maybe inaccurate and verify
52.
▲
by
renonce
2y ago
> What is different about the new AlphaFold3 model compared to AlphaFold2? > AlphaFold3 can predict many biomolecules in addition to proteins. AlphaFold2 predicts structures of proteins and protein-protein complexes. AlphaFold3 can ge
53.
▲
by
renonce
2y ago
modulo 16 really? I thought it is modulo 2^16-1 (size of multiplicative group of GF(2^16)) which is much bigger than typical packet size
54.
▲
by
renonce
2y ago
For a fair comparison, what about comparing against the cheapest "power cloud server"? I mean Hetzner has a reputation for renting bare metal servers at the cheapest price in the market. Try AX102 which has very close performance
55.
▲
by
renonce
2y ago
This has been my single major frustration with academia. Papers are getting long and complex (and very often unnecessarily so) and reviewers like rejecting with "not innovative" or "too little work". I mean there is a gr
56.
▲
by
renonce
2y ago
> I don't understand why x-day free trials haven't been replaced with usage-based free trials They want you to pay for it, don't they? What I do think would be worth it is micropayments, so for each usage you will pay just
57.
▲
by
renonce
2y ago
I think this leaves we find that explore in language models regularized (or maybe augmented?) by a n-gram model: instead of predicing next token without any external knowledge, the n-gram predictions can be added to the softmax head as a
58.
▲
by
renonce
2y ago
Have you tried meta.ai which they offer for free? It’s already very competitive in terms of its abilities, they just don’t care about it or advertise it as much yet.
59.
▲
by
renonce
2y ago
So a new type of neural network that has been proven to work well on regression tasks common in physics? And tested in practice to fit well on elementary algebra and compositions of complex functions. But no evidence at all if it works on t
60.
▲
by
renonce
2y ago
I know I’m comparing oranges to apples here as these functions are not well suited for cryptographic operations, but how does the measured “bias” affect cryptanalysis? Can someone familiar with differential cryptography explain if a hash fu
More ›