Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
psb217
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
61.
▲
by
psb217
2y ago
Well, the ability to identify talent is itself a talent. The skills for identifying talent may be relatively unaligned with demands of the field in which one is identifying talent, but I figure it's helpful to be at least competent in
62.
▲
by
psb217
2y ago
I'll be deep in the cold, cold ground before I recognize Missouri. -- https://www.youtube.com/watch?v=ZoWc6WRHKEE
63.
▲
by
psb217
2y ago
If the corruption can be revealed, it exists. This is true whether or not the decision to reveal it is biased.
64.
▲
by
psb217
2y ago
I figure a reasonable rule of thumb is that if someone got to the top of some system by maximizing some metric X, where X is the main metric of merit in that system, then they're unlikely to push for the system to prefer some other met
65.
▲
by
psb217
2y ago
My point was mainly that this claim: "we keep seeing LLMs building more robust approximations of real world models" is hard to evaluate without a well-formed definition of what it means to have a world model. Eg, a more restrictiv
66.
▲
by
psb217
2y ago
The othello paper is annoying and oversold. Yes, the representations in a model M trained to predict y (the set of possible next moves) conditioned on x (the full sequence of prior moves) will contain as much information about y as there is
67.
▲
by
psb217
2y ago
This is a bit pedantic, but FID score wouldn't really be a viable metric for best of n selection since it's a metric that's only computable for distributions of samples. FID score is also pretty high variance for small sample
68.
▲
by
psb217
2y ago
Why bother with GPS or other "absolute" coordinate systems? Once the rocket's in close, all that matters is relative position and orientation of the rocket with respect to the landing apparatus. Eg, if you had many sensors in
69.
▲
by
psb217
2y ago
It's a classic "Will it work? IDK, maybe. Let's try it and find out..." paper.
70.
▲
by
psb217
2y ago
The technique in this paper would still be rightly described as distillation. In this case it's distillation of "internal" representations rather than the final prediction. This a reasonably common form of distillation. The i
71.
▲
by
psb217
2y ago
In this case I'm talking strictly about his persona, on-stage performance, and overall "vibe". To an extent, these traits are orthogonal to his claimed policy/politics in the traditional sense, which I agree are basicall
72.
▲
by
psb217
2y ago
There's a big practical difference between chewing coca leaves and smoking crack. Also, the shift in personality and tone from Bush to Trump are... not small. The inconoclast and anti-establishment things are intentional, effective and
73.
▲
by
psb217
2y ago
"I highly doubt there are hundreds of researches earning millions." -- by doing purely academic research, maybe not. But, the number of people who have moved from academia to industry off the strength of their research and made mi
74.
▲
by
psb217
2y ago
The one thing that immediately stood out to me in the ghost example was how the face of the ghost had "wobbly geometry" and didn't appear physically coupled to the sheet. This and the way the fruit in the sloth's drink m
75.
▲
by
psb217
2y ago
All information about the past which will be available for predicting future tokens must be stored in the present state. So, if some bits of info about some past tokens at times less than t_p will be used for predicting some future token at
76.
▲
by
psb217
2y ago
Well, that's what Transformer already does... One problem with the scaling you're describing is that there would be a massive amount of redundant information stored in hidden activations during training the RNN. The hidden state a
77.
▲
by
psb217
2y ago
Well, you could certainly train a big-ass model to mimic the distribution of all that physics data. That doesn't mean the model could, eg, formulate interesting new theories which explain why that distribution has its particular struct
78.
▲
by
psb217
2y ago
I think it would be hard to make a solid argument that AR or non-AR is strictly better wrt full sequence error rates, whether or not we place constraints on compute, memory, etc. I'd guess that there's some intrinsic form of compl
79.
▲
by
psb217
2y ago
A sequence of tokens can be converted back to the sequence of tokenized characters without loss of information. Eg, how do you think text is rendered for the user based on sequences of tokens generated by the LLM? Different tokenization sch
80.
▲
by
psb217
2y ago
The per token error of the non-AR model wrapped with MPC is no higher than the per token error of the non-AR model without MPC. Likelihood of the entire sequence being off the true data manifold is just one minus the product of the per toke
81.
▲
by
psb217
2y ago
"The difference between LLMs and other kinds of predictive models, or humans, is that those kinds of systems do not produce their output one token at a time, but all in one go, so their error basically stays constant." -- This is
82.
▲
by
psb217
2y ago
Yes, professional athletes and hollywood stars are famously compensated on seniority-based schedules.
83.
▲
by
psb217
2y ago
Mostly to keep up with website/app bloat, and better cameras if you like taking photos or videos. Other improvements like better screens or battery life are nice too, but only really noticeable if you're jumping a few phone genera
84.
▲
by
psb217
2y ago
Tokenization does not remove information from the input[1]. All the information required for character counting is still present in the input following tokenization. The reasons you give for why counting characters is hard could be applied
85.
▲
by
psb217
2y ago
Tokenization reorganizes information but doesn't remove it. It may be easier/harder to learn stuff like letter counting with different tokenization schemes, but the main reason it's hard is that there's not much text abo
86.
▲
by
psb217
2y ago
Hmm, I wonder if anyone has a simple pipeline for extracting data for "voice cloning" type models from the combination of original audio and transcribed text. It should be possible to chain this with some post-processing to replac
87.
▲
by
psb217
2y ago
It's a tricky balance. Once the signal to noise gets bad enough, even if the cost of ignoring individual events is low enough, the cumulative cost of ignoring so much overwhelms any positive value. Twitter now seems bad enough that it&
88.
▲
by
psb217
2y ago
In a sense, poorly reproducing rare content is a form of compression artifact. Ie, since this content occurs rarely in the training set, it will have less impact on the gradients and thus less impact on the final form of the model. Roughly
89.
▲
by
psb217
2y ago
Academic authors are consistently better at editing away unclear and ambiguous statements which make their work seem less impressive compared to ones which make their work seem more impressive. Maybe it's just a coincidence, lol.
90.
▲
by
psb217
2y ago
It's funny how academic writing works. Authors rarely produce many unclear or ambiguous statements where the most likely interpretation undersells their work...
More ›