Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
krackers
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
91.
▲
by
krackers
3mo ago
>Maybe we're somehow treating f(0) = 0 so that you can apply it directly Hm thinking about it a bit more, I think what's going on is that you treat the baseline of hidden layer L at which you apply the J-lens as 0 activation, t
92.
▲
by
krackers
3mo ago
Part of the annoying thing is that if you're working on a product which uses LLMs, at some level you run out of levers to pull in terms of being able to fix things. At best you're stacking hacks on top of hacks to prevent unwanted
93.
▲
by
krackers
3mo ago
The details seem to be present in the paper (section 2.1). I'm still trying to understand, but it seems instead of computing gradients with respect to cross-entropy loss for the 1-hot "next word" vs output logits, you comput
94.
▲
by
krackers
3mo ago
If they were going to do this, they must have known a few days in advance. Feels intentional.
95.
▲
by
krackers
3mo ago
You can make eggs in a microwave (critically so long as you don't do it in the shell)
96.
▲
by
krackers
3mo ago
Youtube comments are also links given by the site. I think in this case it's not necessarily the prompt injection that's the issue but the fact that untrusted content allows formatted links. YouTube doesn't allow clicabkle li
97.
▲
by
krackers
3mo ago
LLMs can learn to do arithmetic (without tool use), and they can learn a mapping from tokens to the letter counts contained therein (you could imagine trivially training on synthetic data). So there doesn't seem to be any fundamental b
98.
▲
by
krackers
3mo ago
Was that ever solved? It seems that entire retort faded overnight, yet to my knowledge there was never any systematic analysis on cause or tokenizer change that fixed it. Maybe we just decided that this failure mode doesn't have any pr
99.
▲
by
krackers
3mo ago
https://support.apple.com/en-us/102174 >A Threat Notification is displayed at the top of the page after the user signs into account.apple.com. >Apple sends an email and iMessage notification to the email addresse
100.
▲
by
krackers
3mo ago
>I expect that Apple's TextEdit.app is just a wrapper around the rich text control in Cocoa https://developer.apple.com/library/archive/samplecode/TextE...
101.
▲
by
krackers
3mo ago
>discontinued in case you don't know, there's back and made by incase. The first manufacturing run completely sold out I think, it's backordered until sep 2026. The matias is... not a good replacement, see https:/&#x
102.
▲
by
krackers
3mo ago
Isn't that one of the reasons why KL-divergence is used, at least in DPO/RL for LLM? Otherwise the model can effectively cheat and mode collapse. For pre-training against a 1-hot label the KL-divergence should be equivalent to cro
103.
▲
by
krackers
3mo ago
> thought for like 20 minutes then just told me it was all "inevitable" I have in mind an image of ASI as something that's able to seamlessly work across time as if it was weaving cloth. Reasoning about not just first or s
104.
▲
by
krackers
3mo ago
Would you be better off pooling that money with some hackerspace group and then setting up shared inference infra, so that way you at least get better utilization?
105.
▲
by
krackers
3mo ago
Game Theory seems sort of useless in the real world because people are not rational players, and the real challenge is in getting an accurate model of their behavior. The honor system would work probably fine in a tiny close-knit liberal ar
106.
▲
by
krackers
3mo ago
>Energy is force times distance This is not true though, work is only necessary to change KE. An object can have kinetic energy even when no forces are presently acting on it (of course force was needed to bring it from resting to that s
107.
▲
by
krackers
3mo ago
The fact that GLP-1 seems to have roles not just in satiety but that agonists seem to reduce other types of impulsiveness (e.g. gambling, shopping) is interesting. That's not something you'd predict as a consequence, and perhaps i
108.
▲
by
krackers
3mo ago
I think this is covered by the "overdraft" section, if the only way to know for sure is to just submit it.
109.
▲
by
krackers
3mo ago
> The review is behind a paywall, but not expensive. I think the author wrote a twitter post with a summary of the content, and someone on twitter who had read the original Chinese source also chimed in with a summary https://
110.
▲
Thinking to recall: How reasoning unlocks parametric knowledge in LLMs
(research.google)
3 points
by
krackers
3mo ago
|
0 comments
111.
▲
by
krackers
3mo ago
To play devil's advocate, how is a project supposed to distinguish between your patch and "slop" without a reviewer having to put in effort to vet it. Especially since the patch was drafted by LLM, it seems fair to be immedia
112.
▲
Exploring the internal representations of Pangram 3.3.2
(pangram.com)
34 points
by
krackers
3mo ago
|
5 comments
113.
▲
by
krackers
3mo ago
I mean sliding window attention is the most basic way of getting long context window. For the OCR case it seems like it should be even simpler, since you don't even need to have the "sliding" portion, unless I"m missing
114.
▲
by
krackers
3mo ago
>Additionally, a lot of existing web servers by default ignore GET requests with a body. I think the point made is that _all_ existing web servers have no idea what a "QUERY" is anyway, so changes need to be made anyhow.
115.
▲
by
krackers
3mo ago
"Compute the first few terms and plug into OEIS" is very high on the reward:effort scale
116.
▲
by
krackers
3mo ago
>Is there a similar trick to poison an LLMs weights during training? Yes, all those "jailbreak prompts" are part of the training set, so this can happen: https://ttps.ai/procedure/x_bot_exposing_itself_afte
117.
▲
by
krackers
3mo ago
The highest return small local model for me has been the in-built OCR that macOS has. It has finally "solved" OCR by making high-quality results accessible to everyone. Yet the state of art outside the apple ecosystem seems to be
118.
▲
by
krackers
3mo ago
To understand the threat model you need to understand historical decisions browsers made, such as when cookies are sent, and the distinction between actually sending the request versus allowing client-side JS to read back request content. T
119.
▲
by
krackers
3mo ago
You can't just make a post like this and not say what the condition was or what you did!!
120.
▲
by
krackers
3mo ago
This sounds similar to barycentric coordinates?
More ›