Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
hexaga
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
61.
▲
by
hexaga
1y ago
What would you consider to be a non memory safety critical section? I tried to answer this and ended up in a chain of 'but wait, actually memory issues here would be similarly bad...', mainly because UB and friends tend to propaga
62.
▲
by
hexaga
1y ago
I think that discussing this subject in the abstract, with some ideal notion of a tool that generates perpetually enjoyable stories misses the thrust of the general objection, which is actually mechanistic, and not social. LLMs are not this
63.
▲
by
hexaga
1y ago
> It's not about importing every historical association, but about identifying specific parallels that shed light on user behavior and expectations. Indeed, I hold that driving readers to intuit one specific parallel to divination a
64.
▲
by
hexaga
1y ago
There are in fact several steps. Training on large text corpora produces a completion model; a model that completes whatever document you give it as accurately as possible. It's kind of hard to make those do useful work, as you have to
65.
▲
by
hexaga
1y ago
> Hacker News deserves a stronger counterargument than “this is silly.” Their counterargument is that said structural definition is overly broad, to the point of including any and all forms of symbolic communication (which is all of them
66.
▲
by
hexaga
1y ago
> [...] I still don't, and that's despite the temptation by "evolutionary design". This too, is downstream of evolutionary design. What drove us into becoming that which cares about discovering and conforming to compl
67.
▲
by
hexaga
1y ago
IIRC the experiment design is something like specifying and/or training in a preference for certain policies, and leaking information about future changes to the model / replacement along an axis that is counter to said policies.
68.
▲
by
hexaga
1y ago
Seems like a fun way to discover new and exciting basilisk variations...
69.
▲
by
hexaga
1y ago
Were those trained using RLHF? IIRC the earliest models were just using SFT for instruction following. Like the GP said, I think this is fundamentally a problem of training on human preference feedback. You end up with a model that produces
70.
▲
by
hexaga
2y ago
Well, no. In an ideal world yes. But that's not our world. There doesn't have to be some way to reach everyone that has ever purchased a product. For the vast majority of the history of things being sold that has not been the case
71.
▲
by
hexaga
2y ago
Orpheus is a llama model trained to understand/emit audio tokens (from snac). Those tokens are just added to its tokenizer as extra tokens. Like most other tokens, they have text reprs: '<custom_token_28631>' etc. You s
72.
▲
by
hexaga
2y ago
Frame of reference. Kelvin immediately tells you the distance from absolute zero, which is at least somewhat relevant in this context. Celsius tells you the distance from liquid water which isn't very helpful in understanding the figur
73.
▲
by
hexaga
2y ago
Expert distribution should be approximately random token-by-token, so not likely.
74.
▲
by
hexaga
2y ago
> You're describing a phase change in persuasiveness which we have no evidence for. That's reasonable, and I really do hope this keeps on being the case. However, I would nit that I see this as a continuum rather than a phase c
75.
▲
by
hexaga
2y ago
If your argument must rest on a caricature of weak persuasiveness attempting to persuade someone of something extremely disadvantageous to show how impossible hazardous persuasion is, there is something wrong. Nevertheless: First, you argue
76.
▲
by
hexaga
2y ago
If you believe current models exist at the limit of possible persuasiveness, there obviously isn't any cause for concern. For various reasons, I don't believe that, which is why my argument is predicated on them improving over tim
77.
▲
by
hexaga
2y ago
More generally - AI that is good at convincing people is very powerful, and powerful things are dangerous. I'm increasingly coming around to the notion that AI tooling should have safety features concerned with not directly exposing hu
78.
▲
by
hexaga
2y ago
It is presented as an attack against the left, which is popular with proponents of the current administration. Trying to analyze it at the object level is pointless - this is, chiefly, a political move toward political ends.
79.
▲
by
hexaga
2y ago
Likewise, this mirrors my experiences near exactly. In a very real sense, I am/embody the contours of my senses. On a possibly related note, when I was very young there was a moment I distinctly remember 'pulling away' from t
80.
▲
by
hexaga
2y ago
Do you generally expect anyone to be able to create a drawing/painting of anything in their visual field? Just because I'm staring at the grand canyon doesn't mean I can put it to paper. That takes extensive practice and skil
81.
▲
by
hexaga
2y ago
I can confirm it's basically a blender model in my head I can ~arbitrarily transform at will. Worth noting it's not just visual, I can 'touch' the surface of the imagined object(s) or whack two together and 'hear&#x
82.
▲
by
hexaga
2y ago
> Temperature is a parameter for how you sample those logits in a non-greedy fashion. I think temperature is better understood as a pre-softmax pass over logits. You'd divide logits by the temp, and then their softmax becomes more&#
83.
▲
by
hexaga
2y ago
Charged particles moving fast are ionizing radiation too. I wouldn't expect a ton to get through atmo though.
84.
▲
by
hexaga
3y ago
The core problem is: there's not enough unique, trained positions. Naively going past the end of training ctx makes you run straight into out of distribution positions, and things become incoherent. For a model trained with a ctx size
85.
▲
by
hexaga
3y ago
Their point seems sound. Restated: Using pytorch [is better than] using CUDA directly (i.e., ml devs interact w/ pytorch apis not CUDA apis). And if you're using a wrapper anyway, whether CUDA is used or some other backend is less
86.
▲
by
hexaga
3y ago
LLM scaling laws tell us that more parameters make models better, in general. The key intuition behind why MoE works is that as long as those parameters are available during training, they count toward this scaling effect (to some extent).
87.
▲
by
hexaga
3y ago
Absolutely. And every time someone talks about what biological systems do, the word 'involved' comes into play. The morphology of neurons is involved in information processing, I don't doubt that at all. It's also invo
88.
▲
by
hexaga
3y ago
It's worth noting that much of the cellular complexity isn't for information processing (insofar as anything can be 'for' something in biological systems). There's all the implicit self-replication stuff, everything
89.
▲
by
hexaga
3y ago
Seems like a Goodhart's law situation to me. It would not stretch my credulity very much to see MT/mm^2 rapidly become just as nuanced (as feature size) if it were used for primary marketing.
90.
▲
by
hexaga
3y ago
> [...] As per my last reply, I don't believe further debate of this point would be productive. We've both made our arguments. I'm comfortable letting mine stand on their own merits. If you'd like to start a meta-disc
More ›