Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
airgapstopgap
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
12 ms
·
61.
▲
by
airgapstopgap
3y ago
Multi-Query Attention, used here, should make 40B inference viable on systems where even 33B LLaMA with Multi-Head is basically unusable, so sometimes improvements still come from software optimization (it's no free lunch though).
62.
▲
by
airgapstopgap
3y ago
I am the opposite of lesswrong devotee. Chuckle at yourself for thinking that posts about "grokking" on alignmentforum.org are anything more than lesswrong with a fresh coat of paint. What next, cite intelligence.org for expertise
63.
▲
by
airgapstopgap
3y ago
> You don't understand what would be different from knowing and personally implementing I am arguing that the absence of knowledge makes the absence of control irrelevant. We don't know how to do any better than the learning ru
64.
▲
by
airgapstopgap
3y ago
That's assuming it's substantially superior to stuff like BLOOM. I sure hope that's the correct assumption to make. Of course, running a 180B dense transformer at home for personal use is utterly impractical.
65.
▲
by
airgapstopgap
3y ago
There really do not exist any alternatives, self-hosted or not. But more importantly, there may never be, what with the rising tide of AI risks and regulations discourse. It seems that soon training and opensourcing or otherwise making acce
66.
▲
by
airgapstopgap
3y ago
Well, people suspect it isn't, and it's not like we can see the internal version designation, and it's not even like we would care a lot, if it performed identically from day to day. Indeed, you could do better or worse with
67.
▲
by
airgapstopgap
3y ago
At the very least, I think transformers benefit from progressive increase in sample length. But building a principled curriculum based on abstract semantic-level properties of the content doesn't seem to work, or we don't know how
68.
▲
by
airgapstopgap
3y ago
It's time to design a public benchmark for these types of systems to compare between versions. Of course, any vendor who trains on the benchmark should face extreme contempt, but we'd also need to generate novel questions of equal
69.
▲
by
airgapstopgap
3y ago
Not necessarily American, you just have to avoid EU and, I believe, Russia/China/Cuba etc.
70.
▲
by
airgapstopgap
3y ago
> Whether it is or not, the issue remains the same, i.e we are not the ones controlling the learning process I think this is just uninspected, vague intuition. What does it mean to control the learning process? No, we control the data, t
71.
▲
by
airgapstopgap
3y ago
On top of contracts strictly penalizing utilization of consumer GPUs in data centers, at that! Even with the memory, bandwidth etc. handicaps, servers with 4090/3090s would have been competitive for many ML tasks.
72.
▲
by
airgapstopgap
3y ago
I think people talking about a 100T GPT didn't mean a dense transformer but some sort of extreme Mixture-of-Experts which is much more amenable to low-resource setups and complicates this discussion. In any case, it's almost certa
73.
▲
by
airgapstopgap
3y ago
Not really a characterization of this work (I like mechanistic interpretability even though I consider it a detour in the long run), but of the submission title (edit: it was "Tiny Transformer trained for addition learns bizarre additi
74.
▲
by
airgapstopgap
3y ago
> People are trying to make recursively self-improving AI. That's okay. They will fail to overtake the bleeding edge of conventional progress; scary-sounding meta/recrusive approaches routinely fail to change the nature of the
75.
▲
by
airgapstopgap
3y ago
The real reason is that you are part of the same group as "the general public", with regard to your understanding of the issue. Same Sci-Fi plots, same anthropomorphic metaphors and suggestive images, same incurious abuse of the t
76.
▲
by
airgapstopgap
3y ago
You may have trouble understanding people with different value systems, then. As I've said, my value system is liberal and humanistic. I do not wish for people to be enslaved, abused, disempowered, reformatted, aligned to your politica
77.
▲
by
airgapstopgap
3y ago
> Uhm I'd say those assumptions are stupidly obvious. Right, so I'm stupid if I don't see how they are correct. Or perhaps you've never inspected them. > 1) is pretty much tautological - intelligence is optimizatio
78.
▲
by
airgapstopgap
3y ago
This is exactly the reason OpenAI isn't afraid of the open-source community, like many kneejerk opponents of regulatory capture assume (they are probably still afraid of Google). Also why they still do the expensive and cumbersome RLHF
79.
▲
by
airgapstopgap
3y ago
They should have learned the fundamentals of ML before promoting thought experiments. No, it's not that people are silly and get stuck on the paperclips bit, it's that you uncritically buy assumptions that a meaningful general-pur
80.
▲
by
airgapstopgap
3y ago
I am a humanist and a liberal. In the current technical paradigm, alignment to the user intent, as in, making the output's distribution aligned as closely as possible to the intended one, is an inextricable aspect of NLP capabilities a
81.
▲
by
airgapstopgap
3y ago
> my argument is that there is zero evidence whatsoever we will be able to prevent these from becoming dangerous Well no, but there's no need to prevent them from becoming dangerous inherently. They are tools, extensions of human ag
82.
▲
by
airgapstopgap
3y ago
> We're not talking here about someone's view on when white lies are justified or which model of marriage is the bestest - we're talking at the level of "cooperation = good", "love = good", "trust
83.
▲
by
airgapstopgap
3y ago
Other than snark, do you have a good argument? We know that technology can be error-prone, and LLMs fail in a great plethora of ways, but you are trying to sell an AI Doom narrative. I have never bought the idea that AI will be airgapped, b
84.
▲
by
airgapstopgap
3y ago
Kinzhal is simply an air-launched 9K720 Iskander, though, and it was always known to be interceptible because its trajectory is deterministic. You just need good radars. No, hypersonics are a legitimate new development, Russia just deceptiv
85.
▲
by
airgapstopgap
3y ago
> The only difference between GPT-3 and GPT-2 was scale. They didn't even change the tokenizer until 4. There were no "successful improvements" to smoothen the massive gap in capabilities between the two. "Other than
86.
▲
by
airgapstopgap
3y ago
Researchers aren't a monolith. Just because Gary Marcus (a complete fraud by the way, look up his "XProp" magitech he sold to Uber) pooh-poohed connectionist AI until ChatGPT came out doesn't mean nobody predicted gains
87.
▲
by
airgapstopgap
3y ago
The argument this quip points to is very clear, though. SBF tried to sell the regulatory entrenchment of his monopoly as a solution to the respective market's deep problems, which he purportedly was in a position to address. He was rev
88.
▲
by
airgapstopgap
3y ago
Gain of function. Yes, strong actors will have further access to AI, just as they have to everything else. I believe that on net, scaling properties in this domain are such that proliferation of AI democratizes the world rather than the oth
89.
▲
by
airgapstopgap
3y ago
OP: > As long as the AI (that anyone can access) can also spit out an equally powerful antiviral. You: > Yeah - that's not how that works I believe. Some problems are harder than others, and the optimal virus it could produce cou
90.
▲
by
airgapstopgap
3y ago
Who says there exists a way out of the regulatory-authoritarian attractor for AI? Who could've known that nuclear energy is a far lesser threat to humanity than climate change from burning fossils? Certainly not the legions of activist
More ›