4 ms·
> It would be foolish to deny the effectiveness of image recognizers based on generalized adversarial networks (GANs), the key neural network technology underly
by ainch 2mo ago
> It would be foolish to deny the effectiveness of image recognizers based on generalized adversarial networks (GANs), the key neural network technology underlying LLMs
I could be misreading this, but I hope the author doesn't think GANs are used in LLMs. They are cool, though.
- chpatrick 2mo agoYeah, it's hard to take the rest seriously.
- pasquinelli 2mo agomaybe you're looking for an excuse to ignore what's being said.
- llm_nerd 2mo ago[flagged]
- camillomiller 2mo ago[flagged]
- sigilsack 2mo ago[dead]
- llm_nerd 2mo ago[flagged]
- dgellow 2mo agoYou‘re both contributing negatively to the discourse by doing that sport team fighting…
- camillomiller 2mo agoOf course, quality of the discussion above all else, including destroying the world and the fabric of society via a financial grifting scheme that is commoditizing free knowledge.
- deleted 2mo ago[deleted]
- scarmig 2mo agoThings that are vaguely GAN-shaped might play some role in post training. Though no one would call them GANs or identify them as the key underlying technology.
- bonoboTP 2mo agoThere is RLHF with a model that's trained to emulate a human evaluator, but the evaluator is not really trained jointly with the main model to adapt to its distribution and tell it from real text. Though I'm sure there are some niche cases when this is done. But definitely not a prominent thing.
- threethirtytwo 2mo agoHe hallucinated. This is how you verify non AI nowadays. AI is so good that if you see an hallucination as obviously wrong like this one it’s a sign it’s written by a human.
- sxp 2mo agoYeah. Ironically, if he used an LLM to proofread his work, it would have told him that GAN's are generative not "generalized". That LLMs are primarily built on Transformers rather than GANs. And that they're more famous for image generation rather than image recognition.
- FeepingCreature 2mo agoAnd that GANs haven't been relevant in image generation since 2021.
- bonoboTP 2mo agoTransformers are an architecture and GANs are a training method for the architecture. There are GAN Transformers.
- nullstyle 2mo agoExamples for us less educated folk?
- Alpha3031 2mo agoJiang et al. (2021) TransGAN: https://dl.acm.org/doi/10.5555/3540261.3541391 https://dl.acm.org/doi/10.5555/3540261.3541391 In general probably not much of a stretch to get rid of convolutions or recurrence by replacing with attention and see if it works hence the title of the original transformers paper.
- bonoboTP 2mo agoAnalogy "it's not GAN, it's transformer!" - - "it's not a recursive implementation, it's object oriented!" if you are more familiar with CS. Or "it's not 4-wheel-drive, it's diesel!", if familiar with cars.
- andy99 2mo agoIt’s unfortunate - he’s an author, it would have been fine to stick with an authors perspective that LLMs can’t write (which is true), as well as the copyright stuff (which I don’t agree with but he certainly has standing to give an opinion on). But he’s made the error of trying to come at it from a technical perspective, when he clearly knows nothing about that side of things, which discredits the rest.
- magicalist 2mo ago> But he’s made the error of trying to come at it from a technical perspective Huh? That quote appears to be the entirety of the "technical perspective" of the post and is an aside from his larger points that you have blessed as "fine". Literally nothing in the rest of the post relies on that incorrect statement. Let me quibble with what is discredited here, given the entirety of your point is built upon an error.
- red75prime 2mo agoHere's another technical point: > They're word-association mechanisms with no embodiment and no way to associate the text vectors they manipulate with real-world phenomena. This is wrong too. RLVR grounds foundational models in reality.
- Alpha3031 2mo agoAren't RLVR signals typically based off formal systems not natural phenomena? The formal sciences are certainly useful for producing tools used in natural science, but it's not entirely clear they alone are sufficient to associate text to natural (real-world) phenomena. Honestly, the pretraining and RLHF are probably more tied to the world than RLVR, for all that RLVR might be useful (maybe even more useful) for making the model perform better in certain tasks (such as working with formally defined systems).
- azakai 2mo agoIf you want a more concrete example, then LLMs are also trained on visual data these days, which means they do have access to the world in an important way. This directly contradicts the blogpost's claim that LLMs have > no way to associate the text vectors they manipulate with real-world phenomena. Historically, that LLMs were text-only used to be a major argument for why they "lack access to meaning", see the Stochastic Parrot paper and the Octopus paper that it references. But even the authors of those papers have (grudgingly) conceded that the argument no longer holds due to multimodality.