9 ms·
'Attention is all you need' coauthor says he's 'sick' of transformers
- Xcelerate 1y agoHaha, I like to joke that we were on track for the singularity in 2024, but it stalled because the research time gap between "profitable" and "recursive self-improvement" was just a bit too long that we're now stranded on the transformer model for the next two decades until every last cent has been extracted from it.
- ai-christianson 1y agoThere's massive hardware and energy infra built out going on. None of that is specialized to run only transformers at this point, so wouldn't that create a huge incentive to find newer and better architectures to get the most out of all this hardware and energy infra?
- _lyxd 1y ago>None of that is specialized to run only transformers at this point isn't this what [etched](https://www.etched.com/ https://www.etched.com/) is doing?
- imtringued 1y agoOnly being able to run transformers is a silly concept, because attention consists of two matrix multiplications, which are the standard operation in feed forward and convolutional layers. Basically, you get transformers for free.
- kadushka 1y agodevil is in the details
- Davidzheng 1y agohow do you know we're not at recursive self-improvement but the rate is just slower than human-mediated improvement?
- teleforce 1y ago>The project, he said, was "very organic, bottom up," born from "talking over lunch or scrawling randomly on the whiteboard in the office." Many of the breakthrough and game changing inventions were done this way with the back of the envelope discussions, the other popular example was the Ethernet network. Some good stories of similar culture in AT&T Bell lab is well described in the Hamming's book [1]. [1] Stripe Press The Art of Doing Science and Engineering: https://press.stripe.com/the-art-of-doing-science-and-engineering https://press.stripe.com/the-art-of-doing-science-and-engine...
- atonse 1y agoTrue in creativity too. According to various stories pieced together, the ideas of 4 of Pixar’s early hits were conceived on or around one lunch. Bug’s Life, Wall-E, Monsters, Inc
- emi2k01 1y agoThe fourth one is Finding Nemo
- CaptainOfCoit 1y agoAll transformative inventions and innovations seems to come from similar scenarios like "I was playing around with these things" or "I just met X at lunch and we discussed ...". I'm wondering how big impact work from home will really have on humanity in general, when so many of our life changing discoveries comes from the odd chance of two specific people happening to be in the same place at some moment in time.
- DyslexicAtheist 1y agoI'd go back to the office in a heartbeat provided it was an actual office. And not an "open-office" layout, that people are forced to try to concentrate with all the noise and people passing behind them constantly. The agile treadmill (with PM's breathing down our necks) and features getting planned and delivered in 2 week-sprints, has also reduced our ability to just do something we feel needs getting done. Today you go to work to feed several layers of incompetent managers - there is no room for play, or for creativity. At least in most orgs I know. I think innovation (or even joy of being at work) needs more than just the office, or people, or a canteen, but an environment that supports it.
- Proofread0592 1y agoI think a transformer wrote this article, seeing a suspicious number of em dashes in the last section
- DonHopkins 1y agoThe next big AI architectural fad will be "disrupters".
- judge2020 1y agoMaybe even 'terminators'
- yieldcrv 1y agoThese are evolutionary dead ends, sorry that I'm not inspired enough to see it any other way, this transformer based direction is good enough The LLM stack has enough branches of evolution within it for efficiency, agent-based work can power a new industrial revolution specifically around white collar workers on its own, while expanding the self-expression for personal fulfillment for everyone else Well have fun sir
- password54321 1y ago^AI psychosis, never underestimate its effects. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/ https://metr.org/blog/2025-07-10-early-2025-ai-experienced-o...
- yieldcrv 1y agocoding assistants have nothing to do with what I’m talking about, you’re already 15 months behind what’s happening
- password54321 1y agoIs AGI in the room with us right now?
- yieldcrv 1y agodon’t need agi
- TheRealPomax 1y agotl;dr: AI is built on top of science done by people just "doing research", and transformers took off so hard that those same people now can't do any meaningful, real AI research anymore because everyone only wants to pay for "how to make this one single thing that everyone else is also doing, better" instead of being willing to fund research into literally anything else. It's like if someone invented the hamburger and every single food outlet decided to only serve hamburgers from that point on, only spending time and money on making the perfect hamburger, rather than spending time and effort on making great meals. Which sounds ludicrously far-fetched, but is exactly what happened here.
- jjtheblunt 1y agoGood points, and it made me have a mini epiphany... i think you analogously just described Sun Microsystems, where Unixes (BSD originally in their case, generalized to SVR4 (?) hybrid later) worked soooo well, that NT was built as a hybridization for the Microsoft user base and Apple reabsorbed the BSD-Mach-DisplayPostscript hybridization spinoff NeXT, while Linux simultaneously thrived.
- marcel-c13 1y agoDude now I want a hamburger :(
- TheRealPomax 1y agoDo it. I recommend a flame grilled angus with blue cheese.
- hatthew 1y agoThis is a decent analogy, but I think it understates how good transformers are. People are all making hamburgers because it's really hard to find anything better than a hamburger. Better foods definitely exist out there but nobody's been able to prove it yet.
- TheRealPomax 1y agoundersells (no estimating is being done here) But yes: the analogy is already hyperbole, and real life is even more hyperbolic. Transforms might work really well, but no one actually seems to know how to put them to real use, compared to generating billions-of-dollars in loss and burning the planet for it (In this analogy, we can take aim at the hamburger industry either birthing CAFO's, or putting them into hyper-overdrive, destroying the environment with orders of magnitude more CO2, etc. etc. it's a weirdly long-lasting analogy)
- amelius 1y agoOf course he's sick. He could have made billions.
- efskap 1y agoBut attention is all he needs.
- rzzzt 1y agoWhen you have your (next) lightbulb moment, how would you monetize such an idea? Royalties? 1c after each request?
- BoorishBears 1y agoLeave and raise a round right away.
- password54321 1y agoMoney has diminishing returns. Not everyone wants to buy Twitter.
- dekhn 1y agoThe way I look at transformers is: they have been one of the most fertile inventions in recent history. Originally released in 2017, in the subsequent 8 years they completely transformed (heh) multiple fields, and at least partially led to one Nobel prize. realistically, I think the valuable idea is probabilistic graphical models- of which transformers is an example- combining probability with sequences, or with trees and graphs- is likely to continue to be a valuable area for research exploration for the foreseeable future.
- jimbo808 1y agoWhich fields have they completely transformed? How was it before and how is it now? I won't pretend like it hasn't impacted my field, but I would say the impact is almost entirely negative.
- Profan 1y agohah well, transformative doesn't necessarily mean positive!
- econ 1y agoAll we get is distraction.
- dekhn 1y agoGenomics, protein structure prediction, various forms of small molecule and large molecule drug discovery.
- thesz 1y agoNo neural protein structure prediction papers I read have compared transformers to SAT solvers. As if this approach [1] does not exist. [1] https://pmc.ncbi.nlm.nih.gov/articles/PMC7197060/ https://pmc.ncbi.nlm.nih.gov/articles/PMC7197060/
- 1y ago
- bangaladore 1y ago> Now, as CTO and co-founder of Tokyo-based Sakana AI, Jones is explicitly abandoning his own creation. "I personally made a decision in the beginning of this year that I'm going to drastically reduce the amount of time that I spend on transformers," he said. "I'm explicitly now exploring and looking for the next big thing." So, this is really just a BS hype talk. This is just trying to get more funding and VCs.
- IncreasePosts 1y agoIt would be hype talk if he said and my next big thing is X.
- bangaladore 1y agoWell, that's why he needs funding. Hasn't figured out what the next big thing is.
- htrp 1y agoanyone know what they're trying to sell here?
- aydyn 1y agoprobably AI
- gwbas1c 1y agoThe ability to do original, academic research without the pressure to build something marketable.
- YeGoblynQueenne 1y agoScandalous.
- YC3498723984327 1y agoHis AI company is called "Fish AI"?? Does it mean their AI will have the intelligence of a fish?
- mmaunder 1y agoIf anyone has a video if it I think we'd all very much appreciate you posting a link. I've tried and I can't find one.
- InkCanon 1y agoThe other big missing part here is the enormous incentives (and punishments if you don't) to publish in the big three AI conferences. And because quantity is being rewarded far more than quantity, the meta is to do really shoddy and uninspired work really quickly. The people I talk to have a 3 month time horizon on their projects.
- nabla9 1y agoWhat "AI" means for most people is the software product they see, but only a part of it is the underlying machine learning model. Each foundation model receives additional training from thousands of humans, often very lowly paid, and then many prompts are used to fine-tune it all. It's 90% product development, not ML research. If you look at AI research papers, most of them are by people trying to earn a PhD so they can get a high-paying job. They demonstrate an ability to understand the current generation of AI and tweak it, they create content for their CVs. There is actual research going on, but it's tiny share of everything, does not look impressive because it's not a product, or a demo, but an experiment.
- janalsncm 1y agoI have a feeling there is more research being done on non-transformer based architectures now, not less. The tsunami of money pouring in to make the next chatbot powered CRM doesn’t care about that though, so it might seem to be less. I would also just fundamentally disagree with the assertion that a new architecture will be the solution. We need better methods to extract more value from the data that already exists. Ilya Sutskever talked about this recently. You shouldn’t need the whole internet to get to a decent baseline. And that new method may or may not use a transformer, I don’t think that is the problem.
- fritzo 1y agoIt looks like almost every AI researcher and lab who existed pre-2017 is now focused on transformers somehow. I agree the total number of researchers has increased, but I suspect the ratio has moved faster, so there are now fewer total non-transformer researchers.
- janalsncm 1y agoWell, we also still use wheels despite them being invented thousands of years ago. We have added tons of improvements on top though, just as transformers have. The fact that wheels perform poorly in mud doesn’t mean you throw out the concept of wheels. You add treads to grip the ground better. If you check the DeepSeek OCR paper it shows text based tokenization may be suboptimal. Also all of the MoE stuff, reasoning, and RLHF. The 2017 paper is pretty primitive compared to what we have now.
- marcel-c13 1y agoI think you misunderstood the article a bit by saying that the assertion is "that a new architecture will be the solution". That's not the assertion. It's simply a statement about the lack of balance between exploration and exploitation. And the desire to rebalance it. What's wrong with that?
- tim333 1y agoThe assertion, or maybe idea, that a new architecture may be the thing is kind of about building AGI rather than chatbots. Like humans think about things and learn which may require some differences from feed the internet in to pre-train your transformer.
- mcfry 1y agoSomething which I haven't been able to fully parse that perhaps someone has better insight into: aren't transformers inherently only capable of inductive reasoning? In order to actually progress to AGI, which is being promised at least as an eventuality, don't models have to be capable of deduction? Wouldn't that mean fundamentally changing the pipeline in some way? And no, tools are not deduction. They are useful patches for the lack of deduction. Models need to move beyond the domain of parsing existing information into existing ideas.
- hammock 1y agoThey can induct just can’t generate new ideas. Its not going to discover a new quark without a human in the loop somewhere
- nightshift1 1y agomaybe that's a good thing after all.
- eli_gottlieb 1y agoThat sounds like a category mistake to me. A proof assistant or logic-programming system performs deduction, and just strapping one of those to an LLM hasn't gotten us to "AGI".
- mcfry 1y agoA proof assistant is a verifier, and a tool so therefor a patch, so I really fail to see how that could be understood as the LLM having deduction.
- energy123 1y agoI don't see any reason to think that transformers are not capable of deductive reasoning. Stochasticity doesn't rule out that ability. It just means the model might be wrong in its deduction, just like humans are sometimes wrong.
- 1y ago
- wohoef 1y agoI'm tired of feeling like the articles I read are AI generated.
- stevetron 1y agoAnd here I thought this would be about Transformers: Robots in Disguise. The form of transformers I'm tired of hearing about.
- stevetron 1y agoAnd here I thought this would be about Transformers: Robots in Disguise. The form of transformers I'm tired of hearing about. And the decepticons.
- einrealist 1y agoI ask myself how much the focus of this industry on transformer models is informed by the ease of computation on GPUs/NPUs, and whether better AI technology is possible but would require much greater computing power on traditional hardware architectures. We depend so much on traditional computation architectures, it might be a real blinder. My brain doesn't need 500 Watts, at least I hope so.
- alyxya 1y agoI think people care too much about trying to innovate a new model architecture. Models are meant to create a compressed representation of its training data. Even if you came up with a more efficient compression, the capabilities of the model wouldn't be any better. What is more relevant is finding more efficient ways of training, like the shift to reinforcement learning these days.
- marcel-c13 1y agoBut isn't the max training efficiency naturally tied to the architecture? Meaning other architecture have another training efficiency landscape? I've said it somewhere else: It is not about "caring too much about new model architecture" but to have a balance between exploitation and exploration.
- alyxya 1y agoI didn't really convey my thoughts very well. I think of the actual valuable "more efficient ways of training" to be paradigm shifts between things like pretraining for learning raw knowledge, fine-tuning for making a model behave in certain ways, and reinforcement learning for learning from an environment. Those are all agnostic to the model architecture, and while there could be better model architectures that make pretraining 2x faster, it won't make pretraining replace the need for reinforcement learning. There isn't as much value in trying to explore this space compared to finding ways to train a model to be capable of something it wasn't before.
- nextworddev 1y agoIsn’t Sakana the one that got flack for falsely advertising its CUDA codegen abilities?
- deleted 1y ago[deleted]
- Mithriil 1y agoMy opinion on the "Attention is all you need" paper is that its most important idea is the Positional Encoding. The transformer head itself... is just another NN block among many.
- nashashmi 1y agoTransformers have sucked up all the attention and money. And AI scientists have been sucked in to the transformer-is-prime industry. We will spend more time in the space until we see bigger roadblocks. I really wished energy consumption was a very big roadblock that forced them into still researching.
- tim333 1y agoI think it may be a future roadblock quite soon. If you look at all the data centers planned and speed of it, it's going to be a job getting the energy. xAI hacked it by putting about 20 gas turbines around their data center which is giving locals health problems from the pollution. I imagine that sort of thing will be cracked down on.
- dmix 1y agoIf there’s a legit long term demand for energy the market will figure it out. I doubt that will be a long term issue. It’s just a short term one because of the gold rush. But innovation doesn’t have to happen overnight. The world doesn’t live or die on a subset of VC funds not 100xing within a certain timeframe Or it’s possible China just builds the power capabilities faster because they actually build new things
- tippytippytango 1y agoIt's difficult to do because of how well matched they are to the hardware we have. They were partially designed to solve the mismatch between RNNs and GPUs, and they are way too good at it. If you come up with something truly new, it's quite likely you have to influence hardware makers to help scale your idea. That makes any new idea fundamentally coupled to hardware, and that's the lesson we should be taking from this. Work on the idea as a simultaneous synthesis of hardware and software. But, it also means that fundamental change is measured in decade scales. I get the impulse to do something new, to be radically different and stand out, especially when everyone is obsessing over it, but we are going to be stuck with transformers for a while.
- danielmarkbruce 1y agoThis is backwards. Algorithms that can be parallelized are inherently superior, independent of the hardware. GPUs were built to take advantage of the superiority and handle all kinds of parallel algorithms well - graphics, scientific simulation, signal processing, some financial calculations, and on and on. There’s a reason so much engineering effort has gone into speculative execution, pipelining, multicore design etc - parallelism is universally good. Even when “computers” were human calculators, work was divided into independent chunks that could be done simultaneously. The efficiency comes from the math itself, not from the hardware it happens to run on. RNNs are not parallelizable by nature. Each step depends on the output of the previous one. Transformers removed that sequential bottleneck.
- Scene_Cast2 1y agoThere are large, large gaps of parallel stuff that GPUs can't do fast. Anything sparse (or even just shuffled) is one example. There are lots of architectures that are theoretically superior but aren't popular due to not being GPU friendly.
- danielmarkbruce 1y agoThat’s not a flaw in parallelism. The mathematical reality remains that independent operations scale better than sequential ones. Even if we were stuck with current CPU designs, transformers would have won out over RNNs. Unless you are pushing back on my comment "all kinds" - if so, I meant "all kinds" in the way someone might say "there are all kinds of animals in the forest", it just means "lots of types".
- vagab0nd 1y agoIt's pretty common I think. A thing is useful, but not intrinsically interesting (not anymore). So.. "let's move on"? Also, burn out is possible.
- stephc_int13 1y agoThe current AI race has created a huge sunken cost issue, if someone found a radically better architecture it could not only destroy a lot of value but also reset the race. I am not surprised that everyone is trying to make faster horses instead of combustion engines…
- krbaccord94f 1y agoTransformer architecture was the ideation of Google's "Deep Seek," which was licensed to enterprise AI for document processor integration.
- wrecked_em 1y agoYeah there's definitely more than meets the eye here.