3 ms·
Honestly, this is starting to make more and more sense. SOTA models are starting to converge to certain architecture and capabilities. I wouldn’t be surprised w
by syntaxing 2mo ago
Honestly, this is starting to make more and more sense. SOTA models are starting to converge to certain architecture and capabilities. I wouldn’t be surprised we end up with a base model ASIC + “fine tune” card where it’s a physical LoRA style adapter.
- smokel 2mo agoThe technical aspects of SOTA models are not publicly documented. How do you know if something is converging?
- cyanydeez 2mo agoif they were still exponentially increasing, they wouldn't be preparing for an IPO. IPO is where companies go to die and founders escape.
- _aavaa_ 2mo agoIf we had deepseek v4 flash 0731 etched on a chip it would be more than capable enough and fast enough for so many people's needs, even hardcore engineer.
- nurumaik 2mo agoWill be capable and fast enough for 2-3 weeks until new sota drops
- amazingamazing 2mo agoIf it is capable today why would a new model change this?
- FridgeSeal 2mo agoBecause new stuff instantly makes anything prior bad and incapable and garbage of course! Did you forget the hype-machine speaking notes??? /s
- catchnear4321 2mo agoif capability is a commodity then the differentiator becomes taste.
- thombles 2mo agoI think it’s tongue in cheek. When I first got access to Sonnet 4.5 I remember thinking to myself “y’know if they never got any better and I just had access to this forever then that would be pretty okay”. Turns out my expectations have changed since then and I would like a higher baseline now.
- singingtoday 2mo agoInteresting. I've yet to find a model I consider sufficiently intelligent. Fable is nice, but still requires a lot of guidance for large scope tasks.
- syntaxing 2mo agoSOTA American models are not. SOTA Chinese models are. From a physics aspect, closed source models cannot be too far from open source ones in terms of size. There’s only so much you can squeeze out a B100 style cluster even with fancy Dflash style diffusion model for the speculative model.
- cyanydeez 2mo agoI don't think there'll be a fine tune card; you'll have the base model vintage whatever year, and then your GPU will do whatever LoRA layers you want it to do; the LoRA will wrangle older dated models into the current of whatever your looking at. But yeah, for things like programming, if it can do linux and python and some go and sql and javascript, larger domains can be threaded with LORA
- VladVladikoff 2mo agoWouldn't this mean someone with sufficient hardware could lift the SOTA model weights off the chip? Or are you saying that these chips would only be used internally by these companies and not sold to the public?
- syntaxing 2mo agoI don’t get why this is an issue? You can run Claude/OpenAI SOTA models through Amazon bedrock. These weights have to live somewhere to run on Bedrock.
- wmf 2mo agosomewhere = an AWS data center with multiple layers of security and NDAs They won't sell/rent/license the weights to an end user at any price because they don't trust your security.
- syntaxing 2mo agoI work in embedded space. Just because it’s in hardware doesn’t mean you can’t “protect” it. Most modern software (regardless if it’s hardware or not) can be cryptophically signed.
- bluezly 2mo agoSigning protects authenticity and integrity, but it doesn’t really solve confidentiality. If the weights are physically encoded in hardware and the attacker owns the device, the problem becomes hardware extraction: decapping, probing, imaging, side channels, etc. You can make that very expensive, but it’s still a very different security model from keeping the weights in a datacenter.
- snek_case 2mo agoThe weights are very unlikely to be on the chip itself. That wouldn't work for SOTA models that are terabyte scale, even quantized. This is probably an accelerator for specific kernels in the model, but the weights are likely loaded from memory. The chip may have SRAM to store some of the weights temporarily during inference.
- encyclopedism 2mo agoImagine a multi-modal model with 1000's of tokens per second. Realtime inference for a host of applications. This is a BIG deal and will change the landscape in unfathomable ways. The https://chatjimmy.ai https://chatjimmy.ai demo was impressive. Once models settle down this makes sense. Imagine a cartridge with a physical model on it. You purchase a cartridge and stick it in your computer/phone/server. Want to upgrade? By a new 'cartridge'. This should bring inference cost down dramatically, I wonder how OpenAI/Anthropic feel about that.
- Grosvenor 2mo ago> Imagine a cartridge with a physical model on it. I can finally have my own Dixie flatline. Cool.
- mdp2021 2mo ago> Dixie Flatline In case some did not know: also the movie (actually TV series) is finally happening. # Neuromancer - Official Teaser ( https://news.ycombinator.com/item?id=49055037 https://news.ycombinator.com/item?id=49055037 )
- 2001zhaozhao 2mo agoi'm looking forward to Qwen3.8 27B launch to see how much models have peaked at a given size. it might already be time to start burning the best small models onto hardware since it's possible they can't get much better at many tasks like knowledge recall due to the inherent information density limits for models at a given size.
- anthonypasq 2mo agovery interesting idea. i didnt think of that. i was just assuming youd have an additional one of these in your phone for actual lightning fast local inference
- pstuart 2mo agoThe cartridge could be a small mac-mini type unit connected and powered over thunderbolt. If it included like an m5 or m7 with 64GB of memory and a PCIe5/6 4TB Nvme it would be amazeballs. Hopefully when the bubble corrects and hardware advances and prices reset something like that will become available. Just even comparing compute from 10 years ago (Apple silicon vs Intel) and it's significant. 20 years it gets crazy. My first computer was an 8 bit 6502 with 64K RAM and a 128K floppy drive (I think, it's fuzzy). Everything amazing now will look quaint in due time.
- kevin_thibedeau 2mo agoThen we can have machine psychologists pull cards when they run amok.
- all2 2mo agoYou have a robot. You need it to be smarter. You buy a new model cartridge (probably a PCIE 9.x). Now you need some domain specific skills. You'd like it to be able to cook, and you'd like it to not dent your walls anymore. You buy 'improved spatial reasoning LORA' card and 'Gordon Ramsey's Chef ULTRA9000' card. Now your robot can respond sarcastically when you ask for chicken nuggets. Again. It also doesn't dent your walls anymore.
- walrus01 2mo agoHaving a base model ASIC as a physical piece of hardware makes me think of the early days of microcomputer desktop stuff where having a socketed ROM or PROM was a key piece of hardware, and people actually knew/cared what ROM was on their system's motherboard. Imagine if like instead of having a specific Mac Plus ROM, you had a thing that looks like a fat ASIC that can hold models sitting on a slotted daughtercard directly next to the CPU and RAM.
- breadislove 2mo agowe have not converged at all, if you look at how different the chinese models in terms of architecture you can guess that the labs are experimenting a lot as well. we are seeing all different types of hybrid architectures, different attention methods and so on. Of course on a high level its still a transformer but if you take a proper look we are seeing more divergence then a convergence.