11 ms·
Local AI is driving the biggest change in laptops in decades
- deleted 10mo ago[deleted]
- aappleby 10mo agoI predict we will see compute-in-flash before we see cheap laptops with 128+ gigs of ram.
- wkat4242 10mo agoYeah especially since what is happening in the memory market
- noosphr 10mo agoFeast and famine. In three years we will be swimming in more ram than we know what to do with.
- fallat 10mo agoKind of feel that's already the case today... 4GB I find is still plenty for even business workloads.
- autoexec 10mo agoVideo games have driven the need for hardware more than office work. Sadly games are already being scaled back and more time is being spent on optimization instead of content since consumers can't be expected to have the kind of RAM available they normally would and everyone will be forced to make do with whatever RAM they have for a long time.
- znpy 10mo agoThat might not be the case. The kind of memory that will flood the second-hand market could not be the kind of memory we can stuff in laptops or even desktop systems.
- deleted 9mo ago[deleted]
- p1esk 10mo agoWe’ve had “compute in flash” for a few years now: https://mythic.ai/product/ https://mythic.ai/product/
- aitchnyu 10mo agoMemristors are (IME) missing from the news. They promised to act as both persistent storage and fast RAM.
- ACCount37 9mo agoIf only memristors weren't vaporware that has "shown promise" for 3 decades now and went nowhere.
- zamadatix 10mo agoI can't tell if this is optimism for compute-in-flash or pessimism with how RAM has been going lately!
- znpy 10mo agoYou could get 128gb ram laptops from the time ddr4 came around: workstation class laptops with 4 ram slots would happily take 128gb of memory. The fact that nowadays there are little to no laptops with 4 ran slots is entirely artificial.
- mhitza 9mo agoI was mussing this summer if I should get a refurbed Thinkpad P16 with 96GB of RAM to run VMs purely in memory. Now that 96GB of ram cost as much as a second P16.
- znpy 9mo agoI feel you, so much. I was thinking of getting a second 64gb node for my homelab and i thought i’d save those money… now the ram alone cost as much as the node, and I’m crying. Lesson learned: you should always listen to that voice inside your head that say: “but i need it…” lol
- pluralmonad 9mo agoI rebuilt a workstation after a failed motherboard a year ago. I was not very excited about being forced to replace it on a days notice and cheaped out on the RAM (only got 32GB). This is like the third or fourth time I've taught myself the lesson to not pinch pennies when buying equipment/infrastructure assets. It's the second time the lesson was about RAM, so clearly I'm a slow learner.
- znpy 9mo agoYou probably had paged out the lesson to slow storage… you should get more ram :p
- 112233 10mo agoBy "we" do you mean consumers? No, "we" will get neither. This is unexpected, irresistable opportunity to create a new class, by controlling the technology that people are required and are desiring to use (large genAI) with a comprehensive moat — financial, legislative and technological. Why make affordable devices that enable at least partial autonomy? Of course the focus will be on better remote operation (networking, on-device secure computation, advancing narrative that equates local computation with extremism and sociopathy).
- cmxch 9mo agoPush Washington to grill the foundries and their customers. Repeat until prices drop.
- 14113 9mo agoThere was a company that did compute-in-dram, which was recently acquired by Qualcomm: https://www.emergentmind.com/topics/upmem-pim-system https://www.emergentmind.com/topics/upmem-pim-system
- ajb 9mo agoThe thing that is supposed to happen next is high-bandwidth flash. In theory, it could allow laptops to run the larger models without being extortionately costly, by loading directly from flash into the GPU (not by executing in flash) But I haven't seen figures of the actual bandwidth yet, and no doubt to start with it will be expensive. The underlying technology of flash has much higher read latency than dram, so it's not really clear (to me, at least) if they can deliver the speeds needed to remove the need to cache in VRAM just by increasing parallelism.
- deleted 9mo ago[deleted]
- wkat4242 10mo agoThis article is so dumb. It totally ignores the memory price explosion that will make large fast memory laptops unfeasible for years and states stuff like this: > How many TOPS do you need to run state-of-the-art models with hundreds of millions of parameters? No one knows exactly. It’s not possible to run these models on today’s consumer hardware, so real-world tests just can’t be done. We know exactly the performance needed for a given responsiveness. TOPS is just a measurement independent from the type of hardware it runs on.. The less TOPS the slower the model runs so the user experience suffers. Memory bandwidth and latency plays a huge role too. And context, increase context and the LLM becomes much slower. We don't need to wait for consumer hardware until we know much much is needed. We can calculate that for given situations. It also pretends small models are not useful at all. I think the massive cloud investments will put pressure away from local AI unfortunately. That trend makes local memory expensive and all those cloud billions have to be made back so all the vendors are pushing for their cloud subscriptions. I'm sure some functions will be local but the brunt of it will be cloud, sadly.
- vegabook 10mo agoalso, state of the art models have hundreds of _billions_ of parameters.
- omneity 10mo agoIt tells you about their ambitions..
- layer8 9mo agoThe article is from mid-November (and probably was written even earlier), where the RAM price explosion wasn’t as striking yet.
- dcreater 9mo agoHorrible article. Low effort, low knowledge. Had no idea the bar was so low for an IEEE publication
- esses 10mo agoI spent a good 30 seconds trying to figure out what DDS was an acronym for in this context.
- seanmcdirmid 10mo agoI’ve been running LLMs on my laptop (M3 Max 64GB) for a year now and I think they are ready, especially with how good mid sized models are getting. I’m pretty sure unified memory and energy efficient GPUs will be more than just a thing on Apple laptops in the next few years.
- allovertheworld 10mo agoOnly because of Apples unified memory architecture. The groundwork is there, we just need memory to be cheaper so we can fit 512+GB now ;)
- seanmcdirmid 10mo agoMemory prices will rise short term and generally fall long term, even with the current supply hiccup the answer is to just build out more capacity (which will happen if there is healthy competition). I meant, I expect the other mobile chip providers to adopt unified architecture and beefy GPU cores on chip and lots of bandwidth to connect it to memory (at the max or ultra level, at least), I think AMD is already doing UM at least?
- spwa4 9mo ago> Memory prices will rise short term and generally fall long term, even with the current supply hiccup the answer is to just build out more capacity (which will happen if there is healthy competition) Don't worry! Sam Altman is on it. Making sure there never is healthy competition that is. https://www.mooreslawisdead.com/post/sam-altman-s-dirty-dram-deal https://www.mooreslawisdead.com/post/sam-altman-s-dirty-dram...
- seanmcdirmid 9mo agoWe’ve been through multiple cycles of scarcity/surplus DRAM cycles in the last couple of decades. Why do we think it will be different now?
- Morromist 10mo agoI was in the market for a laptop this month. Many new laptops now advertise AI features like this "HP OmniBook 5 Next Gen AI PC" which advertises: "SNAPDRAGON X PLUS PROCESSOR - Achieve more everyday with responsive performance for seamless multitasking with AI tools that enhance productivity and connectivity while providing long battery life" I don't want this garbage on my laptop, especially when its running of its battery! Running AI on your laptop is like playing Starcraft Remastered on the Xbox or Factorio on your steamdeck. I hear you can play DOOM on a pregnancy test too. Sure, you can, but its just going to be a tedious inferior experiance. Really, this is just a fine example of how overhyped AI is right now.
- Legend2440 10mo agoLaptop manufacturers are too desperate to cash on the AI craze. There's nothing special about an 'AI PC'. It's just a regular PC with Windows Copilot... which is a standard Windows feature anyway. >I don't want this garbage on my laptop, especially when its running of its battery! The one bit of good news is it's not going to impact your battery life because it doesn't do any on-device processing. It's just calling an LLM in the cloud.
- bitwize 10mo agoAI PCs also have NPUs which I guess provide accelerated matmuls, albeit less accelerated than a good discrete GPU.
- autoexec 10mo agoEven collecting and sending all that data to the cloud is going to drain battery life. I'd really rather my devices only do what I ask them to than have AI running the background all the time trying to be helpful or just silently collecting data.
- sandworm101 10mo ago>> I'd really rather my devices only do what I ask them to Linux hears your cry. You have a choice. Make it.
- bfrog 10mo agoI suppose it depends on the model, code was useless. As a lossy copy of an interactive Wikipedia it could be ok not good or great just ok. Maybe for creative suggestions and editing it’d be ok.
- socketcluster 10mo agoI feel like there's no point to get a graphics card nowadays. Clearly, graphics cards are optimized for graphics; they just happened to be good for AI but based on the increased significance of AI, I'd be surprised if we don't get more specialized chips and specialized machines just for LLMs. One for LLMs, a different one for stable diffusion. With graphics processing, you need a lot of bandwidth to get stuff in and out of the graphics card for rendering on a high-resolution screen, lots of pixels, lots of refreshes, lots of bandwidth... With LLMs, a relatively small amount of text goes in and a relatively small amount of text comes out over a reasonably long amount of time. The amount of internal processing is huge relative to the size of input and output. I think NVIDIA and a few other companies already started going down that route. But probably graphics cards will still be useful for stable diffusion; especially AI-generated videos as the inputs and output bandwidth is much higher.
- Legend2440 10mo agoLLMs are enormously bandwidth hungry. You have to shuffle your 800GB neural network in and out of memory for every token, which can take more time/energy than actually doing the matrix multiplies. GPUs are almost not high bandwidth enough.
- Zambyte 10mo agoThis doesn't seem right. Where is it shuffling to and from? My drives aren't fast enough to load the model every token that fast, and I don't have enough system memory to unload models to.
- smallerize 10mo agoYou're probably not using an 800GB model.
- p1esk 10mo agoIt is right. The shuffling is from CPU memory to GPU memory, and from GPU memory to GPU. If you don’t have enough memory you can’t run the model.
- fwipsy 10mo agoSeems like wishful thinking. > How many TOPS do you need to run state-of-the-art models with hundreds of millions of parameters? No one knows exactly. Why not extrapolate from open-source AIs which are available? The most powerful open-source AI (which I know of) is Kimi K2 and >600gb. Running this at acceptable speed requires 600+gb GPU/NPU memory. Even $2000-3000 AI-focused PCs like the DGX spark or Strix Halo typically top out at 128gb. Frontier models will only run on something that costs many times a typical consumer PC, and only going to get worse with RAM pricing. In 2010 the typical consumer PC had 2-4gb of RAM. Now the typical PC has 12-16gb. This suggests RAM size doubling perhaps every 5 years at best. If that's the case, we're 25-30 years away from the typical PC having enough RAM to run Kimi K2. But the typical user will never need that much RAM for basic web browsing, etc. The typical computer RAM size is not going to keep growing indefinitely. What about cheaper models? It may be possible to run a "good enough" model on consumer hardware eventually. But I suspect that for at least 10-15 years, typical consumers (HN readers may not be typical!) will prefer capability, cheapness, and especially reliability (not making mistakes) over being able to run the model locally. (Yes AI datacenters are being subsidized by investors; but they will remain cheaper, even if that ends, due to economies of scale.) The economics dictate that AI PCs are going to remain a niche product, similar to gaming PCs. Useful AI capability is just too expensive to add to every PC by default. It's like saying flying is so important, everyone should own an airplane. For at least a decade, likely two, it's just not cost-effective.
- sipjca 10mo ago> It may be possible to run a "good enough" model on consumer hardware eventually 10-15 years?!!!! What is the definition of good enough? Qwen3 8B or A30B are quite capable models which run on a lot of hardware even today. SOTA is not just getting bigger, it's also getting more intelligence and running it more efficiently. There have been massive gains in intelligence at the smaller model sizes. It is just highly task dependent. Arguably some of these models are "good enough" already, and the level of intelligence and instruction following is much better from even 1 year ago. Sure not Opus 4.5 level, but still much could be done without that level of intelligence.
- 9mo ago
- gguncth 10mo agoI have no desire to run an LLM on my laptop when I can run one on a computer the size of six football fields.
- sandworm101 10mo agoI've been playing around with my own home-built AI server for a couple months now. It is so much better than using a cloud provider. It is the difference between drag racing in your own car, and renting one from a dealership. You are going to learn far more doing things yourself. Your tools will be much more consistent and you will walk away with a far greater understanding of every process. A basic last-generation PC with something like a 3060ti (12GB) is more than enough to get started. My current rig pulls less than 500w with two cards (3060+5060). And, given the current temperature outside, the rig helps heat my home. So I am not contributing to global warming, water consumption, or any other datacenter-related environmental evil.
- HelloUsername 9mo ago> I am not contributing to global warming lol
- DamonHD 9mo agoUnless you normally use electric resistance heating (or some kind of fossil fuel with higher gCO2/kWh) then you don't get necessarily a free pass on the global warming thing! Our whole home is heated with <500W on average: at this moment the heat pump is drawing 501W (H4 boundary) at close to freezing outside, and its demand is intermittent.
- theshrike79 9mo agoThe point is that when you run it on your own hardware you can feed the model your health data, bank statements and private journals and can be 5000% sure they’re not going anywhere
- dboreham 9mo agoRegular people don't understand nor care about any of that. They'll happily take the Faustian bargain.
- j45 10mo agoThis must be referring mostly to windows, or non-Apple laptops
- spullara 10mo agoI'm running GPT-OSS 120B on a MacBook Pro M3 Max w/128 GB. It is pretty good, not great, but better than nothing when the wifi on the plane basically doesn't work.
- scotty79 9mo agoI'm running it on PC laptop with mobile 5090 and 64GB of ram. Start is a bit rough, but once it gets going it is perfectly servicable when I'm on a bad connection.
- juancn 9mo agoThe price of RAM is going to throw a wrench at that
- mattas 9mo agoSee: "3D TVs are driving the biggest change in TVs in decades"
- eleventyseven 9mo agoA lazy easy cheap shot. But do you deny these aspects from the article are not coming? Or won't be still here in 5 years? - Addition of more—and faster—memory. - Consolidation of memory. - Combination of chips on the same silicon. All of these are also happening for non AI reasons. The move to SoC that really started with the M1 wasn't because of AI, but unified memory being the default is something we will see in 5 years. Unlike 3D TV.
- blibble 9mo ago> Addition of more—and faster—memory. probably not after scam altman bought up half the world's supply for his shit company
- estimator7292 9mo agoMemory is absolutely not coming in the near future. Nobody can afford it.
- ToucanLoucan 9mo agoIn order: - People wanting more memory is not a novel feature. I am excited to find out how many people immediately want to disable the AI nonsense to free up memory for things they actually want to do. - Same answer. - I think the drive towards SOCs has been happening already. Apple's M-series utterly demolishes every PC chip apart from the absolute bleeding-edge available, includes dedicated memory and processors for ML tasks, and it's mature technology. Been there for years. To the extent PC makers are chasing this, I would say it's far more in response to that than anything to do with AI.
- MisterTea 9mo ago> The move to SoC that really started with the M1 No it did not. There were numerous SoC that came before it and was inevitable in this space.
- seunosewa 9mo ago"How many TOPS do you need to run state-of-the-art models with hundreds of millions of parameters? No one knows exactly." What's he talking about? It's trivial to calculate that.
- fny 9mo agoIt's also been done before...[0] [0]: https://www.edge-ai-vision.com/2024/05/2024-edge-ai-and-vision-product-of-the-year-award-winner-showcase-qualcomm-edge-ai-processors https://www.edge-ai-vision.com/2024/05/2024-edge-ai-and-visi...
- RobotToaster 9mo agoIsn't the ability to run it more dependant on (V)RAM? With TOPS just dictating the speed at which it runs?
- zozbot234 9mo agoStrictly speaking, you don't need that much VRAM or even plain old RAM - just enough to store your context and model activations. It's just that as you run with less and less (V)RAM you'll start to bottleneck on things like SSD transfer bandwidth and your inference speed goes down to a crawl. But even that may or may not be an issue depending on your exact requirements: perhaps you don't need your answer instantly and can wait while it gets computed in the background. Or maybe you're running with the latest PCIe 5 storage which overall gives you comparable bandwidth to something like DDR3/DDR4 memory.
- NitpickLawyer 9mo agoA good rule of thumb is that PP (Prompt Processing) is compute bound while TG (Token Generation) is (V)RAM speed bound.
- swyx 9mo ago> state-of-the-art models > hundreds of millions of parameters lol lmao, even
- cramcgrab 9mo ago
- tehjoker 9mo agoI mean, having a more powerful laptop is great, but at the same time, these guys are calling for a >10x increase in RAM and a far more powerful NPU. How will this affect pricing? How will it affect power management? It made it seem like most of the laptop will be dedicated to gen AI services, which I'm still not entirely convinced are quite THAT useful. I still want a cheap laptop that lasts all day and I also want to be able to tap that device's full power for heavy compute jobs!
- superkuh 9mo agoThe problem with this is that NPU have terrible, terrible support in the various software ecosystems because they are unique to their particular soc or whatever. No consistency even within particular companies.
- tracerbulletx 9mo agoThis mostly just shows you how far behind the M1 (which came out 5 years ago) all the non Apple laptops are.
- properbrew 9mo agoWas never really into Apple hardware (mainly the price), however I recently got an M1 Mac Mini and an iPhone for app development, and the inference speed for as you say, a 5 year old chip is actually crazy. If they made the M series fully open for Linux (I know Asahi is working away) I probably would never buy another non-M series processor again.
- dpedu 9mo agoI got an M1 Mac Mini somewhat recently as well, to replace my ~2012 Mac Mini that I use as a media center PC. And frankly, it's overkill. Used ones can be had for $200-$300 USD, lower side with cosmetic damage. An absolute steal, IMO.
- bnolsen 9mo agoWork gave me an m1 pro with 32gb on it. A year ago I put together one of those minisforum board+laptop apu with 64gb ram and 2tb nvme for not much money at the time, likely 500usd. For the performance sensitive software I was working on the 7935hs ran with about 50x more throughout using compilers with llvm backend.
- jeffbee 9mo agoYou can still get an M1 Macbook Air at retail for $599 ($300 for refurbs), which is a Chromebook price for a laptop that is better in pretty much every respect than any Chromebook.
- nyarlathotep_ 9mo agohttps://slickdeals.net/f/19004236-select-micro-center-stores-apple-mac-mini-desktop-w-m4-chip-16gb-ram-256gb-ssd-400-free-store-pickup https://slickdeals.net/f/19004236-select-micro-center-stores... MicroCenter has(had? OOS near me) M4 Minis for $400! A remarkable bargain, even more so considering the recent hardware price hikes.
- zkmon 9mo agoYou don't understand the needs of a common laptop user. Define the usecases that require reaching out to laptop instead of using the phone that is nearby. Those usecases don't need LLM for a common laptop user.
- TrackerFF 9mo agoWith the wild ram prices, which btw are probably going to last out 2026, I expect 8 GB ram to be the new standard going on forward. 32 GB ram will be for enthusiasts with deep pockets, and professionals. Anything over that, exclusively professionals. The conspiracy theorist inside me is telling me that big AI companies like OpenAI would rather see that people are using their puny laptops as terminals / shells only, to reach sky-based models, than to let them have beefy laptops and local models.
- cmxch 9mo agoNot if a few investigations into the foundries and their datacenter deals stops that.
- andy99 9mo agoThe conspiracy theorist inside me is telling me that big AI companies... I don’t believe in conspiracies but I do believe in incentives sometimes lining up. Now that there is a RAM heavy cloud application, cloud providers are suddenly in direct competition with consumers for scarce resources, with the winner being able to control where people run their models.
- jwr 9mo agoThe author seems unaware of how well recent Apple laptops run LLMs. This is puzzling and puts into question the validity of anything in this article.
- fancyfredbot 9mo agoI think the author is aware of Apple silicon. The article mentions the fact Apple has unified memory and that this is advantageous for running LLMs.
- dangus 9mo agoThen idk why they say that most laptops are bad at running LLMs, Apple has a huge marketshare in the laptop market and even their cheapest laptops are capable in that realm. And their PC competitors are more likely to be generously specced out in terms of included memory. > However, for the average laptop that’s over a year old, the number of useful AI models you can run locally on your PC is close to zero. This straight up isn’t true.
- DANmode 9mo ago> Apple has a huge marketshare in the laptop market Hello, from outside of California!
- dangus 9mo agoGlobal Mac marketshare is actually higher than the US: https://www.mactech.com/2025/03/18/the-mac-now-has-14-8-of-the-u-s-pc-market-and-17-1-of-the-global-pc-market/ https://www.mactech.com/2025/03/18/the-mac-now-has-14-8-of-t...
- DANmode 9mo agoLess than 1 in 5 doesn’t feel like huge market share, but it’s more than I have!
- meisel 9mo agoI think only a small percentage of users care that much about running LLMs locally to pay for extra hardware for it, put up with slower and lower-quality responses, etc. . It’ll never be as good as non-local offerings, and is more hassle.
- tengbretson 9mo agoOutside of Apple laptops (and arguably the Ryzen AI MAX 390), an "AI ready" laptop is simply marketing speak for "is capable of making HTTP requests."
- gamblor956 9mo agoThe "AI laptop" boom is already fading. It turns out that LLMs, local or otherwise, just aren't very useful. Like Big Data, LLMs are useful in a small niche of areas, like poorly summarizing meeting notes, or grammar check at a middle-school level. On LLMs for coding tasks: I asked a programmer why they loved Claude and he showed me the output. Twenty years ago, that kind of code would have gotten someone PIP'd. Today it's considered better than most junior programmers...which is a sign of how far programming standards have fallen, and explains why most programs and apps are such buggy pieces of sh$t these days.
- bad_haircut72 9mo agoMy recent shower thought was the idea that Moores law hasnt slowed at all, we just went multi-core. Its crazy that the intel folks were so interested in optimizing for single thread CPU design they completely misunderstood where the best effort would be spent - if I had been around back then (speaking as an Elixir dev) I would have been way more interested in having 500 theead CPUs than getting down to nanometer scale dies. Thats what you get when everyone on the team is a bunch of C programmers
- ip26 9mo agoBefore LLMs, the use of parallelism on your typical laptop was limited to application level parallelism, e.g. one thread for Outlook and one for each tab in Chrome.
- astrange 9mo agoIntel designed a super high threaded CPU like that, Knightsbridge. It was useless. Single threaded programs are good.
- kristianp 9mo ago"Local AI" could be many different things. NPUs are too puny to run many recent models, such as image generation and llms. The article seems to gloss over many important details like this, for example the creative agency, what AI work are they doing? > marketing firm Aigency Amsterdam, told me earlier this year that although she prefers macOS, her agency doesn’t use Mac computers for AI work.
- Groxx 9mo agore NPUs: they've been a marketing thing for years now, but I really have no idea how many of them are actually used when you run [whatever]. particularly after a year or two of software updates. anyone have numbers? are they just an added expense that is supported for first party stuff for 6 months before they need a bigger model, or do they have staying power? clearly they are capable of being used to save power, but does anything do that in practice, in consumer hardware?
- 0xbadcafebee 9mo agoWirth's Law in action. Eventually it's going to take an entire datacenter to read the news.
- xrd 9mo agoThe takeaway from these comments are that you can really run local models if you use m-series devices from apple. But, can you do that if you install Linux on that hardware? I hate to admit apple hardware is incredible. But, I can't say the same about macos anymore. Can I run Linux and reap the benefits of m-series chips with local inference? Or, are there any alternatives where I can use llms on Linux on a laptop?
- darkreader 9mo ago[dead]
- suprjami 9mo agoExtremely cringe article. The biggest thing to affect laptops in "decades" is solid state storage. No longer do you need to worry about killing your entire device simply by putting it down on a solid surface. There are also plenty of other things like modern dense lithium ion batteries with 12+ hour runtimes, super power friendly CPUs of all architectures, the ultra-thin body and metal body popularised by Apple, LCD panels without ghosting, external power bricks instead of literally a PC power supply in a briefcase. But yeah sure, the infinite slop plagiarism machine is coming. Gotta get some clicks!
- ge96 9mo agoWonder if this relates to/overlaps those Coral Accelerator devices.
- openquery 9mo agoFor 99% of people I don't see the usecase (except for privacy but that ship sailed a decade ago for the aforementioned 99%). If the argument is inference offline - the modern computing experience is basically all done through the browser anyway so I don't buy it. GPUs for video games where you need low latency makes sense. Nvidia GeForce Now works but not for any serious gaming. But when it comes to LLMs at least, the 100ms latency between you and the Gemini API or whichever provider you use is negligible compared to the inference time. What am I missing?
- mginszt 9mo agoI'm sure giants like Microsoft would like to add more AI capabilities, and I'm also sure they would like to avoid running them on their own servers. Another thing is that I wouldn’t expect LLMs to be free forever. One day, CEOs will decide that everyone has become accustomed to them - and that will be the first day of a subscription-based model and the last day of AI companies reporting financial losses.
- chnmig 9mo agoThe power and resource consumption of local large models are problems that laptops have to solve, and new versions of models are constantly being released, which means that laptop configurations will soon become outdated.
- rldjbpin 9mo agoif you focus out of local LLMs (also served using dedicated apps), the title holds a lot of promise. case in point: WASM and WebGPU the edge/on-device AI use cases on smartphones can also extend without user friction through web apps built on the above standards. perhaps one day there will be a "WebNPU" or just get supported through existing standards. there are already some use cases on apps but it usually fallbacks on cpu. perhaps it could be the hw accelerated moment that we saw with video on the web.
- rcarmo 9mo agoKind of ironic that it is a factor that 95% of regular users don’t care about or actively avoid.