7 ms·
The Prompt API
- fg137 6mo ago"sorry, to use our website, you must have at least 22 GB of free disk space."
- cdrini 6mo agoTrue, but arguably better than "sorry, to use our website, you must have a ChatGPT subscription."
- _pdp_ 6mo agothat is ~9% of the total available disk space for baseline phones and laptops for a model that is not that useful.
- jfoster 6mo agoAlso much better than every website wanting its own 22 GB rather than the 22 GB being a shared resource.
- fg137 6mo agoI would very much like not to have to download 22 GB for some inference capability that is way worse than API calls both in terms of quality and speed. I would rather pay money than seeing this thing running in my browser that only prints 5 tps on high-end consumer hardware.
- jfoster 5mo agoWhy are you pretending those are the options? The options are: 1. 22GB per website 2. 22GB per browser 3. 0GB / No AI capabilities By having this in Chrome they are simply ensuring that option 2 replaces option 1. You still have option 3.
- fg137 5mo agoNo. The real options are 1. No AI 2. AI that works and is actually useful 3. AI that is slow, crappy and hallucinates all the time I choose 1 and 2.
- jfoster 5mo agoFair, but actually you'd surely want your choice of those three, right? And what's being discussed here is what the better implementation of option 3 is. My point is that if you're going with one of the possible implementations of option 3, then 22GB per browser is objectively a lot better than 22GB per website.
- fg137 6mo agoMore like "you need to sign up for our website and pay for a subscription", and I'd much rather do that if it's actually providing value. I am absolutely not going to run model locally which slowly churns out words at 5 tps while making the computer hot to touch.
- skybrian 6mo agoStill in origin trial? Looks like they're adding a temperature parameter: https://chromestatus.com/feature/6325545693478912 https://chromestatus.com/feature/6325545693478912
- kbx 5mo agoIt's on track to ship in Chrome 148: https://chromestatus.com/feature/5134603979063296 https://chromestatus.com/feature/5134603979063296 The parameters are not part of this initial release but can be added back with the origin trial you discovered.
- avaer 6mo agoIt works, I've shipped this as a "local inference"/poor person's ollama for low-end llm tasks like search. The main win is that it's free and privacy preserving, and (mostly) transparent to users in that they don't have to do anything, which is great for giving non-technical users local inference without making them do scary native things. But keep in mind the actual experience for users is not great; the model download is orders of magnitude greater than downloading the browser itself, and something that needs to happen before you get your first token back. That's unfixable until operating systems start reliably shipping their own prebaked models that an API like this could plug into.
- Yokohiii 6mo ago> That's unfixable until operating systems start reliably shipping their own prebaked models that an API like this could plug into. Maybe the next big thing will be some software subscription premium offers with a bunch of 5090s as an extra.
- subhobroto 6mo ago> It works, I've shipped this as a "local inference"/poor person's ollama for low-end llm tasks like search fantastic! > the model download is orders of magnitude greater than downloading the browser itself, and something that needs to happen before you get your first token back sure but does this mean the model is lazily downloaded? that is, if I used this and I am the first time the model was called, the user would be waiting until the model was downloaded at that point? that sounds like a horrible user experience - maybe chrome reduces the confusion by showing a download dialog status or similar? also, any idea what the on disk impact is?
- why_is_it_good 6mo ago> Storage: At least 22 GB of free space on the volume that contains your Chrome profile.
- taejavu 6mo agoLmao and here I am still staunchly treating Blazor’s 2MB runtime as a deal-breaker
- deleted 6mo ago[deleted]
- iggerews 6mo ago[dead]
- nl 6mo agoThe model this uses is useless for anything beyond 2 round chat at the most. If you want to do anything interesting you need transformers.js and a decent mode. Qwen 0.9B is where things start working usefully
- iggerews 6mo ago[dead]
- jameslk 6mo agoSeems like a good way for a rogue JS script to offload token generation to a bunch of unsuspecting visitors It would actually be pretty interesting to see if its possible to decentralize the compute to generate something useful from a larger prompt broken down and sent to a bunch of browsers using a subagent pattern or something like RLM, each working on a smaller part of the prompt
- varun_ch 6mo agoThis feels like a lot of work for low reward, the technical/business infrastructure would be wild. And if anyone wants to offload their prompts to users browsers, they might as well just use the Chrome API correctly? How many server side prompts would realistically be useful to offload to a low end model like this? Plus even if you really wanted to do that, WebGPU exists and has for a while right?
- jameslk 6mo ago> How many server side prompts would realistically be useful to offload to a low end model like this? There's a lot of ways this API could go, e.g. more powerful models eventually, or perhaps integration with cloud models. For example, I could see Google trying to default Gemini as the model for users signed into Chrome
- varun_ch 6mo agoI think we’ll get more powerful models when they become reasonable to run on regular people’s computers, in which case the compute costs would hopefully fall enough that people don’t need to resort to this kind of weird stuff. As for cloud models, that would be interesting, although I guess then the fraud would be easier in spoofing whatever parameters (ip address? domain name? some Chrome install identifier?) to get around whatever rate limiting they come up with, rather than actually using people’s computers. Anyways I’m sure if it ends up being abused, they can throw a permissions dialog in front of it. Just need to figure out a way to make normal people understand.
- gorgoiler 6mo agoImagine a Vendor API that adds a way to link from the page straight into a device purchase workflow. As a trial of the API in Chrome you can order a new Google Pixel 9b directly from any page with the word Android in it! Or a LocalNet API that integrates with trusted hardware devices on your local network. As a trial (Chrome beta programme — strictly limited but here’s 3x signup links to share with your friends) you can adjust your Google Next Mini underfloor heating directly from Chrome! Or a DirectCast API that lets you stream <video> elements to a device of your choice even over a VPN. As a Chrome trial, you can use your Google Cloud account to stream directly from YouTube Premium to any linked Google Chromecast devices you own!
- haberman 6mo agoThis API seems perfect for an idea I've had for a while: a de-snarkifier for social media. Social media can be intellectually stimulating and educational, but it's also easy to get sucked into ideological sniping and flamewars, even if you didn't go looking for it. The emotional and intellectual energy spent flaming strangers on the Internet is a complete waste of human capital. With an API like this, I assume you could have a browser extension that could de-snarkify content before showing it to you. You could ask the LLM to preserve all factual content from the post, but to de-claw any aggressive or snarky language. If you really wanted to have fun, you could ask it to turn anything written in an aggressive tone into something that sounds absurd or incompetent, so that the more aggressive the post, the more it would make the author look silly. This could have a double benefit. For the reader, it insulates them from the personal attacks of random strangers on the Internet. Don't get me wrong, there is a time and a place for real, charged arguments about important issues that affect us all. But there is little to be gained from having those fights with strangers; on the contrary, I think it poisons the body politic when strangers are screaming at each other. For the writer, it takes away any incentive to be snarky or rude. If other people filter their content this way, there's no point in trying to be mean to them, and no "race to the bottom" for who can be more nasty.
- coalstartprob 6mo ago[dead]
- UqWBcuFx6NV4r 6mo agoI think the proposed extension would simply hide your comment, and all users would be better for it.
- danny_codes 6mo agoDomain names are a nice candidate for a Georgian tax
- arcknighttech 6mo ago[dead]
- afshinmeh 6mo agohttps://github.com/mozilla/standards-positions/issues/1067 https://github.com/mozilla/standards-positions/issues/1067
- kurtoid 6mo agoSee also: https://github.com/mozilla/standards-positions/issues/1213 https://github.com/mozilla/standards-positions/issues/1213
- rock_artist 6mo agoI think it's a step into a future of proper Model API. But it's just a small step. It reminds me of Apple's Foundation Models [1] While many AI integrations are focused on text communication / chat style. A lot of software benefits from non-text interfaces. I believe at some point OSes and browsers should provide an API to manage models so you'll have access to on-device/remote ones with a simplified interface for the app. Making something standardized that is cross-platform would be fantastic. It also needs to be on mobile devices, so the players that can easily make it happen are mostly Apple and Google. (Meta will follow or vice-versa I guess) Key-point: it shouldn't be exclusive to promoted models. (1) https://developer.apple.com/documentation/foundationmodels https://developer.apple.com/documentation/foundationmodels So the app would be able to query and get the right model(s).
- elpakal 5mo agoApple's Foundation models seem great on paper until you see the 4k context window. (though I know we are still early in this chapter).
- gopalv 6mo agoThe better part of this is having a local-first AI, particularly because it has tool-calling builtin & structured output. I haven't pushed out a full version[1] which uses ducklake-wasm + this to make a completely local SQL answering machine, but for now all it does is retype prompts in the browser. [1] - https://notmysock.org/code/voice-gemini-prompt.html https://notmysock.org/code/voice-gemini-prompt.html
- domenicd 6mo agoI led the design effort on this API, before retiring. Here's my writeup on some of the considerations that went into it: https://domenic.me/builtin-ai-api-design/ https://domenic.me/builtin-ai-api-design/
- comboy 6mo agoHow do you envision short term and long term target usage of it? And do you guys communicate between other browsers when doing something like this to try to settle on something common? I don't mean W3C but practically, it's a small world after all.
- domenicd 6mo agoI can't speak for "you guys" anymore, as I'm retired, but from my personal perspective/recollection: The target usage for the prompt API is anything that would benefit from the general capabilities of a language model, and can't be encompassed by the more-specific APIs for summarization/writing/rewriting. Realistic use cases currently are things like sentiment analysis, keyword extraction, etc. I have a number of ideas on how to integrate it into my current retirement project around Japanese flashcards, e.g. generating example sentences. If the small (~10 GiB) model class keeps getting smarter, the class of things possible on-device in this way gets larger and larger over time. We definitely communicated with other browsers. There were the standing WebML Community Group meetings at the W3C every few weeks. There were async discussions like https://github.com/mozilla/standards-positions/issues/1213 https://github.com/mozilla/standards-positions/issues/1213 and https://github.com/WebKit/standards-positions/issues/495 https://github.com/WebKit/standards-positions/issues/495 . (Side note, I love the contrast between Mozilla's helpful in-depth feedback and WebKit's... less helpful feedback.) There was also a bit of a debacle where the W3C Technical Architecture Group tried to give "feedback" but the feedback ended up being AI-generated slop... https://github.com/w3ctag/design-reviews/issues/1093 https://github.com/w3ctag/design-reviews/issues/1093 . But overall, yeah, the goal with the prompt API, as with all web APIs, is to put something out there for discussion as early as possible, and get input from the broad community, especially including other browsers, to see if it's something that they are interested in collaborating on. https://www.chromium.org/blink/guidelines/web-platform-changes-guidelines/ https://www.chromium.org/blink/guidelines/web-platform-chang... (which I also wrote) goes into how the Chromium project thinks about such collaboration in general.
- tethys 6mo agoSlightly off-topic: Refreshing to see these two authors link to their Bluesky and Mastodon profiles. No Twitter/X in sight!
- izietto 6mo agoCan pass to it the current page contents for a AI-based AdBlock / cookie manager / etc.?
- denniszelada 6mo ago[dead]
- benjaminbenben 6mo agoWe use this for summarising our hack day write ups: https://remotehack.space/previous-hacks/ https://remotehack.space/previous-hacks/ It's a tiny script that looks up the rss feed and uses the content to generate summaries; quite a nice fit with our static site. Sometime I'd like to extend it to ask different questions about the content.
- Ronsenshi 6mo agoNot long before all of the web content will be going through these AI pipelines where user might not even see original webpage.
- timxtokyo 5mo agothe world of agentic ai!
- meander_water 6mo agoThis looks like it uses Gemini Nano under the hood. But the latest Gemma4 E2B and E4B models appear to be much better, so you'd probably be better off deploying quantized versions through an extension for now. - Gemini Nano-1: 46% MMLU, 1.8B - Gemini Nano-2: 56% MMLU, 3.25B - Gemma4 E2B: 60.0% MMLU, 2.3B - Gemma4 E4B: 69.4% MMLU, 4.5B Sources: - https://huggingface.co/google/gemma-4-E2B-it https://huggingface.co/google/gemma-4-E2B-it - https://android-developers.googleblog.com/2024/10/gemini-nano-experimental-access-available-on-android.html https://android-developers.googleblog.com/2024/10/gemini-nan...
- domenicd 6mo agoI no longer have any inside knowledge, but from my time on this team they were very quick about getting the latest small (Google) models into Chrome. I expect that if Gemma 4 (or its equivalent Gemini Nano) isn't already in Chrome, then it will be soon. Note that the article here was last updated 2025-09-21, and as of that time it was already on Gemini Nano 3.
- meander_water 6mo agoThanks for the insider info! Do you know if there are any published benchmarks for Nano 3?
- eis 5mo agoGoogle will soon release Gemini Nano 4 based on Gemma 4. A "Fast" version based on Gemma 4 E2B and a "Full" version based on E4B. https://android-developers.googleblog.com/2026/04/AI-Core-Developer-Preview.html https://android-developers.googleblog.com/2026/04/AI-Core-De...
- kbx 5mo ago[dead]
- ceejayoz 5mo ago> This looks like it uses Gemini Nano under the hood. Yes; "With the Prompt API, you can send natural language requests to Gemini Nano in the browser."
- mudkipdev 6mo agoGemini Nano, unlike Gemma, is not open-weight, right? I would be interested in dumping the model weights, unless someone has done that already
- me551ah 6mo agoI’m just wondering how much more RAM and VRAM chrome will use after these changes
- tom1337 5mo agoThe idea of having local LLMs accessible in the browser for privacy concerning is nice i guess but when each browser has a different model attached to this API testing becomes even more a nightmare then now. I wonder if this will drive more users towards chrome because most of the usages of this API might be just tailored to fit the Gemini Nano model?
- jimmypk 5mo ago[flagged]
- michaelbuckbee 5mo agoFwiw - I did a fairly large comparison of Gemini Nano (the in browser ai model) vs a comparable free hosted model of Gemma (from OpenRouter) and the hosted model absolutely trashed the local model on every aspect of speed, reliability, availability, etc. [1] I'm not particularly happy about that outcome as I wish we had more locally run AI models for reasons of privacy and efficiency, so this is more just a warning that at present there are some severe tradeoffs. 1 - https://sendcheckit.com/blog/ai-powered-subject-line-alternatives https://sendcheckit.com/blog/ai-powered-subject-line-alterna...
- kbx 5mo agoHey, Chrome PM for built-in AI here. Thanks for the write-up and the comparison, but more importantly for using the API in production! You’re highlighting the "state of the art" gap we’re working to close. Cloud models will always have the advantage of massive parameter counts, but our bet is that for a huge class of simpler or high-volume tasks, the upsides of on-device (e.g. zero-cost, permission-less start with no quotas/infra, network-resilience, privacy) make it a compelling trade-off. The models have been getting better at a rapid clip, and the team is heads-down on optimizing performance and reliability. To that end, we're always grateful for feedback. If you hit specific bugs, crashes, or quality regressions, filing a report with repro steps is the best way to help us improve. You can file those on crbug.com under the "Chromium > Blink > AI" component.
- oneeyedpigeon 5mo agoEvery time I see "prompt" nowadays, I'm briefly hopeful that I'm going to read something about $PS1. Then, inevitably, AI disappoints me yet again.
- ilaksh 5mo agoAny chance this will be supported by Firefox or other browsers soon?
- franze 5mo agotrying to wrap it in a unix style CLI see: https://github.com/Arthur-Ficial/fenster https://github.com/Arthur-Ficial/fenster and: https://news.ycombinator.com/item?id=47923692 https://news.ycombinator.com/item?id=47923692 hard work so far
- david_shi 5mo agoWill this API and others like it will be a strong enough incentive to move away from Chromium based browsers and back on to Chrome?
- 6thbit 5mo agoIs “Ship a 22gb model on your product” the new “put a chat window on your product”? I agree with others this fits better in the OS, or hey maybe Apple sells a time-machine sort of NAS with neural engine chips.
- solarkraft 5mo agoInteresting. Questionable from a web standards POV, but interesting. Who‘s gonna make it call tools?
- jussy 5mo agoI get the filtering use case here and I'd say the other one is personalisation of generic marketing copy.