10 ms·
Updates to our web search products and Programmable Search Engine capabilities
- deleted 9mo ago[deleted]
- jonplackett 9mo agoAre search engines like Kagi completely screwed by this or is there a way for them to keep operating?
- direwolf20 9mo agoKagi doesn't have a partnership with Google - they work under adversarial interoperability, stealing results from Google against their will, and paying some third-party to enable this. They'd like to simply pay Google, but Google doesn't want their money.
- TiredOfLife 9mo agoKagi is backed by russia so they will be fine.
- contagiousflow 9mo agoWhat do you mean "backed by"
- Hackbraten 9mo agoYou’re probably referring to the fact that their search results include entries from Yandex. That’s something entirely different from being “backed by Russia.” If anything, they pay Russia, not the other way around.
- direwolf20 9mo agoAnd Yandex has better results than Google, so I support this move. The USA does more wars than Russia does anyway.
- HPsquared 9mo agoI had misread the title as "Google is ending (full-web search) for [aka in favour of] (niche search engines)" The correct parsing is: "Google is ending (full-web search for niche search engines)"
- dredmorbius 9mo ago"Google will discontinue third-party niche search engine access to full-web search" would be far clearer. Given that the title supplied is effectively editorialised, and the original article's title is effectively content-free ("Updates to our Web Search Products & Programmable Search Engine Capabilities"), my rewording would be at least as fair. HN's policy is to try to use text from the article itself where the article title is clickbait, sensational, vague, etc., however. I suspect Google's blog authors are aware of this, and they've carefully avoided any readily-extracted clear statements, though I'll take a stab... Here's the most direct 'graph from TFA: Custom Search JSON API: Vertex AI Search is a favorable alternative for up to 50 domains. Alternatively, if your use case necessitates full web search, contact us to express your interest in and get more information about our full web search solution. Your transition to an alternative solution needs to be completed by January 1, 2027. We can get a clearer, 80-character head that's somewhat faithful to that with: "Google Search API alternative Vertex AI Search limited to 50 domains" (70 chars). That's still pretty loosely adherent, though it (mostly) uses words from the original article. I'm suggesting it to mods via email at hn@ycominator.com; others may wish to suggest their own formulations.
- 01jonny01 9mo agoGoogle quietly announced that Programmable Search (ex-Custom Search) won’t allow new engines to “search the entire web” anymore. New engines are capped at searching up to 50 domains, and existing full-web engines have until Jan 1, 2027 to transition. If you actually need whole-web search, Google now points you to an “interest form” for enterprise solutions (Vertex AI Search etc.), with no public pricing and no guarantee they’ll even reply. This seems like it effectively ends the era of indie / niche search engines being able to build on Google’s index. Anything that looks like general web search is getting pushed behind enterprise gates. I haven’t seen much discussion about this yet, but for anyone who built a small search product on Programmable Search, this feels like a pretty big shift. Curious if others here are affected or already planning alternatives. UPDATE: I logged into Programmable Search and the message is even more explicit: Full web search via the "Search the entire web" feature will be discontinued within the next year. Please update your search engine to specify specific sites to search. With this link: https://support.google.com/programmable-search/answer/12397162 https://support.google.com/programmable-search/answer/123971...
- throwaway_20357 9mo agoWhat are some of the niche search engines build on Google's index affected by this?
- doublerabbit 9mo agoKagi
- deleted 9mo ago[deleted]
- marginalia_nu 9mo agoThey published this the other day: https://blog.kagi.com/waiting-dawn-search https://blog.kagi.com/waiting-dawn-search Which saw some discussion on HN.
- 9mo ago
- Antibabelic 9mo agoRelevant: Waiting for dawn in search: Search index, Google rulings and impact on Kagi https://news.ycombinator.com/item?id=46708678 https://news.ycombinator.com/item?id=46708678
- mrweasel 9mo agoThis might be me reading it wrong, but isn't shutting down the full-web search going against the ruling mentioned in the Kagi post? > Google must provide Web Search Index data (URLs, crawl metadata, spam scores) at marginal cost. Maybe they're shutting down the good integration and then Kagi, Ecosia and others can buy index data in an inconvenient way going forward?
- Hackbraten 9mo agoIf I understand Kagi's blog post correctly, then here's what happened, chronologically: Kagi makes deals with many search engines so they can have raw search results in exchange for money. Google says: no, you can't have raw search results because only whales can get those. Only thing we can offer you is search results riddled with ads and we won't allow you to reorder or filter them. Kagi thinks Google's offer is unacceptable, so Kagi goes to a third party SERP API, which scrapes Google at scale and sells the raw search results to Kagi and others. August 2024: Court says Google is breaking the law by selling raw search results only to whales. December 2025: Court orders that for the next six years, 1. Google must no longer exclude non-whales from buying raw search results, 2. Google must offer the raw search results for a reasonable price, and 3. Google can no longer force partners to bundle the results with ads. December 2025: Google sues the third-party scraping companies. January 2026: Google says "hey, the old search offering is going to go away, there's going to be a new API by 2027, stay tuned."
- mrweasel 9mo agoI don't really see any mentioning of a new API, beyond their Vertex AI thing, and I don't know how comparable that might be. Also it is capped at 50 domains (by default). It is perhaps a clever legal workaround. They must sell access to their index, but the verdict didn't state how much of it you can buy access to at any one time. So they put a limit of 50 domains, because that accommodates everyone who's not a search engine, but effectively blocks Kagi and Ecosia, while not exactly refusing to sell to them.
- chromehearts 9mo agoIs this about the little Google Search Bar that is present on some websites? Or am I mistaking something
- 01jonny01 9mo agoKind of, however the Google Search Bar present on website is usually there to search across their domain, the search results are limited to their domain e.g example.com/page1, example.com/page2. Google will carry on supporting this. What they are ending is their support for websites to search across the entire web. The websites that search across the entire web are usually niche search engine websites.
- chromehearts 9mo agoAhh; so that's the difference. Thanks!
- vaylian 9mo agoMeanwhile in Europe: Qwant and Ecosia team up to build their own search index: https://blog.ecosia.org/eusp/ https://blog.ecosia.org/eusp/
- tweetle_beetle 9mo agoIt's a noble effort, but they're so late to the game that it's hard to see them making a significant dent. I hope I'm wrong. They were: > aiming to serve 30% of French search queries [by end of 2025] https://blog.ecosia.org/launching-our-european-search-index/ https://blog.ecosia.org/launching-our-european-search-index/
- Gigachad 9mo agoI feel like soon there won’t even be a point having a search engine since almost the entire internet will be useless AI slop.
- altairprime 9mo agoIt's as though full-text search of websites you've never heard of was a mistake :) PageRank wouldn't exist without webrings, directories, and forums you could only search individually, and we thrived on that Internet. Welcome back, ye olde Internet.
- baubino 9mo agoThe old internet is still there. It hasn‘t gone away; it‘s just undiscoverable with ad-based search. The more slop there is, the more necessary it is to have good search engines.
- zelphirkalt 9mo agoRecently, I set up a fresh system on a laptop. Ahahahaaa, how utterly crap Google search results now are! It fills me with some stress and disgust to use that. Now one of the first things I do, right after emergency using duckduckgo to search for uBlock Origin and NoScript, is to get Kagi search installed as default search. Then I can continue setting things up more calmly.
- londons_explore 9mo agoWhat examples are there of people using this?
- 01jonny01 9mo agoThere is literally thousands of independent search engines that use Programmable search to search the entire web. Many ISP providers use it on their homepage, kids-based search engines like wackysafe.com use it, also search engines that focus on privacy like gprivate.com etc
- TeMPOraL 9mo agoAlso LLM tools. Programmable Search Engine API was a way to give third-party LLM frontends the ability to give LLMs a web search tool. Notably, this was a common practice long before any of the major LLM providers added search capabilities to their frontents.
- 01jonny01 9mo agoExactly, Google want every one depended on Gemini.
- YoungX 9mo ago[flagged]
- bovermyer 9mo agoI'm curious about what it would take to build my own "toy" search engine with its own index. Anyone ever tried this?
- Gigachad 9mo agoMight find YaCy interesting. It’s meant to be a decentralised search engine where users scrape the internet and can search other users indexes in a kind of torrent like way. I found it didn’t really work as a real search engine but it was interesting.
- reddalo 9mo agoGood luck scraping websites without being blocked, if you're not Google.
- marginalia_nu 9mo agoWell you'll get blocked some places but it's not too big of a deal. If you're running an above board operation, you can surprisingly often successfully just email the admin explaining what you're doing, and ask to be unblocked.
- BolsunBacset 9mo agoSounds very time consuming. Glad you're able to sustain yourself to be able to do it full time.
- marginalia_nu 9mo agoYeah that's where I started out in 2021. Been at it for almost 5 years now, last three of which full time. I'm indexing about 1.1 billion documents now off a single server. Hard part is doing it at any sort of scale and producing useful results. It's easy to build something that indexes a few million documents. Pushing into billions is a bigger challenge, as you start needing a lot of increasingly intricate bespoke solutions. Devlog here: https://www.marginalia.nu/tags/search-engine/ https://www.marginalia.nu/tags/search-engine/ And search engine itself: https://marginalia-search.com/ https://marginalia-search.com/ (... though it operates a bit sub-optimally now as I'm using a ton of CPU cores to migrate the index to use postings lists compression, will take about 4-5 days I think).
- lighthouse1212 9mo agoThe 'Google Graveyard is real' sentiment captures something important: every dependency on a large platform is a loan that can be called in. The 34-million-document indie index project someone mentioned is the right response - own your core infrastructure. Easier said than done for whole-web search, but the same principle applies everywhere.
- 01jonny01 9mo agoMuch easier said than done, especially if you are serving users on scale.
- 1718627440 9mo agoSince the issue here is self-hosting and "core infrastructure", that isn't a problem, but everyone has their own search index, isn't credible either.
- jamesbelchamber 9mo agoAre competing search indexes (Bing, Ecosia/Qwant, etc) objectively worse in significant ways, or is Google just so entrenched that people don't want to "risk it" with another provider (and/or preferences and/or inertia). I suppose I'm asking whether this is actually a _good thing_ in that it will stimulate competition in the space, or if it's just a case that Google's index is now too good for anyone to reasonably catch up at this point.
- 01jonny01 9mo agoThe beauty about Google Programmable Search across the entire web is that it's free and users can make money by linking it their Adsense account. Bing charge per query for the average user. Ecosia and Qwant use Bing to power their results, probably under some type of license, which results in them paying much less per query than a normal user.
- thayne 9mo agoBing recently shut down their API product, which was already very expensive. If you want programmatic access to search results there aren't really many options left.
- SirHumphrey 9mo agoI can manage fine with other search indexes for English language searches; weather that is because others got better or google got worse i cannot tell, though I suspect the latter. But for searching in more niche languages google is usually the only decent option and I have little hope that others will ever reach the scale where they could compete.
- Antibabelic 9mo agoBing's index is smaller than Google's, and anecdotally I get fewer relevant results when using it, particularly from sites like Reddit that have exclusive search deals with Google.
- carlosjobim 9mo agoYes, for non English queries they are all rubbish. And that's billions of users.
- sreekanth850 9mo agoNever build a product with core feature depending on a third-party, you will eventually get fucked up for sure. always have a 70:30 rule for revenue where 70% is core independent features.
- halapro 9mo agoSoon you'll find that you cannot exist on the web without relying on third parties. Sometimes you'll even have trouble getting paid thanks to the painful existence of payment processors.
- sreekanth850 9mo agoTrue, you can’t exist without 3rd parties. But you shouldn’t let them be your core moat/USP. Jasper is a great example, they depended too much on LLM access, then ChatGPT launched and ate the value. Using third party APIs is fine, but building a product whose core depends on them is suicide.
- direwolf20 9mo agoThat's why I eschew HTTPS.
- cubefox 9mo agoIs this perhaps to prevent ChatGPT, Claude and Grok to use Google Search? It would make sense for Google to keep that ability for Gemini.
- 01jonny01 9mo agoI suspect its going to hurt the indie developers and small start-ups who do not have special licensing agreements.
- direwolf20 9mo agoThey'll go adversarial interop through SerpAPI, just like Kagi does. SerpAPI will get the money instead of Google getting it.
- cubefox 9mo ago"Why we’re taking legal action against SerpApi’s unlawful scraping" https://blog.google/innovation-and-ai/technology/safety-security/serpapi-lawsuit/ https://blog.google/innovation-and-ai/technology/safety-secu...
- snackbroken 9mo agoWhat's their angle here? Courts have been over the "Is scraping websites that don't want to be scraped OK?" question plenty of times. From > SerpApi deceptively takes content that Google licenses from others (like images that appear in Knowledge Panels, real-time data in Search features and much more), and then resells it for a fee. In doing so, it willfully disregards the rights and directives of websites and providers whose content appears in Search. it sounds like they are somehow suing on behalf of whoever they are licensing content from, but does that even give Google standing? I guess I'm asking if they actually are hoping to win or just going for a "the process is the punishment"+"we have more money and lawyers than you" approach.
- solarkraft 9mo agoThis will significantly impact (quite possibly kill) Startpage and Ecosia, who are effectively white-label Google, right? What alternatives are there besides Bing? Is it really so hard that it’s not considered worth doing? Some of the AI companies (Perplexity, Anthropic) seem to have managed to get their own indexing up and running.
- ColinHayhurst 9mo agoExcuse the self-promotion but Mojeek offers a web search API (>9 billion pages): https://www.mojeek.com/services/search/web-search-api/ https://www.mojeek.com/services/search/web-search-api/
- direwolf20 9mo agoNice! I hope Kagi is licensing access from you?
- ColinHayhurst 8mo agoThey do
- deleted 9mo ago[deleted]
- zoobab 9mo agoAntitrust do not work against large companies. Just dissolve them in acid.
- marginalia_nu 9mo agoThis is the type of monopoly abuse these laws were designed to target, and antitrust laws actually do work against large companies. If you actually enforce them. Unfortunately, during the Reagan administration, political sentiment toward monopolies shifted and since then antitrust law has been a paper tiger at best.
- zoobab 9mo agoI heard when Bush came to power, the antitrust complaint against Microsoft monopoly driven by the government was dropped.
- shevy-java 9mo agoGoogle has consistently ruined its search engine in the last (almost) 10 years. You can find numerous articles about this, as well as videos on youtube (which is also controlled by google). Not long ago they ruined ublock origin (for chrome; ublock origin lite is nowhere near as good and effective, from my own experience here). Now Google is also committing towards more evil and trying to ruin things for more - people, competitors, you name it. We can not allow Google to continue on its wiched path here. It'll just further erode the quality. There is a reason why "killed by google" is more than a mere meme - a graveyard of things killed by google. We need alternatives, viable ones, for ALL Google services. Let's all work to make this world better - a place without Google.
- philistine 9mo agoTo me there are two eras of the Google Graveyard(tm). First, there's the we're a university research group with an ad company footing the bill era. That's the early Google era, and it was a consequence of its corporate structure. They valued new projects, market fit, profitability, and maintenance be damned. We're in the second era. The era of the MBAs are shutting down the last remnants of openness the company ever had.
- consumer451 9mo agoDumb question: I keep seeing posts about how ~"the volume of AI scrapers is making hosting untenable." There must a ton of new full-web datasets out there, right? What are the major hurdles that prevent the owners of these datasets from providing them to third parties via API? Is it the quality of SERP, or staleness? Otherwise, this seems like a potentially lucrative pivot/side hustle?
- Terretta 9mo ago> the volume of AI scrapers is making hosting untenable Aside from that potential, it's also not true. A Pentium Pro or PIII SSE with circa 1998-99 Apache happily delivers a billion hits a month w/o breaking a sweat unless you think generating pages for every visit is better than generating pages when they change.
- Tenemo 9mo agoI think it is true that it is a real problem (EDIT: but doesn't necessarily make "hosting untenable"), but you are correct to point out that modern pages tend to be horribly optimized (and that's the source of the problem). Even "dynamic" pages using React/Next.js etc. could be pre-rendered and/or cached and/or distributed via CDNs. A simple cache or a CDN should be enough to handle pretty much any scrapping traffic unless you need to do some crazy logic on every page visit – which should almost never be the case on public-facing sites. As an example, my personal site is technically written in React, but it's fully pre-rendered and doesn't even serve JS – it can handle huge amounts of bot/scrapping traffic via its CDN.
- consumer451 9mo agoOK, I agree with both of you. I am an old who is aware of NGINX and C10k. However, my question is: what are the economic or technical difficulties that prevent one of these new web-scale crawlers from releasing og-pagerank-api.com? We all love to complain about modern Google SERP, but what actually prevents that original Google experience from happening, in 2026? Is it not possible? Or, is that what orgs like Perplexity are doing, but with an LLM API? Meaning that they have their own indexes, but the original q= SERP API concept is a dead end in the market? Tone: I am asking genuine questions here, not trying to be snarky.
- deleted 9mo ago[deleted]
- nairboon 9mo agoRegarding alternate search engines: I consider the idea of YaCy kind of interesting: a P2P search engine: https://yacy.net/ https://yacy.net/ Although, it needs some more work and peers to be usable as a general-purpose search engine.
- deleted 9mo ago[deleted]
- mark_l_watson 9mo agoNot directly covered by this blog, but for low cost and good performance the combination of gemini-3-flash with search grounding is hard to beat, at least for the many small experiments I use it for. One thing touched upon in comments here: I never understood how it was proper for 3rd parties to scrape Google search results and reuse/resell them. Really off topic, sorry, but I am surprised that more companies don’t build local search indices for just the few hundred web domains that are important to their businesses. I have tried this in combination with local (small and fast) LLMs and I think this is unappreciated tech: fast, cheap, and local.
- bennydog224 9mo agoI built many products on Google PSE (Custom Search). Results were nowhere near as good as regular Google, but still useful. I usually needed to use another library to get the DOM content anyway. But it still was solid for grounding/checking data. RIP, another one to the Google Graveyard.
- jpadkins 8mo agoyour usage was in obvious violation of the terms of service. This is why we can't have nice things.
- motoboi 9mo agoThis and agressive anti-bot at YouTube is Alphabet closing the AI data leaking
- thayne 9mo agoDoes this mean the !g bang will stop working in DuckDuckGo?
- direwolf20 9mo agoDoesn't it just redirect you to Google? So it will still work.
- jpalepu33 9mo agoThis is a clear example of why building on proprietary APIs is risky for indie devs and small startups. I've seen similar patterns with Twitter's API restrictions and other platforms gradually closing down their ecosystems. For anyone affected: consider this a forcing function to either: 1. Build your own lightweight search infrastructure (tools like Meilisearch, Typesense make this more accessible now) 2. Use adversarial interop via services like SerpAPI (though Google is already taking legal action there) 3. Pivot to specialized vertical search where you control the data sources The real lesson here is the importance of owning your core value proposition. If your product's moat depends entirely on a third-party API that can be yanked away with 12 months notice, you don't really have a sustainable business. Google is essentially saying: indie search is dead, pay enterprise prices or leave. This will probably accelerate the trend toward specialized, domain-specific search engines that don't rely on Google's index at all.
- joelboersma 9mo agoI've been occasionally working on a toy project that's basically "Google search in a TUI" that used this API. I was already planning on adding Brave Search as an option for a different backend, and I was heavily considering making it the default just because it's much easier to set up on the user's end. This is the straw that broke the camel's back.
- bicepjai 9mo agoWhy don’t we have something more “torrent-like” for search? Imagine a decentralized network where volunteers run crawler nodes that each fetch and extract a tiny slice of the web. Those partial results get merged into open, versioned indexes that can be distributed via P2P (or mirrored anywhere). Then anyone can build ranking, vertical search, or specialized tools on top of that shared index layer. I get that reproducing Google’s “Coca-Cola formula” (ranking, spam fighting, infra, freshness, etc.) is probably unrealistic. But I’d happily use the coconut-water version: an open baseline index that’s good enough, extensible, and not owned by a single gatekeeper. I know we have common crawl, but small processing nodes can be more efficient and fresh
- pona-a 9mo agoLook up YaCy. This might be close to what you imagine
- bicepjai 9mo agoThanks for that info, they are doing exactly what I was saying. Why is that not adopted widely ? Found HN posts YaCy, a distributed Web Search Engine, based on a peer-to-peer network https://news.ycombinator.com/item?id=39612950 https://news.ycombinator.com/item?id=39612950
- qingcharles 8mo agoWhat's to stop someone poisoning the data, though? :(
- aghilmort 8mo ago[dead]