5 ms·
> If I run a local script or an LLM that needs to fetch a web page from your site and you block that script or LLM as being a bot, you hurt me, the user. No; i
by hk__2 2mo ago
> If I run a local script or an LLM that needs to fetch a web page from your site and you block that script or LLM as being a bot, you hurt me, the user.
No; in this case you are not a user, you are a bot user.
- ako 2mo agoThe best way to read the information on the internet today is via a LLM. Just like you don't process raw data in a data warehouse or datalake by looking at tables, you use a SQL or BI tool, to process all the information on the internet you need to tool to digest it for you. For many, today that tool is a LLM chat interface or agent.
- deleted 2mo ago[deleted]
- ballooney 2mo agoThis is such a grim thing to read.
- doc_ick 2mo ago100%
- aprdm 2mo agoWhy do you feel that ?
- sdellis 2mo agoBecause AI cannot be trusted as a reliable source of information. Information is more trustworthy when it comes from from primary sources on the open web.
- aprdm 2mo agoWhy not ? Can google be trusted ? Or facebook ? How do people have been accessing the internet for the last decade you reckon and how's this any different ?
- giantrobot 2mo ago> Can google be trusted ? Or facebook ? No. Neither can be trusted as far as you can throw them. They're both incredibly invasive data brokers. Their customer facing products are just vehicles to show ads and Hoover up more PII and behavioral data on everyone.
- antiterra 2mo agoI routinely run into Claude Opus and Fable making basic mistakes like misunderstanding a simple negation, which would be on top of whatever reliability issues there are with a source. I think that means it is functionally very different.
- lunar_mycroft 2mo agoGoogle (as it originally) existed wasn't a source of information, it was an *index* of it. You were trusting google as a source for {search_query}, you were trusting the links it gave as a source (based on your own evaluation). LLMs are fundamentally different, because you are trusting the software to actually generate the information in a truthful way. If you use google (sans AI), you're putting some trust in their page ranking algorithm. If you use it with AI, you're trusting the same algorithm (since that's how the model gets it's sources), but then you're trusting the model to evaluate the sources for credibility and extract the information you actually want.
- ako 2mo agoPagerank is just a popularity contest, doesn't say anything about trust. AI probably does a better job determining trustworthyness than pagerank.
- 2mo ago
- falcor84 2mo agoWell agreed, but it's not new, the web has been actively made user hostile by every website owner and their "273 partners".
- deleted 2mo ago[deleted]
- grumbel 2mo agoHow is that grim? It's the dream of the Semantic Web coming true, just by different means than planed.
- infinitezest 2mo agoBecause the economic, social, and climate impacts are at best uncertain and at worst devastating to the majority of human beings on the planet. It's always surprising to me that people don't intuitively separate the usefulness of AI from its risks. It feels like everyone has to be a doomer or a booster.
- lostmsu 2mo ago> at worst devastating This is such a low bar literally anything would clear it. Even tomato farming.
- infinitezest 2mo agoI think you're going to have to flesh this out a little more, friend. I'm not sure what you're trying to say.
- account42 2mo agoYou know that the reason the semantic web didn't work out is that it only benefits a small number of participants (aggregators), right?
- mmh0000 2mo agoI disagree. I personally consider these my biggest problems with the Web: - Bias, specifically commercial bias - Webpage formatting: every website looks different, hides the information I want in different places. Sure, it "looks pretty", but I don't want pretty; I want info - Scams/SEO/etc... The LLMs are very good at reading from multiple sources, parsing, and presenting only the data in a consistent format. There are many examples of this, but if you want a good one to try for yourself: Google "How to make ham fried rice"; you'll get 10,000 articles, most pretty good recipes. But they're all different; most of them are just bait for ads. And most of them, the 10-line recipe is hidden between 50 useless paragraphs about how serving food is life's most important goal. Now, ask an LLM to search for it, find the best combination, and list only the recipes. You get a perfect, 10-line recipe that doesn't waste your time.
- sethops1 2mo agoAnd if you're lucky, it won't include rocks as an ingredient.
- vel0city 2mo agoThey're an important part of your diet if you happen to have a gizzard. I've also been known to include some ground up rocks in my meals. I'm pretty picky though about them, I mostly just want a couple specific varieties.
- johneth 2mo agoLLMs are not unbiased. Nothing is unbiased.
- tekla 2mo agoYeah I had to laugh at that. LLM's solve commercial bias? Coming from seemingly multiple the most valuable companies of all time?
- altmanaltman 2mo agoWhy even use a LLM, just buy a cookbook and you can even waste less time and know the information is valid. And again if your logic is consistent, google routes your search to an LLM automatically and renders the reply - that will save as much time and do the same thing since LLMs are so good at parsing and presenting information. So why should I pick an LLM over google's inbuilt LLM or an actual cookbook on the desk? What wasted time is the LLM saving?
- ako 2mo agoMaybe, but it's reality. Not too long ago the rule would be "if it doesn't show up in google, it doesn't exist". Today agents+LLMs are better browsers than Chrome/Safari/Firefox/... If a webpage does not show up in an LLM it may as well not exist.
- wulfmann 2mo agoThe worst way to read the information on the internet today is via a LLM.
- roboror 2mo agoThis feels like "the best way to read a book is via CliffsNotes"
- criley2 2mo agoScenario one: you use software to connect to their server and download a webpage. You are a user. Scenario two: you use software to connect to their server and download a webpage. You are a "bot". Make it make sense
- selckin 2mo agobecause they maintain the websites for social status, if the user never sees the specific website or knows they used it, you can't gain any social status
- Flashtoo 2mo ago> they maintain the websites for social status Are you implying that's a bad thing? Doing things for social status has been an integral part of society for millennia. It's a legitimate motivation that can benefit both the recipient of the status and the rest of society. In this specific case, if you assume that the author makes content that is useful to you only because of the social status reward, taking away that reward means the author will stop making content that is useful to you.
- --_-- 2mo agoOne more reason for the bot difference is the llm bot users are doing almost all the traffic, and it feels wasteful, painful, and there are reports of 99.x% being llm bots just hitting webpages over and over pointlessly. This is a change. People put up new caching layers, and it urks people running a website in a single small machine. Yes, this could always happen with slashdotting but it's different now.
- hk__2 2mo agoYou’re missing the part about the human who interacts with the webpage.
- ronsor 2mo agoDoesn't matter. Even by the typical robots.txt definition, a bot has to operate more or less fully autonomously. If the user initiated the request, it doesn't count, and it shouldn't follow robots.txt.
- skinfaxi 2mo agoI assume you are unfamiliar with the concept of user agents? Otherwise your browser would count as a bot user no? And if not, what if it was a custom browser and not Chrome/Firefox/Edge?
- deleted 2mo ago[deleted]
- elorant 2mo agoWhat if I’m a web alerts company? I crawl your content but my clients are all end users who actually see your content at your site.
- deleted 2mo ago[deleted]
- deleted 2mo ago[deleted]
- hk__2 2mo agoFetch my RSS then.
- pessimizer 2mo agoThen you are a uninvited bot that crawls my content in order to send alarms to people to tell them to visit my site? This is a "what if it's for your own good?" argument. What if I break into your house to clean your toilets? Although amusingly, if you're a "web alerts company" that I didn't contract and sends out alerts in batches on your own schedule, you will probably send a ton of your customers to my site at the same time and slashdot me off of the web entirely.
- elorant 2mo agoNo, I won't send a ton, I'll send a few dozens to a few hundreds at most because not everyone is interested in the same things. And they'll visit at their own timeline. If your site can't service a few dozen requests simultaneously then you probably aren't a news site in the first place so the whole argument is moot.
- Terretta 2mo agoyou know what a browser is called by the web site? check the header that tells the version. USER AGENT not user, an agent on behalf of the user. the web was intended to be usable by agents belonging to users. in our path of the multiverse, the main agent we think of happens to be live interactive browser. but that need not have been, nor will it be in the future, the dominant USER AGENT. for a glimpse at one possible futureverse, check out what the home assistant community is up to, how they assemble then update the ambient information displays on their walls. their USER AGENTS are doing what most HN-style "hackers" dreamed of reading scifi as kids. the alternative is all your in home information owned by corpos when the info should be from user agents not corpo feeds. if you want to vote this idea down, you might be a corpo. :-)
- doc_ick 2mo ago[dead]
- charcircuit 2mo agoThis website refers to browsers as browsers. Just because a header in the underlying protocol is named a certain way, that doesn't necessarily reflect what the humans intend. https://www.ycombinator.com/legal https://www.ycombinator.com/legal
- dspillett 2mo ago> not user, an agent on behalf of the user. Identified by the user agent header. Which most bots fake or leave out to increase the chance of getting where they are not wanted. Your bot is a good bot? Great, let us know when you've dealt with all the bad bots and we'll open the doors to the remaining (good) bots again.
- asgraham 2mo agoCan you recommend a specific home assistant community to check out?
- nickthegreek 2mo agor/homeassistant
- matsemann 2mo agoI don't disagree, but there is a sliding scale here. For instance, I wanted to buy a piece of equipment the other day from a local company for a specific usecase. I wanted to find a specific price/weight/specs ratio, and asked an llm to loop through the 20 or so items, fetch their page and calculate and present some values for each. This then led to me going and buying the one I found. So the llm was mainly just an extension of me clicking into every page and making a spreadsheet myself. However, if it were to continuously poll, or just scrape or something with no intention of buying, I would be no better than a bot.
- hluska 2mo agoI don’t see the sliding scale - if the website didn’t want a bot they could have blocked you. They clearly don’t mind so what’s the issue? You still used a bot.
- deleted 2mo ago[deleted]
- baby_souffle 2mo ago> I don’t see the sliding scale - if the website didn’t want a bot they could have blocked you. The point is that not all bots are bad. Assuming as much / implementing policies to that effect won't block all bots but it _will_ block the portion of bots that represent users considering giving you money. I have a series of bots that monitor various eCom sites to monitor prices over time for big-ticket items I am considering as well as staples/groceries and everything in between. I have this little scrape/ingest pipeline because there's no other way to obtain this data... not even an API that I can pay for access to. One of the large appliance sellers that I have in the scrape queue has gotten _hyper_ aggressive with bot detection to the point where even my personal head-full chrome instance doesn't always get to load the page. Guess who I will never buy that ~$2000 appliance from.
- compiler-guy 2mo agoThe tradeoff here is a classic false-positive vs false-negative issue. If the cost of the bad-bot false positive (which blocks your bot out) is lower than the cost of the bad-bot false negatives (which allow bad bots in), then it is still a good tradeoff, if a suboptimal situation.
- pwillia7 2mo agoThere is bot traffic initiated by a human and bot traffic not initiated by a human. I would want to serve the first but not the second if it impacted my cost/performance at all.
- lefra 2mo agoIsn't all bot traffic ultimately initiated by a human? Someone plugged the computer in and gave it instructions. It may result in one http request or billions of them, but the human is still the initiator.
- pwillia7 2mo agoYeah but you know what I mean -- A person specifically interacting with my brand vs anthropic hitting all sites 100000 times a day
- Buttons840 2mo agoI think it's time for people to build a local database of every site they've ever visited, and then they can give their LLMs access to that. I might be willing to pass this data off to a company to store for me. Companies already store all my emails and money--why not trust them with this too? Like, all the comments of this post would go into my personal database simply because I loaded the page, and it would help me find old information I've read, and could also inform LLMs I use. This should be built into browsers.
- godwinson__4-8 2mo ago"Bot" access on behalf of users should be fine, even preferred. The missing piece is some micro transaction layer and some sort of attestation somewhere in the layer that the person driving the bot is not a bad actor. Equating bot with bad actor in 2026 is Luddite behavior. Driving people to your site so you can serve them adspam or just make whatever operation they want to do 10x more difficult is the same. MCP style APIs should eat the web. This doesn't mean the "open" web goes away. > 99% of the time I don't care for a domains particular FE at all. It's a complete waste not only of time, but resources and bloat. Pushing the contract into the agent should become good UX. Making things harder for good faith users should never be the goal.
- scotty79 2mo agoAre you gonna police the means through which I'm browsing the internet? If you want to, you are free to put your stuff behind the paywall and give the key only to people who agree to obey your conditions. If you put it in the open you can't make conditions. That's what publishing means. Author can't make demands in what manner their book should be read.
- ted_dunning 2mo agoThe entire blog was about how the site owner doesn't want to put it in the open. Your pontificating can't change the fact that he can do whatever he wants with his site.
- logn 2mo agoWriting a script to fetch HTML is no different than writing a web browser. I think it's the scale of the operation that distinguishes bots vs human. The browser is the user's agent, but not the only one.
- andai 2mo agoA web browser is an entity that acts on your behalf. That's why it's called a user agent. They're just better at English now.
- deleted 2mo ago[deleted]
- 1718627440 2mo agoVery few user write a HTTP request themself. By that measure everyone that uses an UA, is a bot user.