15 ms·
Just to be clear I'm understanding correctly: This is pulling the content of the RSS feeds of several news sites into the context window of an LLM and then ask
by __jonas 1y ago
Just to be clear I'm understanding correctly:
This is pulling the content of the RSS feeds of several news sites into the context window of an LLM and then asking it to summarize news items into articles and fill in the blanks?
I'm asking because that is what it looks like, but AI / LLMs are not specifically mentioned in this blog post, they just say news are 'generated' under the 'News in your language' heading, which seems to imply that is what they are doing.
I'm a little skeptical towards the approach, when you ask an LLM to point to 'sources' for the information it outputs, as far as I know there is no guarantee that those are correct – and it does seem like sometimes they just use pure LLM output, as no sources are cited, or it's quoted as 'common knowledge'.
- devmor 1y agoThe line “news stories will be generated” throws up red flags across the horizon for me. That’s not news. That’s news-adjacent random slop.
- input_sh 1y agoIt's also a workaround around copyright, news sites would be (rightfully) pissed if you publicly post their articles in full and would argue that you're stealing their viewership. But, if you're essentially doing an automatic mash-up of five stories on the same topic from different sources, all of a sudden you're not doing anything wrong! As an example from one of their sources, you can only re-publish a certain amount of words from an article in The Guardian (100 commercially, 500 non-comercially) without paying them.
- nemomarx 1y agotbh I would take the headline and first hundred words in a news aggregator. that seems fine?
- input_sh 1y agoYes, that is fine! That's how RSS feeds usually work when you follow more "mainstream" news sources. At the very least, you see the name of the author and you actually make a connection to their server that can be measured in the analytics. But instead, Kagi "helpfully" regurgitates the whole story, visits the article once, delivers it to presumably thousands, and it can't even be bothered to display all of the sources it regurgitates unless you click to expand the dropdown. And even then the headline itself is one additional click away, and they straight up don't even display the name of the journalist in the pop-up, just the headline. Incredibly shitty behaviour from them. And then they have the balls to start their about page with this: > Why Kagi News? Because news is broken.
- amarant 1y agoAnd yet, after trying it, I have to admit it's more informative and less provocative than any other news source I've seen since at least 2005. I don't know how they do it, and I'm not sure I care, the result is they've eliminated both clickbait and ragebait, and the news are indeed better off for it!
- input_sh 1y agoSoulless, uncreative, not fact-checked (or read by anyone before clicking publish), not contributing anything back to the original journalists, all of the editorial decisions are done by an undeterministic AI filter. Not gonna call it the worst insult to journalism I've ever seen because I've seen factually(.)so which does essentially the same thing but calls it an "AI fact check", but it's not much better. It's like instead of borrowing a book from the library, there's like a spokesperson at the entrance who you ask a question and then blindly believe whatever they say.
- amarant 1y ago>soulless,uncreative This is exactly how I want my news to be. Nothing worse than a headline about a new vaccine breakthrough, followed by a first paragraph that starts with "it was a cold November morning as I arrived in..." I guess it's a matter of taste, but I prefer it short and to the point
- mvieira38 1y agoYes, that's what it is. Kagi as a brand is LLM-optimist, so you may be fundamentally at odds with them here... If it lessens the issue for you, the sources of each item are cited properly in every example I tried, so maybe you could treat it as a fancy link aggregator
- meowface 1y agoI consider myself a major LLM optimist in many ways, but if I'm receiving a once per day curated news aggregation feed I feel I'd want a human eye. I guess an LLM in theory might have less of the biases found in humans, but you're trading one kind of bias for another.
- mvieira38 1y agoYeah, I agree. The entire value/fact dichotomy that the announcement bases itself on is a pretty hot philosophical topic I lean against Kagi on. It's just impossible to summarize any text without imparting some sort of value judgement on it, therefore "biasing" the text
- xpe 1y ago> It's just impossible to summarize any text without imparting some sort of value judgement on it, therefore "biasing" the text Unfortunately, the above is nearly a cliché at this point. The phrase "value judgment" is insufficient because it occludes some important differences. To name just two that matter; there is a key difference between (1) a moral value judgment; (2) selection & summarization (often intended to improve information density for the intended audience). For instance, imagine two non-partisan medical newsletters. Even if they have the same moral values (e.g. rooted in the Hippocratic Oath), they might have different assessments of what is more relevant for their audience. One could say both are "biased", but does doing so impart any functional information? I would rather say something like "Newsletter A is compromised of Editorial Board X with such-and-such a track record and is known for careful, long-form articles" or "Newsletter B is a one-person operation known for a prolific stream of hourly coverage." In this example, saying the newsletters differ in framing and intended audience is useful, but calling each "biased in different ways" is a throwaway comment (having low informational content in the Shannonian sense). Personally, instead of saying "biased" I tend to ask questions like: (a) Who is their intended audience; (b) What attributes and qualities consistently shine through?; (c) How do they make money? (d) Is the publication/source transparent about their approach? (e) What is their track record about accuracy, separating commentary from factual claims, professional integrity, disclosure of conflicts of interest, level of intellectual honesty, epistemic standards, and corrections?
- BeetleB 1y agoIt's not binary - it's a continuum. When you go to Google News, the way they group together stories is AI (pre-LLM technology). Kagi is merely taking it one step further. I agree with your concern. I see this as a convenient grouping, and if any interests me I can skip reading the LLM summary and just click on the sources they provide (making it similar to Google News).
- philipwhiuk 1y ago> Kagi is merely taking it one step further. I would argue creating your own summary is several steps beyond an ordering algorithm.
- krainboltgreene 1y agoIt cannot be "one step further", because there's a clear break in reality between what Google News provides and Kagi provides. Google News links to an article that exists in our world, 100%, no chance involved. Kagi uses an LLM generate text and thus is entirely up to chance.
- b112 1y agoDevil's advocate here. Do you know that's what they're doing? They are a search engine after all. They do run their own indexer, as well as cache results from other sources. If they're feeding urls to an AI, why can't they validate AI output urls are real? Maybe they do.
- krainboltgreene 1y agoI don't care.
- timeon 1y ago> When you go to Google News You don't and you should not use this one either.
- bigstrat2003 1y agoThanks for pointing out that this is yet more AI slop. Very disappointing for Kagi to do this. I get my money's worth from searches, but if I was looking for more features I would want them to be not AI-based.
- atonse 1y agoI'm genuinely asking, but have you tried it? https://kite.kagi.com https://kite.kagi.com It actually seems more like an aggregator (like ground.news) to me. And pretty much every single sentence cites the original article(s). There are nice summaries within an article. I think what they mean is that they generate a meta-article after combining the rest of them. There's nothing novel here. But the presentation of the meta-article and publishing once a day feel like great features.
- __jonas 1y agoI have yeah, to me it looks like what I described in my comment above, it's LLM generated text, is it not? > And pretty much every single sentence cites the original article(s). Yeah but again, correct me if I'm wrong, but I don't think asking an LLM to provide a source / citation yields any guarantee that the text it generates alongside it is accurate. I also see a lot of text without any citations at all, here are three sections (Historical background, Technical details and Scientific significance) that don't cite any sources: https://kite.kagi.com/s/5e6qq2 https://kite.kagi.com/s/5e6qq2
- cogman10 1y agoOof, that one is particularly bad because it cites 3 sources which are all the same article. Google points to phys and phys is a republish of the MIT article.
- 1y ago
- imiric 1y agoI'm firmly on the side of "AI" skepticism, but even I have to admit that this is a very good use of the tech. LLMs generally do a great job at summarizing text, which is essentially what this is. The sources could be statically defined in advance, given that they know where they pull the information from, so I don't think the LLM generates that content. So if this automates the process of fetching the top news from a static list of news sites and summarizing the content in a specific structure, there's not much that can go wrong there. There's a very small chance that the LLM would hallucinate when asked to summarize a relatively short amount of text.
- threetonesun 1y agoWe used to do this with a human created meta tag but I guess this is better?
- input_sh 1y agoIt's useful for the users, but tragically bad for anyone involved with journalism. Not that they're not used to getting fucked by search engines at this point, be it via AMP, instant answers, or AI overviews. Not that the userbase of 50k is big enough to matter right now, but still...
- atonse 1y agoAll this is doing is aggregating RSS feeds and linking to the original articles. So this might result in lower traffic for "anyone involved in journalism" – but the constant doomscrolling is worse for society. So I think we can all agree that the industry needs to veer towards less quantity and more quality.
- input_sh 1y agoRSS feeds are meant to be used by actual users, not regurgitated publicly. RSS readers at the very least have have author info visible and its users tend to be reported to website's analytics with a special user agent.
- 1y ago
- zwnow 1y agoThis also is a really ignorant approach to data poisoning issues in the LLM space. LLMs can easily be misused as propaganda machines...
- doublerebel 1y agoDisappointing. Non-LLM NLP summarization is actually rather good these days. It works by finding the key sentences in the text and extracting the relevant sections, no possibility for hallucination. No need to go full AI for this feature.
- tene80i 1y agoThat’s interesting. Could you share an example or a resource about this?
- raincole 1y agoKagi is probably the only pro-LLM company praised on HN. Perhaps people's hatred towards Google outweighs that of LLM. Imagine if Google news use LLM to show summaries to the users without explicitly saying it's AI on the UI. Ironically, one of the first LLM-induced mistakes experienced by average people was a news summary: https://www.bbc.com/news/articles/cge93de21n0o.amp https://www.bbc.com/news/articles/cge93de21n0o.amp
- JohnFen 1y ago> Kagi is probably the only pro-LLM company praised on HN. Kagi made search useful again, and their genAI stuff can be easily ignored. Best of both worlds -- it remains useful for people like myself who don't want genAI involved, but there's genAI stuff for people who like that sort of thing. That said, if their genAI stuff gets to be too hard to ignore, then I'd stop using or praising Kagi. That this is about news also makes it less problematic for me. I just won't see it at all, since I don't go to Kagi for news in the first place.
- raincole 1y agoI'm not against AI summaries if they are marked as so. Sneakily sliding LLM under the table is a dark pattern no matter how I interpret their intentions. Even Google calls the overview box AI Overview (not saying it doesn't hurt content hosting sites.)
- kiicia 1y agoIt also publishes normal unabridged rss feed that you can read with rss reader of your choice, it looks like great news source
- jama211 1y agoI am fine with it using AI but it makes me feel pretty icky that they didn’t mention that this was ai/llm generated at any point in this article. That’s a no-no IMO, and has turned me off this pretty strongly.
- stavros 1y agoWhy do you care what technology was used to generate the summaries? What if they had used their old NLP summarizer?
- lukeschlather 1y agoThey don't explicitly say they generate summaries at any point in the article. In fact I read it and though this was just some fancy RSS aggregator. The way they describe the "daily briefing" is extremely ambiguous.
- stavros 1y agoOK, but I'd like to repeat my question here: Why do you care how the summary was generated?
- edaemon 1y agoI'm not the person you asked, but it's useful to know if the summary was generated using a method prone to inaccuracy.
- stavros 1y agoThat's all methods, though. Have you seen humans?
- motoxpro 1y agoIn this situation, humans are more accurate, for now, so it's good information to have. Same as I would like to know if humans self assessed in a study about how well they drive vs the empirical evidence. Humans just aren't that good at that task so it would be good to know coming in. Just call it Kagi Vibes instead of Kagi News as news has a higher bar (at least for me)
- whatamidoingyo 1y ago> when you ask an LLM to point to 'sources' for the information it outputs, as far as I know there is no guarantee that those are correct A lot of times when I ask for a source, I get broken links. I'm not sure if the links existed at one point, or if the LLM is just hallucinating where it thinks a link should exist. CDN libraries, for example. Or sources to specific laws.
- CaptainOfCoit 1y ago> A lot of times when I ask for a source, They'll do pretty much everything you ask of them, so unless the text actually come from some source (via tool calls, injecting content into the context or other way), they'll make up a source rather than doing nothing, unless prompted otherwise.
- vbezhenar 1y agoThey could make up source, but ChatGPT is an actual app with complicated backend, not dumb pipe between textedit and GPU. Surely they could verify on server side every link they output to user before including it in the answer. I'm sure Codex will implement it in no time!
- therein 1y agoThey surely can detect it, but what are they going to do after detecting it? Loop the last job with a different seed and hope that the model doesn't lie through its teeth? They won't be doing it because the model will gladly generate you a fake source on the next retry too.
- doikor 1y agoThis is actually harder then most think. The chances of your app doing this check being bot detected/blocked is very high. (unless you are Google etc which are specifically let in to get the article indexed into search)
- jacobgkau 1y agoMaybe they should be trained on the understanding that making up a source is not "doing what you ask of them" when you ask for a source. It's actually the exact opposite of the "doing what you asked, not what you wanted" trope-- it's providing something it thinks you want instead of providing what you asked for (or being honest/erroring out that it can't).
- skysthelimitt 1y agoi believe an llm output is fine for giving an overview if provided the articles, if you want a detailed overview you should be reading the articles anyways.
- j2kun 1y agoThis seems like the opposite of "privacy by design" > Privacy by design: Your reading habits belong to you. We don’t track, profile, or monetize your attention. You remain the customer and not the product. But the person running the LLM surely does.
- m4r71n 1y agoHow would the LLM provider get any information about your reading habits from the app? The LLM is used _before_ the news content is served to you, the reader.
- ivape 1y agoML did not figure out every solution on planet earth. It figured out the LLM, and that is the giant's shoulder most apps will stand on.
- Harmon758 1y agoJust for concrete confirmation that LLM(s) are being used, there's an open issue on the GitHub repository, on hallucinations with made up information, where a Kagi employee specifically mentions "an LLM hallucination problem": https://github.com/kagisearch/kite-public/issues/97 https://github.com/kagisearch/kite-public/issues/97 There's also a line at the bottom of the about page at https://kite.kagi.com/about https://kite.kagi.com/about that says "Summaries may contain errors. Please verify important information."
- jazzyjackson 1y agoLove how it only took 8 years to go from "Fake News!" to "News May Be Fake"
- pjc50 1y agoThere's too much demand for fake news, plenty of subsidy for it, and it's far easier to make. Non fake news is going to be restricted to pay services like Bloomberg terminals.
- Yoric 1y agoAnd paid newspapers, hopefully.
- byearthithatius 1y agoIt is getting easier and easier to fake stuff and there are becoming less and less fully trusted institutions. So sadly I think you are right. Its scary but we are likely heading towards a future where you need to pay to get verified information and that itself will likely be segmented to different subscriptions for what information you want.
- andrewinardeer 1y ago> It is getting easier and easier to fake stuff This is why the moon landing hoax was revolutionary in the 60's. The size of this project was enormous.
- deleted 1y ago[deleted]
- viraptor 1y ago> when you ask an LLM to point to 'sources' for the information it outputs, Services listing sources, like Kagi news, perplexity and others don't do that. They start with known links and run LLMs on that content. They don't ask LLMs to come up with links based on the question.
- __jonas 1y agoThat is what I mean yeah, I’m not saying it’s fabricating sources from training data, that would obviously be impossible for news articles, I’m saying if you give it a list of articles A, B and C including their content in the context and ask ‘what is the foo of bar?’ and it responds ‘the foo of bar is baz, source: article B paragraph 2’, that does not tell you whether the output is actually correct, or contained in the cited source at all, unless you manually verify it.
- Onavo 1y agoYes, they are not the only player here. Quite a few companies are doing this, if you use Perplexity, they also have a news tab with the exact feature set.
- xpe 1y ago> if you use Perplexity, they also have a news tab with the exact feature set "Exact" is far from accurate. I just did a side-by-side comparison. To name only two obvious differences: A. At the top level, Perplexity has a "Discover" tab [1] -- not titled "News". That leads to a AAF page with the endless-scroll anti-pattern (see [2] [3] for other examples). Kagi News [4] presents a short list of ~7ish items without images. B. At the detail-page level, Kagi organizes their content differently (with more detail, including "sources", "highlights", "perspectives", "historical background", and "quick questions"). Perplexity only has content with sources and "discover more". You can verify for yourself. [1]: https://www.perplexity.ai/discover https://www.perplexity.ai/discover [2]: https://www.reddit.com/r/rant/comments/e0a99k/cnn_app_is_annoying_as_fuck_every_article_you_try/ https://www.reddit.com/r/rant/comments/e0a99k/cnn_app_is_ann... [3]: https://www.tumblr.com/make-me-imagine/614701109842444288/anyone-else-annoyed-as-fuck-that-the-infinite https://www.tumblr.com/make-me-imagine/614701109842444288/an... [4]: https://kite.kagi.com https://kite.kagi.com
- rldjbpin 1y agoAfter using Perplexity, its news tab is US-centric, without much options to get regional content from what i can see. Kagi seems to offer regional news and the sources appear to be from the respective area also. do appreciate public access (for now?) with RSS feeds (ironic but handy).
- coffeefirst 1y agoYeah. I really like Kagi. This is a terrible idea. 1. It seems to omit key facts from most stories. 2. No economic value is returned to the sources doing the original reporting. This is not okay. 3. If your summary device makes a mistake, and it will, you are absolutely on the hook for libel. There seem to be some misunderstandings about what news is and what’s makes it well-executed. It’s not the average, it’s the deepest and most accurate reporting. If anyone from the Kagi team wants to discuss, I’m a paying member and I know this field really, really well.
- scosman 1y agoThank you. Also a paying Kagi user because I like the idea that it’s worth it to pay for a good service. Ripping off journalists/newspapers content goes against that.
- jmenter 1y ago> It’s not the average, it’s the deepest and most accurate reporting. Yes! I'm also a paying member but I'm deeply suspicious of this feature. The website claims "we expose readers to the full spectrum of global perspectives", but not all perspectives are equal. It smacks of "all sides" framing which is just not what news ought to be about.
- mediumsmart 1y agoMaybe you should read the article before you assume how it works. It’s pretty clear and AI is specifically mentioned.
- __jonas 1y agoYou’re being presumptuous. I read the article yesterday and there was no mention of AI or LLMs, they have changed it, which is good. https://web.archive.org/web/20250930154005/https://blog.kagi.com/kagi-news https://web.archive.org/web/20250930154005/https://blog.kagi...
- torben-friis 1y agoI just don’t understand what this brings into the picture. Presumably your newspaper of choice already has A) redacted the news in a format that is read friendly B) set up a page with prioritized news Because _that’s what a newspaper is_. What extra value is gotten from a AI rewrite? At best is a borderline noop, at worst a lossy transformation (?)
- raxxorraxor 1y agoI guess they embed the news of the day and let it summarize it. You can add metadata to the training set, which you should technically query reliably. You don't have to let the model do the summarization of the source, which can be erroneous. Far more interesting is how they aggregate the data. I thought many sources moved behind paywalls already.