6 ms·
Have a look at this article: https://www.washingtonpost.com/technology/interactive/2023/ai-chatbot-learning/ https://www.washingtonpost.com/technology/interacti
by tossandthrow 7mo ago
Have a look at this article: https://www.washingtonpost.com/technology/interactive/2023/ai-chatbot-learning/ https://www.washingtonpost.com/technology/interactive/2023/a...
NY Times is 0.06% of common crawl.
These news media outlets provide a drop in the ocean worth of information. Both qualitatively and quantitatively.
The news / media industry is really just trying to hold on to their lifeboat before inevitably becoming entirely irrelevant.
(I do find this sad, but it is like the reality - I can already now get considerably better journalism using LLMs than actual journalists - both click bait stuff and high quality stuff)
- pimlottc 7mo agoThat seems like a reductive way to consider it. What percent of music was created by Led Zeppelin? What percent of art was painted by Monet? What percent of films by Alfred Hitchcock? It may be a small percentage objectively but they are hugely influential.
- tossandthrow 7mo agoI don't think back propagation care whose text it is back propagating.
- NiloCK 7mo agoThe data sets aren't naively fed into the training runs. Instead, training attempts to sample more heavily from higher quality sources, with, I'm sure, a mix of manual and heuristic labeling.
- ffsm8 7mo agofwiw, no llm ive ever used generated in the writing style newspapers and -sites use - hence i honestly doubt they've been given a meaningful boost in relevancy. their idioms would leak occasionally otherwise
- Gigachad 7mo ago90% of common crawl is complete junk. While the tiny bit of news articles powers almost all the ai answers in Google search.
- Dylan16807 7mo agoNews takes a very different path to get into search results. It's not going through databases or archive passes, that would take far too long. And don't basically all those news sites allow google on purpose?
- datsci_est_2015 7mo agoHow many Reddit, HN, etc. posts are based on NYT articles? How many derivative news articles, blog posts, YouTube videos, TikToks, etc. are responses to those articles? At least NYT is probably on the correct side of Sturgeon’s Law: https://en.wikipedia.org/wiki/Sturgeon%27s_law https://en.wikipedia.org/wiki/Sturgeon%27s_law
- AnthonyMouse 7mo ago> How many Reddit, HN, etc. posts are based on NYT articles? How many derivative news articles, blog posts, YouTube videos, TikToks, etc. are responses to those articles? You may get an inconvenient answer when you ask the question the other way around.
- Melatonic 7mo ago0.06% is way higher than I would expect