12 ms·
Auto-grading decade-old Hacker News discussions with hindsight
Related from yesterday: Show HN: Gemini Pro 3 imagines the HN front page 10 years from now - https://news.ycombinator.com/item?id=46205632 https://news.ycombinator.com/item?id=46205632
- bediger4000 10mo agoLLMs are watching (or humans using them might be). Best to be good. Shades of Roko's Basilisk!
- ambicapter 10mo agoMore like a Panopticon. As the parenthesis notes, this is just as bad when humans are the final link in the eyeball chain.
- deleted 10mo ago[deleted]
- gen6acd60af 10mo agoCommenters of HN: Your past thoughts have been dredged up and judged. For each $TOPIC, you have been awarded a grade by GPT-5.1 Thinking. Your grade is based on OpenAI's aligned worldview and what OpenAI's blob of weights considers Truth in 2025. Did you think well, netizen? Are you an Alpha or a Delta-Minus? Where will the dragnet grading of your online history happen next?
- HighGoldstein 10mo agoOf all the people on the entire internet, I would hope HN posters understand best that anything and everything posted online already has and also will at some point be used in such ways.
- siliconc0w 10mo agoRandom Bets for 2035: * Nvidia GPUs will see heavy competition and most chat-like use-cases switching to cheaper models and inference-specific-silicon but will be still used on the high end for critical applications and frontier science * Most Software and UIs will be primarily AI-generated. There will be no 'App Stores' as we know them. * ICE Cars will become niche and will be largely been replaced with EVs, Solar will be widely deployed and will be the dominate source of power * Climate Change will be widely recognized due to escalating consequences and there will be lots of efforts in mitigations (e.g, Climate Engineering, Climate-resistant crops, etc).
- xattt 10mo agoYou’re about 20 days short or 345 days late for this HN tradition. ;)
- pu_pe 10mo agoThe infamous Dropbox comment might turn out to be right in 10 more years, when LLMs might just build an entire application from scratch for you.
- deleted 10mo ago[deleted]
- rafaelmn 10mo agoI'd take the other side for most of these - Nvidia one is too vague (some could argue it's already seeing "heavy competition" from Google and other players in the space) but something more concrete - I doubt they will fall below 50% market share.
- jasonthorsness 10mo agoIt's fun to read some of these historic comments! A while back I wrote a replay system to better capture how discussions evolved at the time of these historic threads. Here's Karpathy's list from his graded articles, in the replay visualizer: Swift is Open Source https://hn.unlurker.com/replay?item=10669891 https://hn.unlurker.com/replay?item=10669891 Launch of Figma, a collaborative interface design tool https://hn.unlurker.com/replay?item=10685407 https://hn.unlurker.com/replay?item=10685407 Introducing OpenAI https://hn.unlurker.com/replay?item=10720176 https://hn.unlurker.com/replay?item=10720176 The first person to hack the iPhone is building a self-driving car https://hn.unlurker.com/replay?item=10744206 https://hn.unlurker.com/replay?item=10744206 SpaceX launch webcast: Orbcomm-2 Mission [video] https://hn.unlurker.com/replay?item=10774865 https://hn.unlurker.com/replay?item=10774865 At Theranos, Many Strategies and Snags https://hn.unlurker.com/replay?item=10799261 https://hn.unlurker.com/replay?item=10799261
- HanClinto 10mo agoOkay, your site is a ton of fun. Thank you! :)
- SauntSolaire 10mo agoI'd love to see sentiment analysis done based on time of day. I'm sure it's largely time zone differences, but I see a large variance in the types of opinions posted to hn in the morning versus the evening and I'd be curious to see it quantified.
- embedding-shape 10mo agoYeah, I see this constantly any time Europe is mentioned in a submission. Early European morning/day, regular discussions, but as the European afternoon/evening comes around, you start noticing a lot anti-union sentiment, discussions start to shift into over-regulation, and the typical boring anti-Europe/EU talking points.
- nostrebored 10mo ago“Regular” to who? Pro EU sentiment almost only comes from the EU, which is what you’re observing. Pro-US sentiment is relatively mixed (as is anti-US sentiment) in distribution.
- moultano 10mo agoNotable how this is only possible because the website is a good "web citizen." It has urls that maintain their state over a decade. They contain a whole conversation. You don't have to log in to see anything. The value of old proper websites increases with our ability to process them.
- chrisweekly 10mo agoYes! See "Cool URIs Don't Change"^1 by Sir TBL himself. 1. https://www.w3.org/Provider/Style/URI https://www.w3.org/Provider/Style/URI
- jeffbee 10mo agoThere are things that you have to log in to see, and the mods sometimes move conversations from one place to another, and also, for some reason, whole conversations get reset to a single timestamp.
- latexr 10mo ago> for some reason, whole conversations get reset to a single timestamp. What do you mean?
- jeffbee 10mo agoThere is some action that moderators can take that throws one of yesterday's articles back on the front page and when that happens all the comments have the same timestamp.
- consumer451 10mo agoI believe that this is called "the second chance pool." It is a bit strange when it unexpectedly happens to one's own post.
- embedding-shape 10mo agoSubmissions put in the second-chance pool briefly appear (sometimes "again") on the frontpage, and the conversation timestamps are reset so it appears like they were written after the second-chance submission, not before.
- exasperaited 10mo ago> Everything we do today might be scrutinized in great detail in the future because it will be "free". s/"free"/stolen/ The bit about college courses for future prediction was just silly, I'm afraid: reminds me of how Conan Doyle has Sherlock not knowing Earth revolves around the Sun. Almost all serious study concerns itself with predicting, modelling and influence over the future behaviour of some system; the problem is only that people don't fucking listen to the predictions of experts. They aren't going to value refined, academic general-purpose futurology any more than they have in the past; it's not even a new area of study.
- GaggiX 10mo agoI think the most fun thing is to go to: https://karpathy.ai/hncapsule/hall-of-fame.html https://karpathy.ai/hncapsule/hall-of-fame.html And scroll down to the bottom.
- MBCook 10mo agoIt’s interesting, if you go down near the bottom you see some people with both A’s and D’s. According to the ratings for example, one person both had extremely racist ideas but also made a couple of accurate points about how some tech concepts would evolve.
- brian_spiering 10mo agoThat is interesting because of the Halo effect. There is a cognitive bias that if a person is right in one area, they will be right in another unrelated area. I try to temper my tendency to believe the Halo effect with Warren Buffett's notion of the Circle of Competence; there is often a very narrow domain where any person can be significantly knowledgeable.
- deleted 10mo ago[deleted]
- xpe 10mo ago> A circle of competence is the subject area which matches a person's skills or expertise. The concept was developed by Warren Buffett and Charlie Munger as what they call a mental model, a codified form of business acumen, concerning the investment strategy of limiting one's financial investments in areas where an individual may have limited understanding or experience, while concentrating in areas where one has the greatest familiarity. -Wikipedia > I try to temper my tendency to believe the Halo effect with Warren Buffett's notion of the Circle of Competence; there is often a very narrow domain where any person can be significantly knowledgeable. (commenter above) Putting aside Buffett in particular, I'm wary of claims like "there is often a very narrow domain where any person can be significantly knowledgeable". How often? How narrow of a domain? Doesn't it depend on arbitrary definitions of what qualifies as a category? Is this a testable theory? Is it a predictive theory? What does empirical research and careful analysis show? Putting that aside, there are useful mathematical ways to get an idea of some of the backing concepts without making assumptions about people, culture, education, etc. I'll cook one up now... Start with 70K balls split evenly across seven colors: red, orange, yellow, green, blue, indigo, and violet. 1,000 show up demanding balls. So we mix them up and randomly distribute 10 balls to every person. What does the distribution tend to look like? What particulars would you tune and/or definitions would you choose to make this problem "sort of" map to something sort of like assessing the diversity of human competence across different areas? Note the colored balls example assumes independence between colors (subjects or skills or something). But in real life, there are often causally significant links between skills. For example, general reasoning ability improves performance in a lot of other subjects. Then a goat exploded, because I don't know how to end this comment gracefully.
- MBCook 10mo ago#272, I got a B+! Neat. It would be very interesting to see this applied year after year to see if people get better or worse over time in the accuracy of their judgments. It would also be interesting to correlate accuracy to scores, but I kind of doubt that can be done. Between just expressing popular sentiment and the first to the post people getting more votes for the same comment than people who come later it probably wouldn’t be very useful data.
- pjc50 10mo ago#250, but then I wasn't trying to make predictions for a future AI. Or anyone else, really. Got a high score mostly for status quo bias, e.g. visual languages going nowhere and FPGAs remain niche.
- embedding-shape 10mo agoYeah, it be much more interesting to see the people who made (at the time) outrageous claims, but they came to be true, rather than a list of people who could state that the status quo most likely would stay as it is.
- swalsh 10mo agoI have never felt less confident in the future than I do in 2025... and it's such a stark contrast. I guess if you split things down the middle, AI probably continues to change the world in dramatic ways but not in the all or nothing way people expect. A non trivial amount of people get laid off, likely due to a finanical crisis which is used as an excuse for companies scale up use of AI. Good chance the financial crisis was partly caused by AI companies, which ironically makes AI cheaper as infra is bought up on the cheap (so there is a consolidation, but the bountiful infra keeps things cheap). That results in increased usage (over a longer period of time). and even when the economy starts coming back the jobs numbers stay abismal. Politics are divided into 2 main groups, those who are employed, and those who are retired. The retired group is VERY large, and has alot of power. They mostly care about entitlements. The employed age people focus on AI which is making the job market quite tough. There are 3 large political forces (but 2 parties). The Left, the Right, and the Tech Elite. The left and the right both hate AI, but the tech elite though a minority has outsized power in their tie breaker role. The age distributions would surprise most. Most older people are now on the left, and most younger people are split by gender. The right focuses on limiting entitlements, and the left focuses on growing them by taxing the tech elite. The right maintains power by not threatening the tech elite. Unlike the 20th century America is a more focused global agenda. We're not policing everyone, just those core trading powers. We have not gone to war with China, China has not taken over Taiwan. Physical robotics is becoming a pretty big thing, space travel is becoming cheaper. We have at least one robot on an astroid mining it. The yield is trivial, but we all thought it was neat. Energy is much much greener, and you wouln't have guessed it... but it was the data centers that got us there. The Tech elite needed it quickly, and used the political connections to cut red tape and build really quickly.
- 1121redblackgo 10mo agoWe do not currently have the political apparatus in place to stop the dystopian nightmares depicted in movies and media. They were supposed to be cautionary tales. Maybe they still can be, but there are basically zero guardrails in non-progressive forms of government to prevent massive accumulations of power being wielded in ways most of the population disapproves of.
- bgwalter 10mo ago"If LLMs are watching, humans will be on their best behavior". Karpathy, paraphrasing Larry Ellison. The EU may give LLM surveillance an F at some point.
- lapcat 10mo agoDoes anyone else think that HN engages in far too much navel-gazing? Nothing gets upvotes faster than a HN submission about HN.
- yellow_lead 10mo agoIt's weird that HN viewers are interested in HN
- CamperBob2 10mo agoAs moultano suggests, this is likely because most other websites make it completely impossible to navel-gaze. We can't possibly give the HN admins too much praise and credit for their commitment to open and stable availability of legacy data.
- deleted 10mo ago[deleted]
- dang 10mo agoIt's true that meta is the crack of internet forums, so we, er, crack down on it quite a bit. That's a longstanding view: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&query=by%3Adang%20meta%20crack&sort=byDate&type=comment https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu... Alternate metaphor: evil catnip - https://hn.algolia.com/?dateRange=all&page=0&prefix=true&query=by%3Adang%20meta%20catnip&sort=byDate&type=comment https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que... But yesterday's thread and this one are clearly exceptions—far above the median. https://news.ycombinator.com/item?id=46212180 https://news.ycombinator.com/item?id=46212180 was particularly incredible I think!
- latexr 10mo agoI love it when you share some insight about HN or internet communication for which you have relevant searches at the ready to explanations of the concept. A personal favourite is “the contrarian dynamic”. Do you have a list of those at the ready or do you just remember them? If you feel like sharing, what’s your process and is there a list of those you’d make public? I imagine having one would be useful, e.g. for onboarding someone like tomhow, though that doesn’t really happen often.
- artur44 10mo agoInteresting experiment. Using modern LLMs to retroactively grade decade-old HN discussions is a clever way to measure how well our collective predictions age. It’s impressive how little time and compute it now takes to analyze something that would’ve required days of manual reading. My only caution is that hindsight grading can overvalue outcomes instead of reasoning — good reasoning can still lead to wrong predictions. But as a tool for calibrating forecasting and identifying real signal in discussions, this is a very cool direction.
- collinmcnulty 10mo ago> But if intelligence really does become too cheap to meter, it will become possible to do a perfect reconstruction and synthesis of everything. LLMs are watching (or humans using them might be). Best to be good. I cannot believe this is just put out there unexamined of any level of "maybe we shouldn't help this happen". This is complete moral abdication. And to be clear, being "good" is no defense. Being good often means being unaligned with the powerful, so being good is often the very thing that puts you in danger.
- Teever 10mo agoThe time for discussion and action on this was over a 15 years ago when Snowden and the NSA with their Utah data centre was a big story. Governments around the world have profiles on people and spiders that quietly amass the data that continuously updates those profiles. It's just a matter of time before hardware improves and we see another holocaust scale purge facilitated by robots. Surveillance capitalism won.
- doctoboggan 10mo agoI've had the same though as Karpathy over the past couple of months/years. I don't think it's good, exciting, or something to celebrate, but I also have no idea how to prevent it. I would read his "Best to be good." as a warning or reminder that everything you do or say online will be collected and analyzed by an "intelligence". You can't count on hiding amongst the mass of online noise. Imagine if someone were to collect everything you've written or uploaded to the internet and compiled it into a long document. What sort of story would that tell about who you are? What would a clever person (or LLM) be able to do with that document? If you have any ideas on how to stop everyone from building the torment nexus, I am willing to listen.
- karpathy 10mo agoThank you
- collinmcnulty 10mo agoThis is my plan at least 1. Don't build the Torment Nexus yourself. Don't work for them and don't give them your money. 2. When people you know say they're taking a new job to work at Torment Nexus, act like that's super weird, like they said they're going to work for the Sinaloa cartel. Treat rich people working on the Torment Nexus like it's cringe to quote them. 3. Get hostile to bots. Poison the data. Use AdNauseum and Anubis. 4. Give your non-tech friends the vague sense that this stuff is bad. Some might want to listen more, but most just take their sense of what's cool and good from people they trust in the area.
- gaigalas 10mo agoI am not sure if we need a karma precog analogue. It does seem better than just upvotes and downvotes though.
- jeffbee 10mo agoI'm delighted to see that one of the users who makes the same negative comments on every Google-related post gets a "D" for saying Waymo was smoke and mirrors. Never change, I guess.
- modeless 10mo agoThis is a cool idea. I would install a Chrome extension that shows a score by every username on this site grading how well their expressed opinions match what subsequently happened in reality, or the accuracy of any specific predictions they've made. Some people's opinions are closer to reality than others and it's not always correlated with upvotes. An extension of this would be to grade people on the accuracy of the comments they upvote, and use that to weight their upvotes more in ranking. I would love to read a version of HN where the only upvotes that matter are from people who agree with opinions that turn out to be correct. Of course, only HN could implement this since upvotes are private.
- cootsnuck 10mo agoThe RES (Reddit Enhancement Suite) browser extension indirectly does this for me since it tracks the lifetime number of upvotes I give other users. So when I stumble upon a thread with a user with like +40 I know "This is someone whom I've repeatedly found to have good takes" (depending on the context). It's subjective of course but at least it's transparently so. I just think it's neat that it's kinda sorta a loose proxy for what you're talking about but done in arguably the simplest way possible.
- nickff 10mo agoI am not a Redditor, but RES sounds like it would increase the ‘echo-chamber’ effect, rather than improving one’s understanding of contributors’ calibration.
- mistercheph 10mo agoit depends on if you vote based on the quality of contribution to the discussion or based on how much you agree/disagree.
- miki123211 10mo agoI don't think you can change user behavior like this. You can give them a "venting sink" though. Instead of having a downvote button that just downvotes, have it pop up a little menu asking for a downvote reason, with "spam" and "disagree" as options. You could then weigh downvotes by which option was selected, along with an algorithm to discover "user honesty" based on whether their downvotes correlate with others or just with the people on their end of the political spectrum, a la Birdwatch.
- GaggiX 10mo agoI was reading the Anki article on 2015-12-13, and the best prediction was by markm248 saying: "Remember that you read it here first, there will be a unicorn built on the concept of SRS" They were right, Duolingo.
- Bjartr 10mo agoNeat, I got a shout-out. Always happy to share the random stuff I remember exists!
- mvdtnz 10mo agoDo we need more AI slop on the front page?
- hackthemack 10mo agoI noticed the Hall of Fame grading of predictive comments has a quirk? It grades some comments about if they came true or not, but in the grading of comment to the article https://news.ycombinator.com/item?id=10654216 https://news.ycombinator.com/item?id=10654216 The Cannons on the B-29 Bomber "accurate account of LeMay stripping turrets and shifting to incendiary area bombing; matches mainstream history" It gave a good grade to user cstross but to my reading of the comment, cstross just recounted a bit of old history. The evaluation gave cstross for just giving a history lesson or no?
- karpathy 10mo agoYes I noticed a few of these around. The LLM is a little too willing to give out grades for comments that were good/bad in a bit more general sense, even if they weren't making strong predictions specifically. Another thing I noticed is that the LLM has a very impressive recognition of the various usernames and who they belong to, and I think shows a little bit of a bias in its evaluations based on the identity of the person. I tuned the prompt a little bit based on some low-hanging fruit mistakes but I think one can most likely iterate it quite a bit further.
- patcon 10mo agoI think you were getting at this, but in case others didn't know: cstross is a famous sci-fi author and futurist :)
- slg 10mo agoThis is a perfect example of the power and problems with LLMs. I took the narcissistic approach of searching for myself. Here's a grade of one of my comments[1]: >slg: B- (accurate characterization of PH’s “networking & facade” feel, but implicitly underestimates how long that model can persist) And here's the actual comment I made[2]: >And maybe it is the cynical contrarian in me, but I think the "real world" aspect of Product Hunt it what turned me off of the site before these issues even came to the forefront. It always seemed like an echo chamber were everyone was putting up a facade. Users seemed more concerned with the people behind products and networking with them than actually offering opinions of what was posted. >I find the more internet-like communities more natural. Sure, the top comment on a Show HN is often a critique. However I find that more interesting than the usual "Wow, another great product from John Developer. Signing up now." or the "Wow, great product. Here is why you should use the competing product that I work on." that you usually see on Product Hunt. I did not say nor imply anything about "how long that model can persist", I just said I personally don't like using the site. It's a total hallucination to claim I was implying doom for "that model" and you would only know that if you actually took the time to dig into the details of what was actually said, but the summary seems plausible enough that most people never would. The LLM processed and analyzed a huge amount of data in a way that no human could, but the single in-depth look I took at that analysis was somewhere between misleading and flat out wrong. As I said, a perfect example of what LLMs do. And yes, I do recognize the funny coincidence that I'm now doing the exact thing I described as the typical HN comment a decade ago. I guess there is a reason old me said "I find that more interesting". [1] - https://karpathy.ai/hncapsule/2015-12-18/index.html#article-10759879 https://karpathy.ai/hncapsule/2015-12-18/index.html#article-... [2] - https://news.ycombinator.com/item?id=10761980 https://news.ycombinator.com/item?id=10761980
- npunt 10mo agoI'm not so sure; that may not have been what you meant, but that doesn't mean it's not what others read into it. The broader context is HN is a startup forum and one of the most common discussion patterns is 'I don't like it' that is often a stand-in for 'I don't think it's viable as-is'. Startups are default dead, after all. With that context, if someone were to read your comment and be asked 'does this person think the product's model is viable in the long run' I think a lot of people would respond 'no'.
- neilv 10mo ago> I spent a few hours browsing around and found it to be very interesting. This seems to be the result of the exercise? No evaluation? My concern is that, even if the exercise is only an amusing curiosity, many people will take the results more seriously than they should, and be inspired to apply the same methods to products and initiatives that adversely affect people's lives in real ways.
- cootsnuck 10mo ago> My concern is that, even if the exercise is only an amusing curiosity, many people will take the results more seriously than they should, and be inspired to apply the same methods to products and initiatives that adversely affect people's lives in real ways. That will most definitely happen. We already have known for awhile that algorithmic methods have been applied "to products and initiatives that adversely affect people's lives in real ways", for awhile: https://www.scientificamerican.com/blog/roots-of-unity/review-weapons-of-math-destruction/ https://www.scientificamerican.com/blog/roots-of-unity/revie... I guess the question is if LLMs for some reason will reinvigorate public sentiment / pressure for governing bodies to sincerely take up the ongoing responsibility of trying to lessen the unique harms that can be amplified by reckless implementation of algorithms.
- btbuildem 10mo agoI've spent a weekend making something similar for my gmail account (which google keeps nagging me about being 90% full). It's fascinating to be able to classify 65k+ of emails (surprise: more than half are garbage), as well as summarize and trace the nature of communication between specific senders/recipients. It took about 50 hours on a dual RTX 3090 running Qwen 3. My original goal was to prune the account deleting all the useless things and keeping just the unique, personal, valuable communications -- but the other day, an insight has me convinced that the safer / smarter thing to do in the current landscape is the opposite: remove any personal, valuable, memorable items, and leave google (and whomever else is scraping these repositories) with useless flotsam of newsletters, updates, subscription receipts, etc.
- red-iron-pine 10mo agoso then what do you do with the useful stuff?
- btbuildem 10mo agoLocal archive + client for search
- subscriptzero 10mo agoI would love to do something like this, and weirdly I even have a dual 3090 home setup. Any chance you can outline the steps/prompts/tools you used to run this? I've been building a 2nd brain type project, that plugs into all my work places and a custom classifier has been on that list that would enhance that.
- 0xWTF 10mo agoNow: compared to what? Is there a better source than HN? How's it compare to Reddit or lobsters? Compared to what happens next? Does tptacek's commentary become market signal equivalent to the Fed Chair or the BLS labor and inflation reports?
- scosman 10mo agoAnyone have a branch that I can run to target my own comments? I'd love to see where I was right and where I was off base. Seems like a genuinely great way to learn about my own biases.
- xpe 10mo agoI appreciate your intent, but this tool needs a lot of work -- maybe an entire redesign -- before it would be suitable for the purpose you seek. See discussion at [1]. Besides, in my experience, only a tiny fraction of HN comments can be interpreted as falsifiable predictions. Instead I would recommend learning about calibration [2] and ways to improve one's calibration, which will likely lead you into literature reviews of cognitive biases and what we can do about them. Also, jumping into some prediction markets (as long as they don't become too much of a distraction) is good practice. [1]: https://news.ycombinator.com/item?id=46223959 https://news.ycombinator.com/item?id=46223959 [2]: https://www.lesswrong.com/w/calibration https://www.lesswrong.com/w/calibration
- tptacek 10mo ago'pcwalton, I'm coming for you. You're going down. Kidding aside, the comments it picks out for us are a little random. For instance, this was an A+ predictive thread (it appears to be rating threads and not individual comments): https://news.ycombinator.com/item?id=10703512 https://news.ycombinator.com/item?id=10703512 But there's just 11 comments, only 1 for me, and it's like a 1-sentence comment. I do love that my unaccredited-access-to-startup-shares take is on that leaderboard, though.
- kbenson 10mo agoI noticed from reviewing my own entry (which honestly I'm surprised exists) that the idea of what it thinks constitutes a "prediction" is fairly open to interpretation, or at least that adding some nuance to a small aspect in a thread to someone else prediction counts quite heavily. I don't really view how I've participated here over the years in any way as making predictions. I actually thought I had done a fairly good job at not making predictions, by design.
- n4r9 10mo agoYeah, I'm having to pinch myself a little here. Another slightly odd example it picked out from your history: https://news.ycombinator.com/item?id=10735398 https://news.ycombinator.com/item?id=10735398 It's a good comment, but "prescient" isn't a word I'd apply to it. This is more like a list of solid takes. To be fair there probably aren't even that many explicit, correct predictions in one month of comments in 2015.
- mvkel 10mo agoHilariously, it seems you anticipated this happening and copyrighted your comments. Is karpathy's tool in violation of your copyright?!
- tptacek 10mo agoKarpathy, I'm coming for you next.
- mistercheph 10mo agoA majority don't seem to be predictions about the future, and it seems to mostly like comments that give extended air to what was then and now the consensus viewpoint, e.g. the top comment from pcwalton the highest scored user: https://news.ycombinator.com/item?id=10657401 https://news.ycombinator.com/item?id=10657401 > (Copying my comment here from Reddit /r/rust:) Just to repeat, because this was somewhat buried in the article: Servo is now a multiprocess browser, using the gaol crate for sandboxing. This adds (a) an extra layer of defense against remote code execution vulnerabilities beyond that which the Rust safety features provide; (b) a safety net in case Servo code is tricked into performing insecure actions. There are still plenty of bugs to shake out, but this is a major milestone in the project.
- Rperry2174 10mo agoOne thing this really highlights to me is how often the "boring" takes end up being the most accurate. The provocative, high-energy threads are usually the ones that age the worst. If an LLM were acting as a kind of historian revisiting today’s debates with future context, I’d bet it would see the same pattern again and again: the sober, incremental claims quietly hold up, while the hyperconfident ones collapse. Something like "Lithium-ion battery pack prices fall to $108/kWh" is classic cost-curve progress. Boring, steady, and historically extremely reliable over long horizons. Probably one of the most likely headlines today to age correctly, even if it gets little attention. On the flip side, stuff like "New benchmark shows top LLMs struggle in real mental health care" feels like high-risk framing. Benchmarks rotate constantly, and “struggle” headlines almost always age badly as models jump whole generations. I bet theres many "boring but right" takes we overlook today and I wondr if there's a practical way to surface them before hindsight does
- simianparrot 10mo agoInstead of "LLM's will put developers out of jobs" the boring reality is going to be "LLM's are a useful tool with limited use".
- jimbokun 10mo agoThat is at odds with predicting based on recent rates of progress.
- yunwal 10mo ago"Boring but right" generally means that this prediction is already priced in to our current understanding of the world though. Anyone can reliably predict "the sun will rise tomorrow", but I'm not giving them high marks for that.
- SubiculumCode 10mo agoPerhaps a new category, 'highest risk guess but right the most often'. Those is the high impact predictions.
- karmickoala 10mo agoI understand the exercise, but I think it should have a disclaimer, some of the LLM reviews are showing a bias and when I read the comments they turned out not to be as bad as the LLM made them. As this hits the front page, some people will only read the title and not the accompanying blog post, losing all of the nuance. That said, I understand the concept and love what you did here. By this being exposed to the best disinfectant, I hope it will raise awareness and show how people and corporations should be careful about its usage. Now this tech is accessible to anyone, not only big techs, in a couple of hours. It also shows how we should take with a grain of salt the result of any analysis of such scale by a LLM. Our private channels now and messages on software like Teams and Slack can be analyzed to hell by our AI overlords. I'm probably going to remove a lot of things from cloud drives just in case. Perhaps online discourse will deteriorate to more inane / LinkedIn style content. Also, I like that your prompt itself has some purposefully leaked bias, which shows other risks—¹for instance, "fsflover: F", which may align the LLM to grade worse the handles that are related to free software and open source). As a meta concept of this, I wonder how I'll be graded by our AI overlords in the future now that I have posted something dismissive of it. ¹Alt+0151
- smugma 10mo agoI believe that the GPA calculation is off, maybe just for F's. I scrolled to the bottom of the hall of fame/shame and saw that entry #1505 and 3 F's and a D, with an average grade of D+ (1.46). No grade better than a D shouldn't average to a D+, I'd expect it to be closer to a 0.25.
- deleted 10mo ago[deleted]
- deleted 10mo ago[deleted]
- ComputerGuru 10mo agoLooking at the results and the prompt, I would tweak the prompt to * ignore comments that do not speculate on something that was unknown or had not achieved consensus as of the date of yyyy-mm-dd * at the same time, exclude speculations for which there still isn’t a definitive answer or consensus today * ignore comments that speculate on minor details or are stating a preference/opinion on a subjective matter * it is ok to generate an empty list of users for a thread if there are no comments meeting the speculation requirements laid out above * etc
- janalsncm 10mo agoYou would also need to exclude “predictions” for things which already happened at the time they were predicted.
- xpe 10mo agoGood points. To summarize: for a given comment, one presumably must downselect to the ones that can reasonably be interpreted as forecasts. I see some indicators that the creator of the project (despite his amazing reputation) skated over this part.
- losvedir 10mo agoAgreed. I feel like it's more just a collection of good comments. It doesn't surprise me to see tptacek, patio11, etc there. I think the "prediction" aspect is under weighted. But it reminds me that I miss Manishearth's comments! What ever happened to him? I recall him being a big rust contributor. I'd think he'd be all over the place, with rust's adoption since then. I also liked tokenadult. interesting blast from the past.
- godelski 10mo ago> I was reminded again of my tweets that said "Be good, future LLMs are watching". You can take that in many directions, but here I want to focus on the idea that future LLMs are watching. Everything we do today might be scrutinized in great detail in the future because doing so will be "free". A lot of the ways people behave currently I think make an implicit "security by obscurity" assumption. But if intelligence really does become too cheap to meter, it will become possible to do a perfect reconstruction and synthesis of everything. LLMs are watching (or humans using them might be). Best to be good. Can we take a second and talk about how dystopian this is? Such an outcome is not inevitable, it relies on us making it. The future is not deterministic, the future is determined by us. Moreso, Karpathy has significantly more influence on that future than your average HN user. We are doing something very *very* wrong if we are operating under the belief that this future is unavoidable. That future is simply unacceptable.
- jacquesm 10mo agoGiven the quality of the judgment I'm not worried, there is no value here. To properly execute this idea rather than to just toss it off without putting in the work to make it valuable is exactly what irritates me about a lot of AI work. You can be 900 times as productive at producing mental popcorn, but if there was value to be had here we're not getting it, just a whiff of it. Sure, fun project. But I don't feel particularly judged here. The funniest bit is the judgment on things that clearly could not yet have come to pass (for instance because there is an exact date mentioned that we have not yet reached). QA could be better.
- godelski 10mo agoI think you're missing the actual problem. I'm not worried about this project but instead harvesting, analyzing all that data and deanonymizing people. That's exactly what Karparthy is saying. He's not being shy about it. He said "behave because the future panopticon can look into the past". Which makes the panopticon effectively exist now. Be good, future LLMs are watching ... or humans using them might be That's the problem. Not the accuracy of this toy project, but the idea of monitoring everyone and their entire history. The idea that we have to behave as if we're being actively watched by the government is literally the setting of 1984 lol. The idea that we have to behave that way now because a future government will use the Panopticon to look into the past is absolutely unhinged. You don't even know what the rules of that world will be! Did we forget how unhinged the NSA's "harvest now, decrypt later" strategy is? Did we forget those giant data centers that were all the news talked about for a few weeks? That's not the future I want to create, is it the one you want? To act as if that future is unavoidable is a failure of *us*
- dschnurr 10mo agoNice! Something must be in the air – last week I built a very similar project using the historical archive of all-in podcast episodes: https://allin-predictions.pages.dev/ https://allin-predictions.pages.dev/
- sanex 10mo agoI'll use this as evidence supporting my continued demand for a Friedberg only spinoff.
- LeroyRaz 10mo agoI am surprised the author thought the project passed quality control. The LLM reviews seem mostly false. Looking at the comment reviews on the actual website, the LLM seems to have mostly judged whether it agreed with the takes, not whether they came true, and it seems to have an incredibly poor grasp of it's actual task of accessing whether the comments were predictive or not. The LLM's comment reviews are of often statements like "correctly characterized [program language] as [opinion]." This dynamic means the website mostly grades people on having the most confirmist take (the take most likely to dominate the training data, and be selected for in the LLM RL tuning process of pleasing the average user).
- hathawsh 10mo agoAre you sure? The third section of each review lists the “Most prescient” and “Most wrong” comments. That sounds exactly like what you're looking for. For example, on the "Kickstarter is Debt" article, here is the LLM's analysis of the most prescient comment. The analysis seems accurate and helpful to me. https://karpathy.ai/hncapsule/2015-12-03/index.html#article-10667041 https://karpathy.ai/hncapsule/2015-12-03/index.html#article-... phire > “Oculus might end up being the most successful product/company to be kickstarted… > Product wise, Pebble is the most successful so far… Right now they are up to major version 4 of their product. Long term, I don't think they will be more successful than Oculus.” With hindsight: Oculus became the backbone of Meta’s VR push, spawning the Rift/Quest series and a multi‑billion‑dollar strategic bet. Pebble, despite early success, was shut down and absorbed by Fitbit barely a year after this thread. That’s an excellent call on the relative trajectories of the two flagship Kickstarter hardware companies.
- jacquesm 10mo agoPredictions are only valuable when they're actually made ahead of the knowledge becoming available. A man will walk on mars by 2030 is falsifiable, a man will walk on mars is not. A lot of these entries have very low to no predictive value or were already known at the time, but just related. Would be nice if future 'judges' put in more work to ensure quality judgments. I would grade this article B-, but then again, nobody wrote it... ;)
- SequoiaHope 10mo agoThis is great! Now I want to run this to analyze my own comments and see how I score and whether my rhetoric has improved in quality/accuracy over time!
- sigmar 10mo agoGotta auto grade every HN comment for how good it is at predicting stock market movement then check what the "most frequently correct" user is saying about the next 6 months.
- Rychard 10mo agoAs the saying goes, "past performance is not indicative of future results"
- xpe 10mo agoI hope this is a joke. Forecasting and the meta-analysis of forecasters is fairly well studied. [1] is a good place to start. [1]: https://en.wikipedia.org/wiki/Superforecaster https://en.wikipedia.org/wiki/Superforecaster
- sigmar 10mo ago> The conclusion was that superforecasters' ability to filter out "noise" played a more significant role in improving accuracy than bias reduction or the efficient extraction of information. >In February 2023, Superforecasters made better forecasts than readers of the Financial Times on eight out of nine questions that were resolved at the end of the year.[19] In July 2024, the Financial Times reported that Superforecasters "have consistently outperformed financial markets in predicting the Fed's next move" >In particular, a 2015 study found that key predictors of forecasting accuracy were "cognitive ability [IQ], political knowledge, and open-mindedness".[23] Superforecasters "were better at inductive reasoning, pattern detection, cognitive flexibility, and open-mindedness". I'm really not sure what you want me to take from this article? Do you contend that everyone has the same competency at forecasting stock movements?
- xpe 10mo ago> I'm really not sure what you want me to take from this article? I linked to the Wikipedia page as a way of pointing to the book Superforecasters by Tetlock and Gardner. If forecasting interests you, I recommend using it as a jumping off point. > Do you contend that everyone has the same competency at forecasting stock movements? No, and I'm not sure why you are asking me this. Superforecasters does not make that claim. > I'm really not sure what you want me to take from this article? If you read the book and process and internalize its lessons properly, I predict you will view what you wrote above in a different different light: > Gotta auto grade every HN comment for how good it is at predicting stock market movement then check what the "most frequently correct" user is saying about the next 6 months. Namely, you would have many reasons to doubt such a project from the outset and would pursue other more fruitful directions.
- dw_arthur 10mo agoReading this I feel the same sense of dread I get watching those highly choreographed Chinese holiday drone shows.
- huflungdung 10mo ago[dead]
- deleted 10mo ago[deleted]
- pierrec 10mo ago"the distributed “trillions of Tamagotchi” vision never materialized" I begrudgingly accept my poor grade.
- Sophira 10mo agoIt somehow feels right to see what GPT-5 thinks of the article titled "Machine learning works spectacularly well, but mathematicians aren’t sure why" and its discussion: https://karpathy.ai/hncapsule/2015-12-04/index.html#article-10672276 https://karpathy.ai/hncapsule/2015-12-04/index.html#article-...
- intheitmines 10mo agoInteresting that for the "December 16 2015 geohot is building Comma" it graded geohot's comments on the thread as only B
- snowwrestler 10mo agoPresumably because of how things went with Comma since then.
- npunt 10mo agoOne of the few use cases for LLMs that I have high hopes for and feel is still under appreciated is grading qualitative things. LLMs are the first tech (afaik) that can do top-down analysis of phenomena in a manner similar to humans, which means a lot of important human use cases that are judgement-oriented can become more standardized, faster, and more readily available. For instance, one of the unfortunate aspects of social media that has become so unsustainable and destructive to modern society is how it exposes us to so many more people and hot takes than we have ability to adequately judge. We're overwhelmed. This has led to conversation being dominated by really shitty takes and really shitty people, who rarely if ever suffer reputational consequence. If we build our mediums of discourse with more reputational awareness using approaches like this, we can better explore the frontier of sustainable positive-sum conversation at scale. Implementation-wise, the key question is how do we grade the grader and ensure it is predictable and accurate?
- Arodex 10mo agoThis is wrong, just look at this comment here: https://news.ycombinator.com/item?id=46222523 https://news.ycombinator.com/item?id=46222523 LLM can't grade reliably human text. It doesn't understand it.
- throwaway984393 10mo ago[dead]
- Uptrenda 10mo agodude, please do this for every year until today. This idea is actually amazing. If you need more money for API credits im sure people here could help donate.
- DonHopkins 10mo agoI'd love to see an "Annie Hall" analysis of hn posts, for incidents where somebody says something about some piece of software or whatever, and the person who created it replies, like Marshall McLuhan stepping out from behind a sign in Annie Hall. https://www.youtube.com/watch?v=vTSmbMm7MDg https://www.youtube.com/watch?v=vTSmbMm7MDg
- apparent 10mo ago> And then when you navigate over to the Hall of Fame, you can find the top commenters of Hacker News in December 2015, sorted by imdb-style score of their grade point average. Now let's make a Chrome extension that subtly highlights these users' comments when browsing HN.
- popinman322 10mo agoIt doesn't look like the code anonymizes usernames when sending the thread for grading. This likely induces bias in the grades based on past/current prevailing opinions of certain users. It would be interesting to see the whole thing done again but this time randomly re-assigning usernames, to assess bias, and also with procedurally generated pseudonyms, to see whether the bias can be removed that way. I'd expect de-biasing would deflate grades for well known users. It might also be interesting to use a search-grounded model that provides citations for its grading claims. Gemini models have access to this via their API, for example.
- khafra 10mo agoYou can't anonymize comments from well-known users, to an LLM: https://gwern.net/doc/statistics/stylometry/truesight/index https://gwern.net/doc/statistics/stylometry/truesight/index
- WithinReason 10mo agoThat's an overly strong claim, an LLM could also be used to normalise style
- wetpaws 10mo agoHow would you possibly grade comments if you change them?
- koakuma-chan 10mo agoYou don’t need comments, just facts in them to see if they’re accurate.
- strken 10mo agoExtract the concrete predictions, evaluate them as true/false/indeterminate, and grade the user on the number of true vs false?
- NooneAtAll3 10mo agoUX feedback: I wish clicking on a new thread scrolled right side to the top again reading from the end isn't really useful, y'know :)
- anshulbhide 10mo agoI often summarise HN comments (which are sometimes more insightful than the original article) using an LLM. Total game-changer.
- Tossrock 10mo agoSo where do I collect my prize for this 2015 comment? https://news.ycombinator.com/item?id=9882217 https://news.ycombinator.com/item?id=9882217
- johncolanduoni 10mo agoNever call a man happy until he is dead. Also I don’t think your argument generalizes well - there are plenty of private research investment bubbles that have popped and not reached their original peaks (e.g. VR).
- Tossrock 10mo agoIt wasn't a generalized argument, though, it was a specific one, about AI.
- johncolanduoni 10mo agoOkay, but the only part that’s specific to AI (that the companies investing the money are capturing more value than they’re putting into it) is now false. Even the hyperscalers are not capturing nearly the value they’re investing, though they’re not using debt to finance it. OpenAI and Anthropic are of course blowing through cash like it’s going out of style, and if investor interest drops drastically they’ll likely need to look to get acquired.
- xpe 10mo agoHere is one sentence from the referenced prediction: > I don't think there will be any more AI winters. This isn't enough to qualify as a testable prediction, in the eyes of people that care about such things, because there is no good way to formulate a resolution criteria for a claim that extends indefinitely into the future. See [1] for a great introduction. [1]: https://www.astralcodexten.com/p/prediction-market-faq https://www.astralcodexten.com/p/prediction-market-faq
- DeathArrow 10mo ago>I believe it is quite possible and desirable to train your forward future predictor given training and effort. That's interesting. I wouldn't have thought that a decent generic forward future predictor would be possible.
- nixpulvis 10mo agoQuick give everyone colors to indicate their rank here and ban anyone with a grade less than C-. Seriously, while I find this cool and interesting, I also fear how these sorts of things will work out for us all.
- pnt12 10mo agoOn the site itself: it's great that this was produced in 1h with 60$. This is amazing to create small utilities, explore your curiosity, etc. But the site is also quite confusing and messy. OK for a vibe coded experiment, sure, but wouldn't be for a final product. But I fear we're gonna see more and more of this. Big companies downsizing their tech departments and embracing vibe coded. Comparing to inflation, shrinkflation and skimpflation/ enshittification , will we soon adopt some word for this? AIflation? LLMflation? And how will this comment score in a couple of years? :)
- alister 10mo ago> https://karpathy.ai/hncapsule/2015-12-24/index.html#article-10786492 https://karpathy.ai/hncapsule/2015-12-24/index.html#article-... I wonder why ChatGPT refused to analyze it? The HN article was "Brazil declares emergency after 2,400 babies are born with brain damage" but the page says "No analysis available".
- bspammer 10mo agoMy guess is that it’s because there’s a lot of very negative comments about Brazil in that article. Trying to grade people for their opinions on a topic like that gets into dangerous territory.
- bbcisking 10mo agoWhy not rank ESP for each HN user, with evidence?
- jeffnappi 10mo agoThe analysis of the 2015 article about Triplebyte is fascinating [1]. Particularly the Awards section. 1. https://karpathy.ai/hncapsule/2015-12-08/index.html#article-10698009 https://karpathy.ai/hncapsule/2015-12-08/index.html#article-...
- tgtweak 10mo agoCool - now make it analyze all of those and come up with the 10 commandments of commenting factually and insightfully on HN posts...
- nomel 10mo ago> I realized that this task is actually a really good fit for LLMs I've found the opposite, since these models still fail pretty wildly at nuance. I think it's a conceptual "needle in the haystack sort of problem. A good test is to find some thread where there's a disagreement and have it try to analyze the discussion. It will usually strongly misrepresent what was being said, by each side, and strongly align with one user, missing the actual divide that's causing the disagreement (a needle).
- gowld 10mo agoAs always, which model versions did you use in your test?
- nomel 10mo agoClaude Opus 4.5, Gemini 3 Pro, ChatGPT 5.1. Haven't tried ChatGPT 5.2. It requires that the discussion has nuance, to see the failure. Gemini is, by far the, worst at this (which fits my suspicion that they heavily weighted reddit posts). I don't think this is all that strange though. The human, on one side of the argument, is also missing the nuance, which is the source of the conflict. Is there a belief that AI has surpassed the average human, with conversational nuance!?
- bretpiatt 10mo ago10 Years Ago, December 11, 2015 - Introducing Open AI -- very meta: https://karpathy.ai/hncapsule/2015-12-11/index.html#article-10720176 https://karpathy.ai/hncapsule/2015-12-11/index.html#article-... The company has changed and it seems the mission has as well.
- bspammer 10mo agoYes very funny to see their own model betray them like this: > The original “non‑profit, open, patents shared” promise now reads almost like an alternate timeline. Today OpenAI is a capped‑profit entity with a massive corporate partner, closed frontier models, and an aggressive product roadmap.
- xpe 10mo agoMany people are impressed by this, and I can see why. Still, this much isn't surprising: the Karpathy + LLM combo can deliver quickly. But there are downsides of blazing speed. If you dig in, there are substantial flaws in the project's analysis and framing, such as the definition of a prediction, assessing comments, data quality overall, and more. Go spelunking through the comments here and notice people asking about methodology and checking the results. Social science research isn't easy; it requires training, effort, and patience. I would be very happy if Karpathy added a Big Flashing Red Sign to this effect. It would raise awareness and focus community attention on what I think are the hardest and most important aspects of this kind of project: methodology, rigor, criticism, feedback, and correction.
- deleted 10mo ago[deleted]
- abhinav_sk 10mo agoWhat's interesting is that the hindsight it has now is not going to be what it has in 10 years either. Some of the most wrong and most prescient comments could switch as stuff unfolds. In a way some could both still be wrong and right just at different points in time.
- JetSetWilly 10mo agoIt would be great to run this on a collection of interesting threads over different periods and not just one snapshot. For example, the thread from the day Trump got elected in 2016, the thread from the day of brexit and so on. Those are the times when people make many passionate predictions about how the future will play out, be good to see them retroactively scored.
- rkuykendall-com 10mo agoAssuming this keeps running, I suppose we just have to wait about a year.