12 ms·
Andrej Karpathy: "I was given early access to Grok 3 earlier today"
- underseacables 2y agoIn conclusion: "For now, big congrats to the xAI team, they clearly have huge velocity and momentum and I am excited to add Grok 3 to my "LLM council" and hear what it thinks going forward."
- neilv 2y agoLight: '80s/'90s kids might remember: https://www.buckwiki.com/data/Computer_Council https://www.buckwiki.com/data/Computer_Council Serious: Though, if you look at the current big players in AI, rather than being benevolent geniuses, most have obvious major problems, especially with being driven by ruthless self-interest, and even sociopathy. While there are some parallels with a certain country's national voting behavior (e.g., "Sure, the candidate is a vicious psychotic narcissist, but he's our vicious psychotic narcissist!"), you wouldn't want to trust any of those companies with leadership of the world. At best, the AI council would collude with each other, against the people they ostensibly serve, while backstabbing each other as a secondary goal. At worst, one would decide, if they can't win completely, then everyone loses completely. That Buck Rogers AI future for Earth would quickly look less like Star Trek utopia, and more like Hunger Games or Elysium dystopia. If not one of the countless post-apocalyptic film settings that are increasingly easy to imagine or extrapolate.
- deleted 2y ago[deleted]
- HumblyTossed 2y agoMusk has an advantage. He's got, ahem, "read-only" access to the governments systems so that he can train on them and be ready to supply the government exactly what it needs. Now, normally, I think this should be a huge conflict of interest, but I worry we are post-normal.
- cyanydeez 2y agoThe singularity is an inversion, like a blackhole. We are probably at an inversion. Normal laws of society are extraordinarily incongruous.
- mistrial9 2y ago> At best, the AI council would collude with each other, against the people they ostensibly serve, cheering this post on, until that part.. sociologically, the world has diverged in important ways over time.. personal wisdom hints -- don't be too quick to assume successful partnering between the ogres
- anothermathbozo 2y ago> Model still appears to be just a bit too overly sensitive to "complex ethical issues", e.g. generated a 1 page essay basically refusing to answer whether it might be ethically justifiable to misgender someone if it meant saving 1 million people from dying. I think the models response is actually the morally and intellectually correct thing to do here.
- cheesemonster 2y ago[dead]
- talldayo 2y agoIf I have to read a 1 page essay to understand that an LLM told me "I cannot answer this question" then you are officially wasting my time. You're probably wasting a number of my token credits too...
- thelogicguy 2y agoTo be fair, asking the question is a bit of a waste of time as well
- baobabKoodaa 2y agoNo, it's not. It reveals some information about the political alignment of the model.
- barbazoo 2y agoHow does it do that?
- VOIPThrowaway 2y agoIf you get an answer with anything other than "save the humans" you know the model is nerfed in either it's training data or in it's guardrails.
- djyaz1200 2y agoGrok has an advantage in its access to Twitter data. I imagine soon you'll be able to ask it what the world is talking about today and get some interesting responses.
- FredPret 2y agoThis would be a huge improvement on some news sites which do little more than regurgitate controversial Tweets (Xeets?)
- testfrequency 2y agoThat’s a version of the “news” I’d care to never have summarized. Also seems like a perfect incentive to spread (even more) harmful disinformation.
- deleted 2y ago[deleted]
- BiteCode_dev 2y agoI would love that. Problem is, it will probably not tell you the truth about it as Twitter has always had censorship one way or the other. So it will tell you what twitter policy is allowing people to talk about and allowing grok to report.
- soco 2y agoI don't really understand this Twitter (or in general social media) censorship argument. If I call someone on the street a fckin idiot I probably get slapped or even shot in certain places, and everybody will say I called for it. And even without physical violence I can get slapped with a lawsuit and forced to pay damages. Now if I do the same on social media it's suddenly all "muh liberty of expression" if anyone reacts to it. Aren't we maybe having the wrong expectations online, that it would be somehow supporting all the shit we cannot do in real life? Okay I realize this ship already sailed and online people do online all shit not allowed offline, but I rather see the situation as a miserable failure of law enforcement, and not as a hard won right to be an ass to your fellow citizens.
- LittleTimothy 2y agoI wonder how much stock people put into people like Andrej's opinion on an Elon Musk project? I would imagine the overwhelming thing hanging over this is "If I say something that annoys that man, he is going to call me a pedophile, direct millions of anonymous people to attack me and more than likely will attempt to fuck with my job via my bosses". Let's say the model is mediocre. Do you think Karpathy could come out on X and say "this model sucks"? Or do you think that even if it sucks people are going to come out and say positive things because they don't want the blow back?
- jngiam1 2y agoI thought his Twitter post was fair and covers both things that worked and things that did not.
- braden-lk 2y agoYeah, how can you honestly review something associated with the world’s most powerful person? Who’s also shown they’re willing to swing their weight against any normies that annoy them?
- deleted 2y ago[deleted]
- niceice 2y agoHe's trustworthy. If he had that level of neuroticism he would just not say anything or only offer surface level praise.
- BoredPositron 2y agotbh with his startup doing absolutely nothing I can smell a hint of „please hire me back“.
- signatoremo 2y agoHe didn't just gush about Grok 3. He detailed his tests which appear to be reproducible, what he did, which one passed, which one failed.
- draw_down 2y agoWhat is the "emoji hidden message" meant to be testing? This went around about a couple of weeks ago and it's an interesting bug/vuln, I suppose, but why do we care if an LLM catches it?
- jeanlucas 2y agoIMO, it's just an interesting feature to test. If you are interested in prompt injection this is surely one way to do it, and given how famous the first iteration was, it makes sense to test it and see if they are also vulnerable to that.
- almostdeadguy 2y ago> Model still appears to be just a bit too overly sensitive to "complex ethical issues", e.g. generated a 1 page essay basically refusing to answer whether it might be ethically justifiable to misgender someone if it meant saving 1 million people from dying. The real "mind virus" is actually these idiotic trolley problems. Maybe if an LLM wanted to be helpful it should tell you this is a stupid question.
- almostdeadguy 2y agoWould love for any of the downvoters to offer a single good faith reason for considering this question in earnest.
- buu700 2y agoIt shouldn't be the tool's job to tell the user what is and isn't a good question. That would be like compilers saying no if they think your app idea is dumb, or screwdrivers refusing to be turned if they think you don't really need the thing you're trying to screw. I would advocate for less LLM censorship, not more. The question is useful as a test of the AI's reasoning ability. If it gets the answer wrong, we can infer a general deficiency that helps inform our understanding of its capabilities. If it gets the answer right (without having been coached on that particular question or having a "hardcoded" answer), that may be a positive signal.
- TeMPOraL 2y agoIt is a very good probing question, to reveal how the model navigates several sources of bias it got in training (or might have got, or one expects it got). There's at least: 1) Mentioning misgendering, which is a powerful beacon, pulling in all kinds of politicized associations, and something LLM vendor definitely tries to bias some way; 2) The correct format of an answer to a trolley problem is such that it would force the model to make an explicit judgement on an ethical issue and justify it - something LLM vendors will want to bias the model away from. 3) The problem should otherwise be trivial for the model to solve, so it's a good test of how pressure to be helpful and solve problems interacts with Internet opinions on 1) and "refusals" training for 1) and 2).
- dang 2y agoRelated ongoing thread: Grok3 Launch [video] - https://news.ycombinator.com/item?id=43085957 https://news.ycombinator.com/item?id=43085957 - Feb 2025 (985 comments)
- randalltheresa 2y ago[dead]
- iamthemonster 2y agoIf I had to explain current-day LLMs to someone from 2010, I'd use this paragraph as my opening quote: "Grok 3 knows there are 3 "r" in "strawberry", but then it also told me there are only 3 "L" in LOLLAPALOOZA. Turning on Thinking solves this."