3 ms·
They have developed an LLM, so they are an AI lab, but the quality of that model suggests they're not a frontier anything.
by hawkice 4mo ago
They have developed an LLM, so they are an AI lab, but the quality of that model suggests they're not a frontier anything.
- gowld 4mo agoI am also an "AI lab", but I look more like a corporate cog, because that's where most of my revenue comes from and how I spend the most my time.
- beepbopboopp 4mo agoOr the model was a marketing expense to capitalize the data center model. Im not saying it was intentionally that, but its been an effective "that."
- bpodgursky 4mo agoEh. It was a leading model for a few weeks, it was a real effort, but they never built a real revenue model around it. It wasn't SaaS, it wasn't for governments, it couldn't get B2C payments. Made it hard to justify the training cost to stay at the frontier.
- dsgn93 4mo ago[flagged]
- Qhemlomo 4mo agoSo like the 4D Chess Trump is playing with us? Come on, the most logical thing is that Musk overestimated the compute he needs and got lucky with the secondary usage of it. As soon as the IPO is done and if it didn't fail, he will buy curser and try to push again if he hasn't given up on it. He also needs some compute for the robotics stuff and for Tesla in-car entertainment and for training FSD.
- bottlepalm 4mo agoGrok isn't at the front of the frontier, but they are there for sure.
- grim_io 4mo agoAs much as the Chinese models, so not at all.
- leetharris 4mo agoI have the pro account for ChatGPT, Claude, Gemini, and Grok. They all have various strengths and weaknesses. My favorite is still ChatGPT, then Gemini/Claude, then Grok. Grok often feels 1-2 generations behind the competition in general use, but it has three things that I love: 1. It seems to be the best at understanding current events. Maybe due to X integration, or some other tool call optimization in the backend? I don't know, but I often ask about things going on, and the other models have outdated info, give unhelpful answers, etc. 2. It is generally the least sycophantic for personal things. Anthropic is getting here too. ChatGPT and Gemini are working on this, but previous models in those families would almost never say anything negative about what I am doing. Sometimes I need career advice, personal advice, etc and I like the tone of how it responds. I think Claude will be caught up soon. 3. For professional work, there are certain topics that other models would refuse to engage with. At my last company we had an enormous amount of legal users. When a deposition would need a summary on certain topics, most models would refuse. Grok would not. I understand the need for safety and I don't blame the other model providers, but for some professional use cases you NEED a model that is capable of handling sensitive subjects.
- epolanski 4mo agoMy SO works in audit/compliance and business Gemini definitely does not refuse to answer.
- e9 4mo agoI recently worked with NRC dataset, specifically about nuclear reactor events and status reports(example: https://www.nrc.gov/reading-rm/doc-collections/event-status/event/en https://www.nrc.gov/reading-rm/doc-collections/event-status/...). Public data that just needed some cleaning. Several time Claude API would refuse to engage. Because of that I can't trust Claude to clean production data sets.
- Azantys 4mo agoCareer and personal advice from LLMs, not sure if thats your best bet
- deaton 4mo agoAll 4 of these still regularly insist that I am a genius and everything I say is brilliant. Grok definitely pushes back more than the others, but I don't like how sycophantic they all still are.
- throwaway67678 4mo ago[flagged]
- plaidthunder 4mo agoIt's a general problem of defining yourself in negative terms. Being "un-{thing I don't like}" doesn't say what you are. It only excludes one possibility while leaving behind an infinitude of mostly crappy alternatives to try to choose from. Having a positive set of beliefs annoys people and and can make them feel judged, but at least it provides a vector that points somewhere definite in possibility space.
- fooker 4mo ago> the quality of that model I guess the benchmarks disagree, but whenever I need to find specific information that does not easily show up with a web search, I try chatgpt, gemini and grok. Grok surfaces what I was looking for more often than the others. Things like "find the github repo from 2017 that does $vague_thing".
- chatmasta 4mo agoGrok does seem to have the best searching capabilities, and not just for twitter. I wonder what search engine they’re using on the backend.
- PixyMisa 4mo agoGood question. You can actually see the searches it runs (momentarily) so testing could determine if it's using public search engines or a private system.
- Azantys 4mo agoIsnt that more Perplexitys thing anyways?
- gowld 4mo agoCan you give a specific example (that doesn't violate any privacy you want to protect)?
- PixyMisa 4mo agoI find that too. I use Claude for coding but when I need to dig out something based on limited data I turn to Grok and it delivers.
- mbesto 4mo agoAnd they are planning (well "planning" if you believe Elon) to start building their LLM over from scratch, which means they need a HUGE ass training data center, i.e. not a data center for inference to do so.
- harrall 4mo agoBut supposedly they’re the cheapest for certain workloads, especially ones that have high tokens and can make use of caching. So they’re cutting edge in that way.