7 ms·
Thoughts[^0] from Theo, who had early access: > It's a damn good model. Not quite as "smart" as Fable, but it is incredibly capable. Fixed all the problems I h
by aarvin_roshin 3mo ago
Thoughts[^0] from Theo, who had early access:
> It's a damn good model. Not quite as "smart" as Fable, but it is incredibly capable. Fixed all the problems I had with GPT-5.5.
> It is incredibly determined. Will run for a day without even using a /goal. It understands subagents incredibly well and is great at orchestrating. It's super pleasant in use cases like OpenClaw and Hermes Agent. It knows iOS dev incredibly well.
> It has rough edges too, but FAR fewer than 5.5 did.
> For many things, gpt-5.6-sol will become my obvious defaults.
> It is better about [following instructions] than 5.5 was. Understands intent well and hammers until it gets there. Sometimes a bit too hard.
Also[^1]:
> gpt-5.6-sol is world leading in computer use. It made me use it 100x more. When we lost access to 5.6, I quickly started to go insane without it
[^0]: https://nitter.net/theo/status/2074708892341481755 https://nitter.net/theo/status/2074708892341481755
[^1]: https://nitter.net/theo/status/2074720467395756499 https://nitter.net/theo/status/2074720467395756499
- isoprophlex 3mo agoIncreased tenacity & goal following is exactly what I want in this model, to make it compete with Claude models. (A little toning down of the goblin fetish would be nice too, haha.)
- bashtoni 3mo agoI feel like listening to Theo about anything technical is like consulting a Labrador retriever for advice on quantum physics. Every time I've ever seen one of his videos it's pretty clear he has very little understanding of development or engineering. I first became aware of him from his early "unit tests are a waste of time" stuff, and it seems his skillset is building a personal brand. Fair play, he's clearly talented at that, but that doesn't make his opinion on anything else worthwhile.
- ttoinou 3mo agoThere is a whole religion about tests that is worth attacking though
- bashtoni 3mo agoSure. If his take was "100% unit test coverage is a waste of time" I think that's not unreasonable. You could make a case that the "you must write tests before you write code, every single time!" stuff is needlessly dogmatic. I also think that sometimes people focus too much on unit tests to the detriment of end to end tests that better model actual system interactions. None of these were Theo's take. He was pushing the idea that unit tests in general were a waste of time because you could be shipping new features instead. https://www.youtube.com/watch?v=pvBHyip4peo https://www.youtube.com/watch?v=pvBHyip4peo for an example of this. The nicest possible interpretation on this is that he's deliberately saying something he knows is wrong to attract attention.
- theshrike79 3mo agoTests before code makes sense when fixing bugs. Red-green specifically. 1: get bug 2: write tests that should work, but don’t because of bug 3: fix bug 4: confirm fix by running tests Makes things a LOT easier for people checking the PR, they can just confirm the tests are correct pretty much. As a bonus the same bug can’t surface again.
- bashtoni 3mo agoYep, I'm in full agreement. When extending functionality of some already existing code it also generally makes sense to write tests first. I think the value is much lower (maybe even negative) when you're still trying to work out what shape the code will take, in an initial implementation. Of course, as others have pointed out, nuanced opinion doesn't get clicks or YouTube views.
- ttoinou 3mo agoOh I do that naturally as my rational problem investigation. Sometimes you can’t write a test for that, you need to test it yourself
- scotty79 3mo agoHe doesn't believe that unit tests are complete waste of time. Just a relative waste of time. He doesn't mind AI agents writing tests. It's just mostly waste of time for humans. Because the value you get for them is not worth developer time in most cases. It's worth agent time.
- patates 3mo ago> it's pretty clear he has very little understanding of development or engineering I cannot prove it but I have a feeling that you may be conflating "he clearly has different opinions on things I consider non-negotiable" to "he doesn't know what he's talking about". I also watched a lot of his videos. I wildly disagree with him a lot of times, but he has his reasoning, and I can see (and verify!) that those ideas are coming from an engineering perspective.
- steve_adams_86 3mo agoHe's clearly very knowledgeable about some things, but I think he has harmed his credibility be becoming a 'tuber who prioritizes thumbnails and hot takes over engineering.
- patates 3mo ago> he has demolished his credibility be becoming a 'tuber who prioritizes thumbnails and hot takes over engineering I don't agree that he has demolished his credibility. I also dislike the youtube face and sensationalization but I personally don't hold it against him, given the Youtube algorithm. Regardless of his style, I like hearing the take from an engineer who's working in a different country/culture and has a completely different perspective. edit: it seems you changed "demolished" to "harmed". I still don't agree but it reads more defensible IMHO, thank you.
- satvikpendem 3mo agoNot really. If you're a YouTuber it's necessary to follow the algorithm which includes making such a YouTuber face, clickbait actually works and has a direct financial correlation as Linus Tech Tips has shown.
- incrudible 3mo agoIf you are a competent engineer already why do you need to create self demeaning clickbait content on YouTube? Narcissism? I aggressively block any such content, because if you click on any of it, you easily get sent down the YouTube spiral of crap.
- Havoc 3mo agoAnd half his videos are him coming up with indirect ways of saying look how amazing I am.
- scotty79 3mo agoIt's curious how so many people get triggered by a smart person saying what he believes to be true. Yes, he is pretty amazing. Yes, he is rarely wrong. No, it doesn't affect you or me in any way because he is not in competition with you. Go do something else if you don't enjoy his takes. I don't get many programmer influencers in my feed that deal with newsworthy relevant stuff. Theo is the least wrong and most humble one in my perception.
- ttul 3mo agoI disagree with this assessment. Theo is nerdy and, yes, he has a healthy ego. But, he provides insightful commentary on his channel and he works very hard to present what he believes is the truth. Compared to many YouTubers, whose content is vacuous, Theo is mostly the real deal.
- zarzavat 3mo agoThere's a simpler explanation. Social media rewards surprise and hype, not truth. Don't expect objectivity from someone who gets paid by the view.
- torginus 3mo agoHe has big 'theatre kid' energy (at least certainly had, watched him years ago) - he desperately wants to make clear that there's a group of cool kids and he's in it. His youtube channel used to be about talking about the new FOTM Javascript framework/technology - not presented as 'here's a cool thing, let's check it out' but 'everyone worth a damn already uses this, get with the times grandpa'
- satvikpendem 3mo ago"Average Theo video be like": https://youtu.be/h1p9zdUtUdo https://youtu.be/h1p9zdUtUdo It's shocking how many accurate tropes this hits.
- pdantix 3mo agoi already found his clear shilling of nextjs a bit distasteful, but his whole gpt-5 thing really just made it clear he's just not worth listening to.
- stingraycharles 3mo ago“Understands intent well and hammers until it gets there. “ If there’s anything I learned over the past 12-18 months is that this is a recipe for disaster, except for throwaway stuff. I thought most senior engineers settled on the fact that steering a model yields much better results?
- oefrha 3mo agoI wouldn’t call it a recipe for disaster, but oh boy if you leave an agent that “hammers until it gets there” on its own with an underlying bug in a dependency…
- icepush 3mo agoIt's very possible that would be the best strategy over the last 12-18 months and now that this is released it is no longer the best strategy.
- stingraycharles 3mo agoThat would be an extremely massive leap if agents could suddenly make nuanced architectural decisions and prevent technical debt. In my experience even Fable still requires guidance (although the options it provides are generally better).
- hodgehog11 3mo agoFor some tasks, there is no amount of "steering" that will produce sensible code. The model needs to be sufficiently capable as a baseline; this is the "intent" that people are referring to with Fable.
- stingraycharles 3mo agoThat doesn’t sound like the “it hammers until it’s done”-type of intent. Just last night Fable decided to get into a rabbit hole of debugging a database driver issue by packet sniffing the network traffic instead of just adding debug statements to the code. Definitely needed steering, and I don’t know many people whose first intuition would be to use pcap when they have a segfault.
- jychang 3mo ago> Not quite as "smart" as Fable, but it is incredibly capable. THIS IS BECAUSE GPT-5.6 SOL IS... just a more posttrained version of GPT-5.5, not a brand new bigger model than GPT-5.5. It's not like how Mythos is bigger than Opus. OpenAI switching to Sol/Terra/Luna renaming is just a way to rip off people and charge more usage for the same sized model. GPT-5.6 --------> GPT-5.6 Sol GPT-5.6-mini ---> GPT-5.6 Terra GPT-5.6-nano ---> GPT-5.6 Luna Except OpenAI is about to advertise GPT-5.6 Sol and GPT-5.6 Terra as a whole tier better, than if they named it GPT-5.6 and GPT-5.6-mini.
- ppaattrriicckk 3mo ago> OpenAI switching to Sol/Terra/Luna renaming is just a way to rip off people and charge more money for the same sized model. Excuse me, but what are you on about? Unless I'm mistaken, they have literally(1) stated that it will cost $5 per 1M tokens in, and $30 for 1M output tokens. The same as GPT-5.5. [1] https://openai.com/index/previewing-gpt-5-6-sol/ https://openai.com/index/previewing-gpt-5-6-sol/
- deleted 3mo ago[deleted]
- threatripper 3mo agoMy feeling is that GPT-5.5 doesn't lack the raw intelligence so much as it lacks "methodology". I don't know how exactly to put it... how to approach a problem, how to take care of the details and side effects, how to handle unexpected difficulties and bugs, how to not spin out of control, how to write solid code, how to clean up afterwards, how to document, how to give useful feedback... the things that you learn on the job. So, if they improved a lot in those areas, then GPT-5.6 could become a lot more useful compared to GPT-5.5 even though it might score lower in many benchmarks. It's possible but unlikely since their approach was mostly brute force in the past.
- gck1 3mo agoIs Fable really that much different? I almost instinctively create elaborate processes, workflows, set up a bunch of linters and dump research docs any time I bootstrap a new project regardless of what model I'm using. They all spiral out of control if they're not following a predefined process.
- zuzululu 3mo agohere is the original x post https://x.com/theo/status/2074708892341481755 https://x.com/theo/status/2074708892341481755 5.6 sol seems to hit a lot of the gaps with 5.5 sucks its not "mythos" but i will take it
- jeswin 3mo ago> Thoughts[^0] from Theo, who had early access: I looked at his YouTube, and found a stream of industry gossip and beginner content like "web dev tutorials". I have nothing against such content and it may be useful and good fun to watch. But does that say anything about this particular model? People have been using models effectively for web code since Gpt 3.x.
- satvikpendem 3mo agoThat's sad to see Sol not beating Fable as it was explicitly stated by OpenAI that Sol benchmarks and overall performance were better than Fable.
- xyzsparetimexyz 3mo agoThis guy is enthusiastic about everything. Not a great benchmark
- rafaelmn 3mo ago[dead]
- nullbio 3mo agoNext up: Thoughts from an OnlyFans model.
- Copenjin 3mo ago> Thoughts[^0] from Theo I will stop here, sorry but I think we have limited time to listen to opinions and nowadays since they are abundant on social media we should give preference to the substantiated ones.