5 ms·
Everything is still based on 4 4o still right? is a new model training just too expensive? They can consult deepseek team maybe for cost constrained new models.
by sfmike 10mo ago
Everything is still based on 4 4o still right? is a new model training just too expensive? They can consult deepseek team maybe for cost constrained new models.
- catigula 10mo agoThe irony is that Deepseek is still running with a distilled 4o model.
- blovescoffee 10mo agoSource?
- verdverm 10mo agoApparently they have not had a successful pre training run in 1.5 years
- fouronnes3 10mo agoI want to read a short scify story set in 2150 about how, mysteriously, no one has been able to train a better LLM for 125 years. The binary weights are studied with unbelievably advanced quantum computers but no one can really train a new AI from scratch. This starts cults, wars and legends and ultimately (by the third book) leads to the main protagonist learning to code by hand, something that no human left alive still knows how to do. Could this be the secret to making a new AI from scratch, more than a century later?
- armenarmen 10mo agoI’d read it!
- barrenko 10mo agoMonsieur, if I may offer a vaaaguely similar story on how things may progress https://www.owlposting.com/p/a-body-most-amenable-to-experimentation https://www.owlposting.com/p/a-body-most-amenable-to-experim...
- verdverm 10mo agoYou can ask 2025 Ai to write such a book, it's happy to comply and may or may not actually write the book https://www.pcgamer.com/software/ai/i-have-been-fooled-reddit-user-endures-the-roasting-of-a-lifetime-after-asking-how-to-download-a-487mb-book-they-worked-on-with-chatgpt-for-over-2-weeks/ https://www.pcgamer.com/software/ai/i-have-been-fooled-reddi...
- WhyOhWhyQ 10mo agoThere's a scifi short story about a janitor who knows how to do basic arithmetic and becomes the most important person in the world when some disaster happens. Of course after things get set up again due to his expertise, he becomes low status again.
- bradfitz 10mo agoI had to go look that up! I assume that's https://en.wikipedia.org/wiki/The_Feeling_of_Power https://en.wikipedia.org/wiki/The_Feeling_of_Power ? (Not a janitor, but "a low grade Technician"?)
- WhyOhWhyQ 10mo agoHmm it could be a false memory, since this was almost 15 years ago, but I really do remember it differently than the text of 'Feeling of Power'.
- YouAreWRONGtoo 10mo ago[dead]
- ssl-3 10mo agoSounds good. Might sell better with the protagonist learning iron age leatherworking, with hides tanned from cows that were grown within earshot, as part of a process of finding the real root of the reason for why any of us ever came to be in the first place. This realization process culminates in the formation of a global, unified steampunk BDSM movement and a wealth of new diseases, and then: Zombies. (That's the end. Zombies are always the end.)
- wafflemaker 10mo agoSorry, but compared with the parent, my money is in you ssl-3. Do you get better results from prompting by being more poetic?
- ssl-3 10mo ago> Do you get better results from prompting by being more poetic? Is that yet-another accusation of having used the bot? I don't use the bot to write English prose. If something I write seems particularly great or poetic or something, then that's just me: I was in the right mood, at the right time, with the right idea -- and with the right audience. When it's bad or fucked-up, then that's also just me. I most-assuredly fuck up plenty. They can't all be zingers. I'm fine with that. --- I do use the hell out of the bot for translating my ideas (and the words that I use to express them) into languages that I can't speak well, like Python, C, and C++. But that's very different. (And at least so far I haven't shared any of those bot outputs with the world at all, either.) So to take your question very literally: No, I don't get better results from prompting being more poetic. The responses to my prompts don't improve by those prompts being articulate or poetic. Instead, I've found that I get the best results from the bot fastest by carrying a big stick, and using that stick to hammer and welt it into compliance. Things can get rather irreverent in my interactions with the bot. Poeticism is pretty far removed from any of that business.
- wafflemaker 10mo agoNo. I just genuinely liked your style, and didn't notice previous posts by you. I haven't yet learned to look at names on hn, it's mostly anonymous posts for me. No snark here. And was also genuinely curious if better writing style yields better results. I've observed that using proper grammar gives slightly better answers. And using more "literacy"(?) kind of language in prompts sometimes gives better answers and sometimes just more interesting ones, when bots try to follow my style. Sorry for using the word poetic, I'm travelling and sleep deprived and couldn't find the proper word, but didn't want to just use "nice" instead either.
- georgefrowny 10mo agoAn software version of Asimov's Holmes-Ginsbook device? https://sfwritersworkshop.org/node/1232 https://sfwritersworkshop.org/node/1232 I feel like there was a similar one about software, but it might have been mathematics (also Asimov: The Feeling of Power)
- ijl 10mo agoWhat kind of issues could prevent a company with such resources from that?
- verdverm 10mo agoDrama if I had to pick the symptom most visible from the outside. A lot of talent left OpenAI around that time, most notably in this regard would be Ilya in May '24. Remember that time Ilya and the board ousted Sam only to reverse it almost immediately? https://arstechnica.com/information-technology/2024/05/chief-scientist-ilya-sutskever-leaves-openai-six-months-after-altman-ouster/ https://arstechnica.com/information-technology/2024/05/chief...
- Wowfunhappy 10mo agoI thought whenever the knowledge cutoff increased that meant they’d trained a new model, I guess that’s completely wrong?
- brokencode 10mo agoTypically I think, but you could pre-train your previous model on new data too. I don’t think it’s publicly known for sure how different the models really are. You can improve a lot just by improving the post-training set.
- rockinghigh 10mo agoThey add new data to the existing base model via continuous pre-training. You save on pre-training, the next token prediction task, but still have to re-run mid and post training stages like context length extension, supervised fine tuning, reinforcement learning, safety alignment ...
- astrange 10mo agoContinuous pretraining has issues because it starts forgetting the older stuff. There is some research into other approaches.
- elgatolopez 10mo agoWhere did you get that from? Cutoff date says august 2025. Looks like a newly pretrained model
- SparkyMcUnicorn 10mo agoIf the pretraining rumors are true, they're probably using continued pretraining on the older weights. Right?
- deleted 10mo ago[deleted]
- FergusArgyll 10mo ago> This stands in sharp contrast to rivals: OpenAI’s leading researchers have not completed a successful full-scale pre-training run that was broadly deployed for a new frontier model since GPT-4o in May 2024, highlighting the significant technical hurdle that Google’s TPU fleet has managed to overcome. - https://newsletter.semianalysis.com/p/tpuv7-google-takes-a-swing-at-the https://newsletter.semianalysis.com/p/tpuv7-google-takes-a-s... It's also plainly obvious from using it. The "Broadly deployed" qualifier is presumably referring to 4.5
- ric2b 10mo agoHow is that a technical hurdle if they obviously were able to do it before? It's probably just a question of cost/benefit analysis, it's very expensive to do, so the benefits need to be significant.