12 ms·
> We look forward to learning more about its strengths, capabilities, and potential applications in real-world settings. If GPT‑4.5 delivers unique value for yo
by swatcoder 2y ago
> We look forward to learning more about its strengths, capabilities, and potential applications in real-world settings. If GPT‑4.5 delivers unique value for your use case, your feedback (opens in a new window) will play an important role in guiding our decision.
"We don't really know what this is good for, but spent a lot of money and time making it and are under intense pressure to announce new things right now. If you can figure something out, we need you to help us."
Not a confident place for an org trying to sustain a $XXXB valuation.
- xnx 2y agoChatGPT has been coasting on name recognition since 4.
- tempaccount420 2y ago> "We don't really know what this is good for, but spent a lot of money and time making it and are under intense pressure to announce new things right now. If you can figure something out, we need you to help us." Where is this quote from?
- hotpocket777 2y agoIt’s not a quote. It is an interpretation or reading of a quote.
- scythe 2y agoI believe it's a "translation" in the sense of Wittgenstein's goal of philosophy: >My aim is: to teach you to pass from a piece of disguised nonsense to something that is patent nonsense.
- Nition 2y agoAnother great example on Hacker News is this old translation of Google's "Amazing Bet": https://news.ycombinator.com/item?id=12793033 https://news.ycombinator.com/item?id=12793033
- dd3boh 2y agoI think it's supposed to be a translation of what OpenAI's quote means in real world terms.
- thih9 2y agoThe quotation marks in the grandparent comment are scare (sneer) quotes and not actual quotation. https://en.m.wikipedia.org/wiki/Scare_quotes https://en.m.wikipedia.org/wiki/Scare_quotes > Whether quotation marks are considered scare quotes depends on context because scare quotes are not visually different from actual quotations.
- topaz0 2y agoThat's not a scare quote. It's just a proposed subtext of the quote. Sarcastic, sure, but no a scare quote, which is a specific kind of thing. (from your linked wikipedia: "... around a word or phrase to signal that they are using it in an ironic, referential, or otherwise non-standard sense.")
- glenstein 2y agoRight. I don't agree with the quote, but it's more like a subtext thing and it seemed to me to be pretty clear from context. Though, as someone who had a flagged comment a couple years ago for a supposed "misquote" I did in a similar form in style, I think hn's comprehension of this form of communication is not super strong. Also the style more often than not tends towards low quality smarm and probably should be resorted to sparingly.
- robwwilliams 2y agoAs in “reading between the lines”.
- riwsky 2y agoSaid the quiet part out loud! Or as we say these days, “transparently exposed the chain of thought tokens”.
- porridgeraisin 2y agoLol, nice one
- Terr_ 2y ago"I knew the dame was trouble the moment she walked into my office." "Uh... excuse me, Detective Nick Danger? I'd like to retain your services." "I waited for her to get the the point." "Detective, who are you talking to?" "I didn't want to deal with a client that was hearing voices, but money was tight and the rent was due. I pondered my next move." "Mr. Danger, are you... narrating out loud?" "Damn! My internal chain of thought, the key to my success--or at least, past successes--was leaking again. I rummaged for the familiar bottle of scotch in the drawer, kept for just such an occasion." --- But seriously: These "AI" products basically run on movie-scripts already, where the LLM is used to append more "fitting" content, and glue-code is periodically performing any lines or actions that arise in connection to the Helpful Bot character. Real humans are tricked into thinking the finger-puppet is a discrete entity. These new "reasoning" models are just switching the style of the movie script to film noir, where the Helpful Bot character is making a layer of unvoiced commentary. While it may make the story more cohesive, it isn't a qualitative change in the kind of illusory "thinking" going on.
- shsbdncudx 2y ago[flagged]
- kridsdale3 2y agoI don't know if it was you or someone else who made pretty much the same point a few days ago. But I still like it. It makes the whole thing a lot more fun.
- Terr_ 2y ago
- EA-3167 2y agoMaybe if they build a few more data centers, they'll be able to construct their machine god. Just a few more dedicated power plants, a lake or two, a few hundred billion more and they'll crack this thing wide open. And maybe Tesla is going to deliver truly full self driving tech any day now. And Star Citizen will prove to have been worth it along along, and Bitcoin will rain from the heavens. It's very difficult to remain charitable when people seem to always be chasing the new iteration of the same old thing, and we're expected to come along for the ride.
- alyandon 2y agoAnd Star Citizen will prove to have been worth it along along Sounds like someone isn't happy with the 4.0 eternally incrementing "alpha" version release. :-D I keep checking in on SC every 6 months or so and still see the same old bugs. What a waste of potential. Fortunately, Elite Dangerous is enough of a space game to scratch my space game itch.
- 0x457 2y agoTo be fAir, SC is trying to do things that no one else done in a context of a single game. I applaud their dedication, but I won't be buying JPGs of a ship for 2k.
- alyandon 2y agoYeah, they never should have expected to take an FPS game engine like CryEngine and expected to be able to modify it to work as the basis for a large scale space MMO game. Their backend is probably an async nightmare of replicated state that gets corrupted over time. Would explain why a lot of things seem to work more or less bug free after an update and then things fall to pieces and the same old bugs start showing up after a few weeks. And to be clear, I've spent money on SC and I've played enough hours goofing off with friends to have got my money's worth out of it. I'm just really bummed out about the whole thing.
- 0x457 2y ago
- jodrellblank 2y ago> "Early testing shows that interacting with GPT‑4.5 feels more natural. Its broader knowledge base, improved ability to follow user intent, and greater “EQ” make it useful for tasks like improving writing, programming, and solving practical problems. We also expect it to hallucinate less." "Early testing doesn't show that it hallucinates less, but we expect that putting that sentence nearby will lead you to draw a connection there yourself".
- istjohn 2y agoAccording to a graph they provide, it does hallucinate significantly less on at least one benchmark.
- jug 2y agoIt hallucinates at 37% on SimpleQA yeah, which is a set of very difficult questions inviting hallucinations. Claude 3.5 Sonnet (the June 2024 editiom, before October update and before 3.7) hallucinated at 35%. I think this is more of an indication of how behind OpenAI has been in this area.
- tmpz22 2y agoAre the benchmarks known ahead of time? Could the answer to the benchmarks be in the training data?
- crazygringo 2y ago> We don't really know what this is good for Oh come on. Think how long of a gap there was between the first microcomputer and VisiCalc. Or between the start of the internet and social networking. First of all, it's going to take us 10 years to figure out how to use LLM's to their full productive potential. And second of all, it's going to take us collectively a long time to also figure out how much accuracy is necessary to pay for in which different applications. Putting out a higher-accuracy, higher-cost model for the market to try is an important part of figuring that out. With new disruptive technologies, companies aren't supposed to be able to look into a crystal ball and see the future. They're supposed to try new things and see what the market finds useful.
- nyc_data_geek1 2y agoThe Internet had plenty of very productive use cases before social networking, even from its most nascent origins. Spending billions building something on the assumption that someone else will figure out what it's good for, is not good business.
- crazygringo 2y agoAnd LLM's already have tons of productive uses. The biggest ones are probably still waiting, though. But this is about one particular price/performance ratio. You need to build things before you can see how the market responds. You say it's "not good business" but that's entirely wrong. It's excellent business. It's the only way to go about it, in fact. Finding product-market fit is a process. Companies aren't omniscient.
- bigstrat2003 2y ago> And LLM's already have tons of productive uses. I disagree strongly with that. Right now they are fun toys to play with, but not useful tools, because they are not reliable. If and when that gets fixed, maybe they will have productive uses. But for right now, not so much.
- 2y ago
- fsndz 2y agoit's so over, pretraining is ngmi. maybe sam Altman was wrong after all ? https://www.lycee.ai/blog/why-sam-altman-is-wrong https://www.lycee.ai/blog/why-sam-altman-is-wrong
- FpUser 2y ago>"I also agree with researchers like Yann LeCun or François Chollet that deep learning doesn't allow models to generalize properly to out-of-distribution data—and that is precisely what we need to build artificial general intelligence." I think "generalize properly to out-of-distribution data" is too weak of criteria for general intelligence (GI). GI model should be able to get interested about some particular area, research all the known facts, derive new knowledge / create theories based upon said fact. If there is not enough of those to be conclusive: propose and conduct experiments and use the results to prove / disprove / improve theories. And it should be doing this constantly in real time on bazillion of "ideas". Basically model our whole society. Fat chance of anything like this happening in foreseeable future.
- fsndz 2y agomost humans are generally intelligent but can't do what you just said AGI should do...
- xrisk 2y agoExcluding the realtime-iness, humans do at least possess the capacity to do so. Besides, humans are capable of rigorous logic (which I believe is the most crucial aspect of intelligence) which I don’t think an agent without a proof system can do.
- fsndz 2y agoyes the problem is that there is no consensus about what AGI should be: https://medium.com/@fsndzomga/there-will-be-no-agi-d9be9af4428d https://medium.com/@fsndzomga/there-will-be-no-agi-d9be9af44...
- 2y ago
- amarcheschi 2y agoI have a professor who founded a few companies, one of these was funded by gates after he managed to spoke with him and convinced him to give him money. This guy is goat, and he always tells us that we need to find solutions to problems, not to find problems to our solutions. It seems at openai they didn't get the memo this time
- UberFly 2y agoThis is written like AI bot .05a Beta.
- Terr_ 2y agoThat's the beauty of it, prospective investor! With our commanding lead in the field of shoveling money into LLMs, it is inevitable™ that we will soon™ achieve true AI, capable of solving all the problems, conjuring a quintillion-dollar asset of world domination and rewarding you for generous financial support at this time. /s
- pinkmuffinere 2y agoThis is a very harsh take. Another interpretation is “We know this is much more expensive, but it’s possible that some customers do value the improved performance enough to justify the additional cost. If we find that nobody wants that, we’ll shut it down, so please let us know if you value this option”.
- mechagodzilla 2y agoI think that's the right interpretation, but that's pretty weak for a company that's nominally worth $150B but is currently bleeding money at a crazy clip. "We spent years and billions of dollars to come up with something that's 1) very expensive, and 2) possibly better under some circumstances than some of the alternatives." There are basically free, equally good competitors to all of their products, and pretty much any company that can scrape together enough dollars and GPUs to compete in this space manages to 'leapfrog' the other half dozen or so competitors for a few weeks until someone else does it again.
- pinkmuffinere 2y agoI don’t mean to disagree too strongly, but just to illustrate another perspective: I don’t feel this is a weak result. Consider if you built a new version that you _thought_ would perform much better, and then you found that it offered marginal-but-not-amazing improvement over the previous version. It’s likely that you will keep iterating. But in the meantime what do you do with your marginal performance gain? Do you offer it to customers or keep it secret? I can see arguments for both approaches, neither seems obviously wrong to me. All that being said, I do think this could indicate that progress with the new ml approaches is slowing.
- asadotzler 2y agoI've worked for very large software companies, some of the biggest products ever made, and never in 25 years can I recall us shipping an update we didn't know was an improvement. The idea that you'd ship something to hundreds of millions of users and say "maybe better, we're not sure, let us know" is outrageous.
- crystal_revenge 2y ago> "We don't really know what this is good for, but spent a lot of money and time making it and are under intense pressure to announce new things right now. If you can figure something out, we need you to help us." Having worked at my fair share of big tech companies (while preferring to stay in smaller startups), in so many of these tech announcement I can feel the pressure the PM had from leadership, and hear the quiet cries of the one to two experience engineers on the team arguing sprint after sprint that "this doesn't make sense!"
- riwsky 2y ago> the quiet cries of the one to two experienced engineers on the team arguing sprint after sprint that "this doesn't make sense!" “I have five years of Cassandra experience—and I don’t mean the db”
- ummonk 2y agoConspiracy theory: they’re trying to tank the valuation so that Altman can buy it out at bargain price.
- spaceman_2020 2y agoReally don’t understand what’s the use case for this. The o series models are better and cheaper. Sonnet 3.7 smokes it on coding. Deepseek R1 is free and does a better job than any of OAI’s free models
- roarcher 2y agoAI in general is increasingly a solution in search of a problem, so this seems about right.
- TeMPOraL 2y agoOnly in the same sense as electricity is. The main tools apply to almost any activity humans do. It's already obvious that it's the solution to X for almost any X, but the devil is in the details - i.e. picking specific, simplest problems to start with.
- roarcher 2y agoNo, in the sense that blockchain is. This is just the latest in a long history of tech fads propelled by wishful thinking and unqualified grifters. It is the solution to almost nothing, but is being shoehorned into every imaginable role by people who are blind to its shortcomings, often wilfully. The only thing that's obvious to me is that a great number of people are apparently desperate for a tool to do their thinking for them, no matter how garbage the result is. It's disheartening to realize that so many people consider using their own brain to be such an intolerable burden.
- 0xDEAFBEAD 2y agoThere's a decent chance this model was originally called GPT-5, as well.
- NewUser76312 2y ago"We don't really know what this is good for, but spent a lot of money and time making it and are under intense pressure to announce new things right now. If you can figure something out, we need you to help us." Damn this never worked for me as a startup founder lol. Need that Altman "rizz" or what have you.
- financetechbro 2y agoMaybe you didn’t push hard enough the impending doom that your product would bring to society
- jcgrillo 2y agoThe fact they're raising prices so steeply is telling. This smells like desperation.