23 ms·
OpenAI's plans according to sama
- heyzk 3y agoGreat writeup, this helps us understand where to spend our time vs what OpenAI's progress will solve.
- boringuser2 3y ago>Dedicated capacity offering is limited by GPU availability. OpenAI also offers dedicated capacity, which provides customers with a private copy of the model. To access this service, customers must be willing to commit to a $100k spend upfront. How many shell corporations are intelligence agencies seeding right now?
- cwkoss 3y agoLast night I was musing how many different countries' intelligence agencies have moles working at OpenAI currently. Gotta be at least 6, maybe as high as two dozen?
- boringuser2 3y agoAgent Lee Chen Huwang, reporting for duty.
- m3kw9 3y agoI bet the NSA has dossier on every employee there as well
- boringuser2 3y ago"Cooperate or we'll kill your family". (Just to be clear, this is a hypothetical intelligence agent saying this, not me.) I mean, it's not exactly rocket science, who wouldn't instantly fold to that?
- layer8 3y agoSomeone without family?
- boringuser2 3y agoYou know the next step, right?
- CSMastermind 3y agoUS, France, Israel ... then who? Maybe another five eyes country like the UK? Possibly China? I'm pretty skeptical Russia would be able to get someone in there but maybe.
- invaliduser 3y agoHi. French here. I may be wrong, but I really feel like you are overestimating us.
- CSMastermind 3y agoDGSE essentially puts all of it's money/effort into industrial espionage and they're the best in the world at it.
- boringuser2 3y agoYou said Isr*el twice.
- m3kw9 3y agoThey are not gonna give the weights for sure but it still will be inferencable, I’m not sure how but it’s be self destructive if they did
- boringuser2 3y agoExactly, with a private model you could easily extract the weights.
- ftxbro 3y agoI had been putting theories in comments but they kept getting flagged or banned or downvoted to oblivion, but maybe its time has come. I'll keep it tame. If you are curious you can google connections of OpenAI board of directors, Will Hurd, In-Q-Tel trustees, Allen and Company, etc. There is more but whatever. The conspiracy theory is that 'the govt stepped in' during the six month pause after gpt-4 was trained and before it was released.
- refulgentis 3y agoIt probably keeps getting flagged because it’s ahistorical, source: OpenAI engineers, and #2 somewhat obviously so. You heard of RLHF?
- ftxbro 3y ago> You heard of RLHF? The conspiracy theory isn't that every employee of OpenAI spent 8 hours every day for six months in meetings with govt agencies.
- refulgentis 3y agonot sure what you mean. anyways, the reason why they don’t release GPT4 when they’re “”“done””” training in June is they have to RLHF
- MacsHeadroom 3y agoPrivate instance means a dedicated endpoint fully managed by OpenAI. You do not get model access or anything a regular API user doesn't already get, except your API url will be something like customer123.openai.com/api instead of api.openai.com/api
- ryanmercer 3y agoWhy bother with shell corps when they already back companies in the clear: look at In-Q-Tel.
- hervature 3y agoI never know if I have an inside scoop or an outside scoop. Has Hyena not addressed the scaling of context length [1]? I know this version is barely a month old but it was shared to me by a non-engineer the week it came out. Still, giving interviews where the person takes away that the main limitation is context length and requires a big breakthrough that already happened makes me seriously question whether or not he is qualified to speak on behalf of OpenAI. Maybe he and OpenAI are far beyond this paper and know it does not work but surely it should be addressed? [1] - https://arxiv.org/pdf/2302.10866.pdf https://arxiv.org/pdf/2302.10866.pdf
- deleted 3y ago[deleted]
- arugulum 3y agoAs someone who is in the field: papers proposing to solve the context length problem come out every month. Almost none of the solutions stick or work as well as a dense or mostly dense model. You'll know when the problem is solved when model after consistently use a method. Until then (and especially if you're not in the field as a researcher), assume that every paper claiming to tackle context length is simply a nice proposal.
- dr_dshiv 3y agoWhat about Meta’s megabyte? Also nice proposal?
- visarga 3y agoYes. Solving context length has been tried in hundreds of different approaches, and yet most LLMs are almost identical to the original one from 2017. Just to name a few families of approaches: Sparse Attention, Hierachical Attention, Global-Local Attention,Sliding Window Attention, Locality sensitive hashing Attention, State space model, EMA gated attention.
- Loquebantur 3y ago
- londons_explore 3y ago> is limited by GPU availability. Which is all the more curious, considering OpenAI said this only in January: > Azure will remain the exclusive cloud provider for all OpenAI workloads across our research, API and products [1] So... OpenAI is severely GPU constrained, it is hampering their ability to execute, onboard customers to existing products and launch products. Yet they signed an agreement not to just go rent a bunch of GPU's from AWS??? Did someone screw up by not putting a clause in that contract saying "exclusive cloud provider, unless you cannot fulfil our requests"? [1]: https://openai.com/blog/openai-and-microsoft-extend-partnership https://openai.com/blog/openai-and-microsoft-extend-partners...
- londons_explore 3y agoPerhaps they are cash flow constrained, which in turn means they are GPU constrained, since GPU's are their biggest expense?
- catchnear4321 3y agothis has nothing to do with sama clamoring for regulation. that absolutely isn’t an attempt to slow down all competition. which isn’t necessary because nobody made such a mistake. this won’t lead to any hasty or reckless internal decisions in a feckless effort to stay in front. not that any have already been made. not that that could lead to disaster.
- chaostheory 3y agoEven if they weren’t exclusive with Azure, aren’t GPU prices reasonable again?
- verdverm 3y agoThey have to be a available to buy, regardless the price. My understanding is there is a distinct lack of supply
- zamalek 3y agoBarring a revolution in chip manufacture, there likely will always be a lack of supply relative to consumer GPUs. The size of the die results in terrible yields.
- dontupvoteme 3y agoHaving been recently taken aboard by the mothership I expect they'll start trying to tune out anything related to programming to push people towards co-pilot X.. It's pretty hilarious and annoying to see bing start to write code only to self censor itself after a few lines (deleting what was there! no wonder these guys love websockets and dynamic histories) Whoops!
- _boffin_ 3y agoWait… what? Can you elaborate.
- mistymountains 3y agoHe’s speculating that Microsoft is nerfing OpenAI / chatGPT to funnel narrow capabilities to silos like CoPilot.
- _boffin_ 3y agoI understand that... I should have specified a bit more that i'm interested in knowing more about the removal of answers as its writing them, if they're code.
- dontupvoteme 3y agohttps://en.wikipedia.org/wiki/Embrace,_extend,_and_extinguish https://en.wikipedia.org/wiki/Embrace,_extend,_and_extinguis...
- _boffin_ 3y agoyes... I know about this, but that's not what I'm asking about. I'm asking about it removing partial answers as it's writing them. Please make more effort next time than to provide me with a Wiki article.
- dontupvoteme 3y ago
- sovietmudkipz 3y agoI’m hoping GPT will remove the information cutoff date. I write plenty of terraform/AWS and it’s a bit of a pain that the latest API isn’t accessible by GPT yet. There’s been quite a bit happening in the programming space since sept 2021. I use GPT to keep things high level and then do my normal research methodology for implementation details.
- furyofantares 3y agoIt's not like an arbitrary imposition, that's the data it was trained on and it's expensive to train. I hope they find a way to continually train in new information too but it's not like they can just remove the cutoff date.
- kmod 3y agoNot disagreeing, but a fascinating thing they did (as a one-off fine-tune?) was teach ChatGPT about the openai python client library, including the features that were added after the cutoff date.
- mustacheemperor 3y agoI enjoy using GPT4 as a co-programmer, and funny enough it is very challenging to get advice on Microsoft's own .NET MAUI because that framework was in prerelease at the time the model was trained. My understanding is right now they essentially need to train a new model on a new updated corpus to fix this, but maybe some other techniques could be devised...or they'll train something more up to date.
- buildbot 3y agoContext drift! https://qntm.org/mmacevedo https://qntm.org/mmacevedo
- ilaksh 3y agoYou might actually get pretty far if you just went through the Microsoft docs and created a bunch of really concise examples and fed that as the start of the prompt. Use like 6-7kb for that and then the question at the end.
- atemerev 3y agoAll AI companies (OpenAI included) are now working full tilt on making AIs improve themselves (writing their own code, inventing new pipelines etc). I don't know why choose anything else to work on. This is a prime directive, that will bring the greatest payoff.
- huijzer 3y agoI disagree since GPUs are a major constraint currently and that skilled specialists outperform GPT-4 almost always as long as they stay in their domain. Will they use copilot(s) to improve the models? Yes, but they have been doing that since 2021 already (the release year of GitHub Copilot).
- ren_engineer 3y agoif this was currently possible wouldn't it lead to sentient/superhuman AI rapidly? >tell AI to make itself more efficient by finding performance improvements in human written code >that newly available processing power can now be used to find more ways to improve itself >flywheel effect of AI improving itself as it gets smarter and smarter eventually you'd turn it loose on improving the actual hardware it runs on. I think the question now is really how far transformers can be taken and if they are really the path to "real" AI.
- ilaksh 3y agoWithin a couple of years of improvement processes like you suggest will actually be really dangerous and stupid. Also don't confuse all other types of human/animal characteristics like sentience with intelligence. They are different things. Things like sentience, subjective stream of experience, or other aspects of being alive don't just accidentally fall out of larger training datasets. And we should be glad. The models are going to be orders of magnitude faster (and perhaps X times higher IQ) than humans within a few years. It is incredibly foolish to try to make something like that into a living creature (or emulation of living).
- visarga 3y ago
- yesimahuman 3y agoThe bit about plugins not having PMF is interesting and possibly flawed. I, like many others, got access to plugins but not the browsing or code interpreter plugins which feel like the bedrock plugins that make the whole offering useful. I think there's also just education that has to happen to teach users how to effectively use other plugins, and the UX isn't really there to help new users figure out what to even do with plugins.
- furyofantares 3y agoPMF meaning "product market fit"? I had to look it up, curious if I found the right thing or not.
- gistbug 3y agoYea, seems weird to allow people to use plugins, but not all of them. Then have the gall to say that no one is using plugins, yea because half of them don't have any context outside of America.
- lumost 3y ago
- cryptoz 3y agoGreat content and great answers except for open source question. Sam is saying that he doesn’t think anyone would be able to run the code at scale so they didn’t bother? Seems like a nonsense answer, maybe I’m misunderstanding. The ability for individuals or businesses to effectively run and host the code shouldn’t have an impact on the ability to open source.
- cwkoss 3y agoI love the tongue-in-cheek paradox myth that the Bitcoin whitepaper was written by a future god-AI to increase demand for GPUs (and thus boost supply) so we are able to assemble the future god-AI.
- barbazoo 3y agoI'd read that book!
- baq 3y agoWatch Tenet.
- skulk 3y agoI'm from the future, traveling backwards in time to tell you to not watch Tenet.
- barbazoo 3y agoSomehow I'm super sensitive to the audio (or might be video) and start feeling nauseous after a short time. Is there an explanation to this? I think it's that scratchy humming background sound.
- skulk 3y agoI don't know, except that I agree that it's horribly mixed. It's almost impossible to make out what people are saying in certain parts, and that causes cognitive load that degrades the experience. Also, the story is completely nonsensical. The only good thing about that movie is that time-reversed fight scene and even that was kind of questionable.
- KineticLensman 3y agoAlso lots of similar wisdom, from Percival Dunwoody, Idiot Time Traveller from 1909 [0]. [0] https://www.gocomics.com/tomthedancingbug/2022/06/17 https://www.gocomics.com/tomthedancingbug/2022/06/17
- asnyder 3y agoHis statements on open sourcing in this interview/write-up is somewhat in conflict with his recent statement made last week in Munich https://youtu.be/uaQZIK9gvNo?t=1170 https://youtu.be/uaQZIK9gvNo?t=1170, where he explicitly said the Frontier of GPT won't be open sourced due to what they perceive as safety reasons, https://youtu.be/uaQZIK9gvNo?t=1170 https://youtu.be/uaQZIK9gvNo?t=1170 (19:30 - 22:00).
- ftxbro 3y agoit's legal to make contradictory statements that's one of the job of a ceo and it's why they aren't usually overly literal types you know the kind i'm talking about
- muskmusk 3y agoI don't see the conflict. They see current models as mostly harmless, but what comes next is dangerous. It sounds a little too much sci-fi for me, but I guess he knows better.
- wintogreen74 3y agoplus this conveniently pairs with "we don't need to regulate current models, but future models... oh boy do those need to be regulated!"
- ilaksh 3y agoHe was talking about open sourcing GPT-3. That is not the frontier. The frontier is the multimodal versions of GPT-4 which he just said wasn't even going to public release until next year. Or whatever they are on now which they are carefully not calling GPT-5.
- sharkjacobs 3y ago> He reiterated his belief in the importance of open source and said that OpenAI was considering open-sourcing GPT-3. Part of the reason they hadn’t open-sourced yet was that he was skeptical of how many individuals and companies would have the capability to host and serve large LLMs. Am I reading this right? "We're not open sourcing GPT-3 because we don't think it would be useful to anyone else"
- greenie_beans 3y agolmao i had the same reaction. sounds like some bullshit.
- bibanez 3y agoI agree, this is so bizarre
- ftxbro 3y agoyes i also can't wrap my head around how a ceo of a billion dollar company isn't sincere in his public statements
- wintogreen74 3y agoReally? Even after saying this? "While Sam is calling for regulation of future models, he didn’t think existing models were dangerous and thought it would be a big mistake to regulate or ban them."
- cinntaile 3y agoIt was a tongue in cheek reaction.
- ethanbond 3y agoWhy couldn’t that be true? E.g. even scientists who worked on the Manhattan Project (justifiably) had antipathy toward the much more powerful hydrogen bomb. It’s possible to think squirt guns shouldn’t be regulated but AR-15s should, or AR-15s shouldn’t but cruise missiles should. Or driving at 25mph should be allowed but driving 125mph shouldn’t.
- deleted 3y ago[deleted]
- purplecats 3y ago> Cheaper and faster GPT-4 — This is their top priority. In general, OpenAI’s aim is to drive “the cost of intelligence” down as far as possible and so they will work hard to continue to reduce the cost of the APIs over time. this certainly aligns with the massive (albeit subjective and anecdotal) degradation in quality i've experienced with ChatGPT GPT-4 over the past few weeks. hopefully a superior (higher quality) alternative surfaces before its unusable. i'm not considering continuing my subscription at this rate.
- nonethewiser 3y agoI wonder of it actually is because they’re tuning it to make it less offensive (by their standards). Thats the only explanation I keep seeing repeated.
- reaperman 3y agoI would be very surprised. Things that are very, very far from that are also much worse. I'm having difficulty finding the difference between GPT-3.5 and GPT-4 for a lot of my programming tasks lately. It's noticeably degraded.
- littlestymaar 3y agoThat's a convenient explanation that's been repeated over and over by certain people, but cost is a much more likely explanation: inference for large models is extraordinarily expensive when you have millions of users and their pricing model always seemed way too low to pay for that. They have likely been subsidizing their users since the launch of their commercial offering (and this is pretty common strategy for SV startups) but they've been so successful that they now need to scale the cost down in order not to burn all their cash too fast.
- brucethemoose2 3y agoAnthropic's Claude is said to be very good. Instruction tuned LLaMA 65B/Falcon 40B are good, especially with an embeddings database. ...But OpenAI has all the name recognition and ease of use now, so it might not even matter if others ambiguously surpass OpenAI models.
- ftxbro 3y agowhy should I believe what someone says their plans are
- flakiness 3y agoIf they open up fine-tuning API for their latest models, I wonder how the enthusiasm around the open source model is impacted. One of the advantages of the open source models is the ability to be fine-tuned. Are other benefits enough to keep the momentum going?
- m3kw9 3y agoYou better have deep pockets, have you check the prices and then the rates for using the tuned models? They sure 10x to 100x more expensive then nontuned models
- sashank_1509 3y agoReally great news to give at cheaper and faster GPT4. As a GPT+ subscriber, the most annoying thing is the 25 message limit every 3 hours, I really want that removed. A bit sad to hear that the multimodal model will only come next year, was hoping to get it this year 100k to 1 Million context length, sounds phenomenal especially if it comes to GPT4. I've used Claudes 100k context length and I found it so useful that when I have large documents I just default to Claude now
- alchemist1e9 3y agoDid you have any tips on how you got access to Claude? I do the request access but never get any email or any contact.
- refulgentis 3y agoPoe, I’m in the same boat btw
- sashank_1509 3y agoI'm a graduate student doing AI research in a US University. And I applied pretty early (Last year December I think) , those might be two factors that got me access to Claude. I think getting access to Claude through slack is much easier and I recently got it by just downloading it as a Slack App
- floydnoel 3y agoI use Poe and got access to Claude 100k as soon as it was released. I think it's a better deal than paying OpenAI for sure, since you have access to GPT-4, Claude+, and others. They also have community bots, etc.
- teetertater 3y agoClaude has a free slack client that I briefly was able to access by creating a new slack workspace and adding it there. But as of yesterday it wasn't working for me
- m3kw9 3y agoIf you look at their API limits, no serious company can use this to scale up beyond say 10k users. 3500 Reqs per min for gpt3.5 turbo. They have a long way to go to make it usable for the rest of the 95%
- thorax 3y agoI've had to move to using Azure OpenAI service during business hours for the API-- much more stable unless the prompts stray into something a little odd and their API censorship blocks the calls.
- nonfamous 3y agoYou can opt out of the safety filtering, btw.
- legendofbrando 3y agoI’ve been working directly with OpenAI’s access, are there any other advantages to doing this through Azure?
- BeenAGoodUser 3y agoNice to see they are working on reducing the pricing. GPT-4 is just too expensive right now imo. A long conversation would quickly end up costing tens of dollars if not more, so less expensive model costs + stateful API is urgently needed. I think even OpenAI will actually gain a lot by reducing the pricing, right now I wouldn't be surprised if many uses of GPT-4 weren't viable just because of the costs.
- Terretta 3y agoThis is off by probably x10 or more. Dozens of people using it daily for coding and conversations and review in a month might be a couple hundred bucks. All day convo, constantly, as fast as it can respond, might add up to $5. Not sure what kind of convo you're having that you could hit $10 unless you're parallelizing with something like the "guidance" tool or langchain.
- refulgentis 3y agoAbsolutely not. Dinner just got here, but tl;dr gpt4 is 0.03 per 750 words in 0.06 per 750 words out. People except the history to be included as well
- jiggawatts 3y agoThe version of GPT 4 with 32K token context length is the enabler for a huge range of "killer apps", but is even more expensive than the 8K version. And yes, parallelism and loops are also key enablers for advanced use-cases. For example, I have a lot of legacy code that needs uplifting. I'd love to be able to run different prompts over reams of code in parallel, iterating the prompts, etc... The point of these things is that they're like humans you can clone at will. The ability to point thousands of these things at a code base could be mindblowing.
- naillo 3y ago> Plugins “don’t have PMF” Probability mass functions? Anyone know what this means in this context?
- simonbutt 3y agoProduct market fit
- deleted 3y ago[deleted]
- dr_dshiv 3y ago> OpenAI will avoid competing with their customers — other than with ChatGPT. Quite a few developers said they were nervous about building with the OpenAI APIs when OpenAI might end up releasing products that are competitive to them. Sam said that OpenAI would not release more products beyond ChatGPT. He said there was a history of great platform companies having a killer app and that ChatGPT would allow them to make the APIs better by being customers of their own product. The vision for ChatGPT is to be a super smart assistant for work but there will be a lot of other GPT use-cases that OpenAI won’t touch. Can anyone elaborate on this? This is a big issue for me.
- ilaksh 3y agoI think the tricky part for me is that "work" is extremely broad and now that ChatGPT has plugins, it can kind of do anything. Heh.
- jiggawatts 3y agoIs this guy Aes Sedai? Technically he can claim that OpenAI will not release competing products while Microsoft plugs AI into everything. Microsoft just announced at Build 2023 that they'll have OpenAI tech integrated with: Windows, Bing, Outlook, Word, Teams, Visual Studio, Visual Studio Code, Microsoft Fabric, Dynamics, GitHub, Azure DevOps, and Logic Apps. I probably missed a bunch. Very soon now, everything Microsoft sells will have OpenAI integration. Unless you're selling a niche product too small for Microsoft to bother with, you're competing directly against OpenAI. Oh, and to top it off: Microsoft can use GPT 4 all they want, via API access. Third parties have to beg and plead to get rate-limited access. That access can be withdrawn at any time if you're doing something unsafe to OpenAI's profit margins. "Please Sir Sam, may I have some GPT please?" "No."
- edanm 3y ago> Is this guy Aes Sedai? Haha having just finished the Wheel of Time, I'm super tickled by this reference. It doesn't seem to be too common, only two uses of it on HN in the past year (at least, found by searching for the phrase "Aes Sedai")
- simse 3y ago> A stateful API This would be huge for many applications, as "chatting" with GPT-4 gets really, really expensive very quickly. I've played with API with friends, and winced as I watched my usage hit several dollars for just a bit of fun.
- vb-8448 3y ago7. find a sustainable business model and make some money
- jasmer 3y agoIt's absurd that people are still thinking that a language model which a bunch of tokens are indexed is some kind of 'AGI'.
- bobbyi 3y agoThe roadmap here is completely focused on ChatGPT and GPT-4. I wonder what portion of their resources is still going to other areas (DALL-E, audio/ video processing, etc.)
- cmelbye 3y agoMaybe some of those things that are currently separate projects will eventually converge with a multimodal model.
- hoschicz 3y ago- they are working on a stateful API - they are working on a cheaper version of GPT-4 Most probably this is driven by their use of it in ChatGPT, which is on fire from PMF. Clearly they're experimenting with the cheaper GPT-4 in ChatGPT right now as it's fairly turbo now, as discussed earlier today.
- twobitshifter 3y agoLeft the best part until the end. Scaling models larger is still paying off for openai. It’s not AGI yet, but how much bigger will a model need to get to max out? >The scaling hypothesis is the idea that we may have most of the pieces in place needed to build AGI and that most of the remaining work will be taking existing methods and scaling them up to larger models and bigger datasets. If the era of scaling was over then we should probably expect AGI to be much further away. The fact the scaling laws continue to hold is strongly suggestive of shorter timelines.
- tikwidd 3y agoAfter training my physics simulator on thousands of hours of video footage of trees moving in the wind, arborists tell me the trees are much more realistic (they are getting worried that I might put them out of business). But the physicists are still not satisfied. How many more videos do I need to generate the laws of motion?
- bigyikes 3y agoThrow in the videos from the rest of the internet, and you might actually do it…
- bongobingo1 3y ago> The simulations incredible, but I have to ask, why do the trees all have breasts?
- jasmer 3y ago'It's not AGI yet' - the implication is insufferable. It's a language model that is incapable of any kind of reasoning, the talk of 'AGI' is a glib utopianism, a very heavy kind of koolaid. If we were to have referred to this tech as anything other than 'intelligence' - for example, if we chose 'adaptive algorithms' or 'weighted node storage' etc. we'd likely have a completely different popular mental model for it. There will be no 'AI model' that is 'AGI', rather, a large swath of different technologies and models, operating together, will give the appearance of 'AGI' via some kind of interface. It will not appear as an 'automaton' (aka single processing unit) and it certain will not be an 'aha moment'. In 10 years, you'll be able to ask various agents, of different kinds, which will use varying kinds of AI to interpret speech, to infer context, which will interface with various AI APIs, in many ways it'll resemble what we have today but with more nuance. The net appearance will evolve over time to appear a bit like 'AGI' but there won't be an 'entity' to identify as 'it'.
- jeffybefffy519 3y agoTitle needs an update as Sama is also the name of the company which helped classify training data for ChatGPT: https://time.com/6247678/openai-chatgpt-kenya-workers/ https://time.com/6247678/openai-chatgpt-kenya-workers/
- pindab0ter 3y agoI agree. Not using the actual name is also gate keeping for anyone not familiar enough. The fact that “sama” isn’t even capitalised adds to this.
- Sharlin 3y agoNever mind the fact that it’s against HN guidelines to modify original titles for no reason. Changing Sam Altman to sama is just ridiculous.
- Imnimo 3y ago>The fact that scaling continues to work has significant implications for the timelines of AGI development. The scaling hypothesis is the idea that we may have most of the pieces in place needed to build AGI and that most of the remaining work will be taking existing methods and scaling them up to larger models and bigger datasets. If the era of scaling was over then we should probably expect AGI to be much further away. The fact the scaling laws continue to hold is strongly suggestive of shorter timelines. If you understand the shape of the power law scaling curves, shouldn't this scaling hypothesis tell you that AGI is not close, at least via a path of simply scaling up GPT-4? For example, the GPT-4 paper reports a 67% pass-rate on the HumanEval benchmark. In Figure 2, they show a power-law improvement on a medium-difficulty subset as a function of total compute. How many powers of ten are we going to increase GPT-4 compute by just to be able to solve some relatively simple programming problems?
- sanxiyn 3y agoSomeone did that calculation and the result is here: https://www.reddit.com/r/slatestarcodex/comments/13u40yf/ https://www.reddit.com/r/slatestarcodex/comments/13u40yf/ 100x GPT-4 to 85%.
- Imnimo 3y agoAnd, if I'm reading their calculation right, that's 85% on the medium-difficulty bucket, not even the entire HumanEval benchmark? (quoting from the GPT-4 paper): >All but the 15 hardest HumanEval problems were split into 6 difficulty buckets based on the performance of smaller models. The results on the 3rd easiest bucket are shown in Figure 2
- PoignardAzur 3y agoThat does seem to support the idea that we're two or three major breakthroughs away from superintelligent AGI, assuming these scaling curves keep holding as they have.
- askkk 3y agoI always enjoy reading some of your comments, they ameliorate the hype about LLM and give a critical review. Anyway, I think a stronger model than GPT-4 could improve the way to use tools, so that the model is able to self-improve using tools. For example using all kind of solvers and heuristics to guide the model. I don't know how to estimate that risk just now. Edited: Don't know if is a good thing to study the weak points of closed LLMs. Even asking LLMs can give hints about possible ways to improve. In my case I am happy I am certainly old and my mind is a lot weaker than before, but even in this case I prefer not to use LLMs for gaining insight because she will someday get a better insight than myself. But the lust of knowledge is a mortal sin.
- throwaway1777 3y agoSurprised no mention of them developing their own chips.
- artichokeheart 3y agoOff-topic note Humanloop might want to redesign their logo. It's been the Australian Broadcasting Corporation logo since 1963. Maybe pick a different Lissajous curve.
- udev4096 3y agoIf they open source it, everyone would know that they used fuck ton of pirated content to train their models
- andybak 3y agoAs far as I'm aware training does not currently constitute "piracy". It's fine to advocate for a redefinition but be explicit about it.
- gnomewascool 3y agoI think the point here is about the procurement of the training data, in violation of copyright laws ("piracy"), rather than that the training itself is piracy. The suspicion[0] is that OpenAI trained their models on a large text dump including libgen (in the so-called "books2"). If a person downloads a book from Library Genesis, they're a pirate; if OpenAI does it, so are they. [0] https://twitter.com/theshawwn/status/1320282152689336320 https://twitter.com/theshawwn/status/1320282152689336320
- braindead_in 3y agoI am looking forward to faster GPT-4, larger context windows and finetuning APIs. The combination of these can solve most of problems that I currently face with my LLM apps. It looks like a good roadmap for 2023.
- jamesfisher 3y agohttps://archive.ph/uwaCp https://archive.ph/uwaCp (original page is now a 404)
- deqwer 3y ago[dead]
- dbmnt 3y agoWould someone please explain like I'm five which components of LLM's like ChatGPT are still closed source? What are the specific technologies that OpenAI is holding on to? I know a lot of LLM stuff has either been released or leaked out, but don't have enough expertise in this area to understand the competitive advantages or breakthroughs OpenAI has obtained.
- pnt12 3y agoAs far as I understand, it's mostly the weights. If you only have the models, you're still gonna need to get a massive amount of training material, huge costs for training it and fine tuning the hyper params until it works.
- pavelstoev 3y agoThis content has been removed at the request of OpenAI. Just when I went back to the post for some quote material...
- webmaven 3y ago"This content has been removed at the request of OpenAI."
- vagab0nd 3y agohttps://archive.ph/rcbem https://archive.ph/rcbem The page now says "This content has been removed at the request of OpenAI." I wonder why they did it.
- bodecker 3y agoRelated thread "OpenAI's plans according to Sam Altman removed at OpenAI's request": https://news.ycombinator.com/item?id=36177895 https://news.ycombinator.com/item?id=36177895