27 ms·
I feel like this just killed a few small startups who were trying to offer more context. Also, I pay for ChatGPT but I have none of the new features except for
by paulmendoza 3y ago
I feel like this just killed a few small startups who were trying to offer more context.
Also, I pay for ChatGPT but I have none of the new features except for GPT4. Very frustrating.
- arbuge 3y agoSame here: https://twitter.com/arbuge/status/1654288169397805057 https://twitter.com/arbuge/status/1654288169397805057 The really odd thing is that I was given GPT-4 with browsing alpha enabled - for a single session last week. As soon as I reloaded the page, it was gone. Since then the picture has reverted back to the above. Twitter has become a bit painful to read these days, with all the AI influencers posting about what GPT-4 and plugins, code interpreter etc. can do.
- MichaelZuo 3y agoConsidering that increasing context length is O(n^2), and that current 8k GPT-4 is already restricted to 25 prompts/3 hours, I think they will launch it at substantially higher pricing.
- choeger 3y agoIt will be interesting to see how far this quadratic algorithm carries in practice. Even the longest documents can only have hundreds of thousands of tokens, right?
- sebzim4500 3y agoIdeally you'd be able to put your entire codebase + documentation + jira tickets + etc. into the context. I think there is no practical limit to how many tokens would be useful for users, so the limits imposed by the model (either hard limits or just pricing) will always be a bottleneck.
- jtbayly 3y agoI'm confused by this. Would you want to just include your codebase, documentation, etc. in some last-mile training? That way you don't need the expense of including huge amounts of context in every query. It's baked in.
- sdenton4 3y agoYeah there's really three options here... Throw everything in context, fine tune, or add external search a la RETRO. The latter is definitely the cheapest option; updates are trivial.
- mlyle 3y agoYah... we really need some kind of architecture that juggles concept vectors around to external storage and does similarity search, etc, instead of forcing us to encode everything into giant tangles of coefficients. GPT-4 seems to show that linear algebra definitely can do the job, but training is so expensive and the model gets so huge and inflexible. It seems like having fixed format vectors of knowledge that the model can use-- denser and more precise than just incorporating tool results as tokens like OpenAI's plugin approach-- is a path forward towards extensibility and online learning.
- sebzim4500 3y agoI haven't tried this myself, but it is my understanding that finetuning does not work well in practice as a way of acquiring new knowledge. There may be a middle ground between these two approaches though. If every query used the same prompt prefix (because you only update the codebase + docs occasionally) then you could put it into the model once and cache the keys and values from the attention heads. I wonder if OpenAI does this with whatever prefix they use for ChatGPT?
- totoglazer 3y agoIt’s been available on Azure in preview. Pricing is double the 8K model.
- cubefox 3y agoO(n^2) seems unlikely: https://cognitiverevolution.substack.com/p/openais-foundry-leaked-pricing-says#:~:text=Aside,price https://cognitiverevolution.substack.com/p/openais-foundry-l.... https://news.ycombinator.com/item?id=34977194#:~:text=Sparse,%29%29 https://news.ycombinator.com/item?id=34977194#:~:text=Sparse...
- MichaelZuo 3y agoYour second link has the immediate comment "Gpt3 includes dense attention layers that are n^2". So it's not at all unlikely.
- space_fountain 3y agoGPT3 was released 3 years ago now. There have been major advancements in scaling attention so it would be strange if they didn't use some of them
- MichaelZuo 3y agoIt doesn't matter how many major advancements they made in scaling, as long as one component is O(n^2) or higher.
- cubefox 3y agoIt's not the scale itself, it's the scaling architecture.
- MichaelZuo 3y agoThe same applies.
- Keyframe 3y agosome of the context length will be lost to waste spent on truncated posts, or are replies not considered part of context on ChatGPT? In both cases, might be worth designing a prompt, every so often, to get a reply with which to re-establish the context, thus compressing it.
- tempaccount420 3y ago> current 8k GPT-4 is already restricted to 25 prompts/3 hours I'm pretty sure they're using a 4k GPT-4 model for ChatGPT Plus, even though they only announced 8k and 32k... It can't handle more than 4k of tokens (actually a little below that, starts ignoring your last few sentences if you get close). If you check developer tools, the request to an API /models endpoint says the limit for GPT-4 is 4096. It's very unfortunate.
- reaperman 3y agoAh this explains a lot. I couldn't understand why I couldn't get close to the ~12 pages that everyone was saying 8,000 tokens implied.
- tempaccount420 3y agoAs far as I know it's not documented anywhere and there is no way to ask the team at ChatGPT questions. I sent them an email about it a few days after GPT-4 release and still haven't received a reply. Another thing that annoys me is how most updates don't get a changelog entry. For whatever reason, they keep little secrets like that.
- jiggawatts 3y agoTheir PR is terrible and I get the impression that they wish their own users would “just go away”. Every time I see a company act like this, more responsive and truly open competition eventually eats their lunch.
- int_19h 3y agoThe raw chat log has the system message on top, plus "user:" and "assistant:" for each message, and im_start/im_end tokens to separate messages, hence why the visible chat context is slightly under 4k.
- YetAnotherNick 3y ago32k context is $1.92 for each request.
- achandlerwhite 3y agoIs it prorated for the actual context used for each request?
- nico 3y agoFor reference, a dev making USD 100k/year and working about 240 days a year, 8 hours/day = total of 1920 hours, or about USD 52/hour, USD 416/day 52/1.92 = 27 416/1.92 = 217 So using GPT-4 with 32k tokens, 27 times per hour, or 217 times per day, in terms of cost, is approximately the equivalent of another dev
- KeplerBoy 3y agoThat's a lot of requests. Not that it matters for the calculation, but i wonder how long such a request (ingesting 32k tokens and responding with a similar amount) would take. At the speed of regular ChatGPT take would take a good while.
- atq2119 3y agoBatch processing scales quadratically with the context size (assuming OpenAI is still using standard transformer architecture) but the batch processing of the prompt is also fast compared to generating tokens because it's batched (parallel). So I wouldn't expect effective response times to go up quadratically. At most linearly, depending on the details of how they implement inference.
- jerrygenser 3y agoPaying for chatgpt I believe is separate from API access
- maxdaten 3y agoIt is. For API access you have to create an account at https://platform.openai.com https://platform.openai.com. You pay per 1k token. For API access to GPT-4 put your organization (org id) on the waitlist.
- VeninVidiaVicii 3y agoAgain, frustrating. I’m an antibiotics researcher with oodles of data and I need ChatGPT plugins/API to make any real progress. (I’m kind of in this intellectual space on my own, so other people can’t really help that much) I’m not sure why I’ve been on the waiting list for so long now.
- ZephyrBlu 3y agoDude, chill. Plugins are insanely new. Barely anyone has access to them. It just seems like they are widespread because they've been going viral. The initial blog post was only just over a month ago, and it was announcing alpha access for a few users and developers: > Today, we will begin extending plugin alpha access to users and developers from our waitlist. While we will initially prioritize a small number of developers and ChatGPT Plus users, we plan to roll out larger-scale access over time. https://openai.com/blog/chatgpt-plugins https://openai.com/blog/chatgpt-plugins We are literally 1 month into the alpha of plugins.
- mptest 3y agoI think part of the anxiety, at least for me, is how fast progress is being made too. Can begin to feel like the "LET ME IN" meme, when you're watching all day the cool things those inside the magic shop can do lol. Layman btw just looking to use it to automate some volunteer work I do. Thanks for this perspective on how new this stuff is.
- fbrncci 3y agoThose startups will move on to open source models because OpenAI api calls with 32k token contexts are way too expensive.
- cced 3y agoWhat is the latest in conversational models that allow GPT3 like (or close) performance w.r.t running things locally?
- modernpink 3y agoGPT3 is dated so many open source models are competitive with it, but Vicuna 13b is supposed to be competitive with GPT4
- speedgoose 3y agoAgainst GPT3.5 perhaps the gaps aren’t too big for your use cases, but I wouldn’t say it’s in the GPT4 league. It looks close in the benchmarks but the difference in quality feels (to me) huge in practice. The others models are simply a lot worse.
- modernpink 3y agoInteresting. Have you tried StableVicuna?
- speedgoose 3y agoNo, is it worth a try? I didn’t see a lot of hype about it so I didn’t try it.
- noman-land 3y agoApparently Vicuna 13B is quite good according to Google's own leaked docs. https://twitter.com/jelleprins/status/1654197282311491592 https://twitter.com/jelleprins/status/1654197282311491592
- decompiled_dev 3y agoI was waiting for a while, but then I found there was a page where if you selected "I want to build plugins" then you would have never seen the option to request them. Once I filled that in I got access within a few days. https://openai.com/waitlist/plugins https://openai.com/waitlist/plugins If you are the person to say: "I am a developer and want to build a plugin" Then it is likely you missed the option to request which plugins you want access to.
- TeMPOraL 3y agoDepends on the use case. Performance quickly tanks when you get to high token count; it's a slowdown I believe the various summarizers/context extenders mostly avoid. (Also UI probably tanks too. I dread what the OpenAI Playground will do when you start actually using 32k model for real, like throwing a 15k token long prompt at it. ChatGPT UI has no chance.)
- toxicFork 3y agoIt's Hella expensive so I think they are ok for now Until they cut down the cost then they should worry yeah
- fakedang 3y agoHonestly for the firms that would use it, for example finance or legal, it's very reasonable.
- HarHarVeryFunny 3y agoMosaicML StoryWriter 65K model just released a day or two ago.
- chrisMyzel 3y agohttps://www.mosaicml.com/blog/mpt-7b https://www.mosaicml.com/blog/mpt-7b 65k+ context window, open source, open weights
- chillfox 3y agoSame. No plugins or GPT-4 API for me despite signing up for the waiting lists on the day they were announced.
- saulpw 3y agoHave you been using the API with GPT-3.5? I wonder if they're prioritizing access to 'active' users who appear to be trying to make something with it, over casual looky-loos.
- ShamelessC 3y ago> I feel like this just killed a few small startups who were trying to offer more context. Those startups killed themselves. A 32K context was advertised as a feature to be rolled out the same day GPT-4 came out. Also - what startups are getting even remotely close to 32K context at GPT-4’s parameter count? All I’ve seen is attempts to use KNN over a database to artificially improve long term recall.
- danjc 3y agoTry OpenAI services in Azure. We were added to a waitlist but got approved a week later. Had 32k for a few weeks now but still on the waitlist for plugins.