78 ms·
GPT-4 details leaked?
- shahules 3y agoThis guy doesn't have any idea what he is talking about. He consistently posts such bullshit on twitter. Mostly copy paste with added spice mix.
- mk_stjames 3y agoI noted several things that don't seem consistent with what people have been assuming from before. For instance - MoE yes, but 16 experts at 111B parameters? Doesn't make sense. GPT 3 had 175B parameters. I doubt they would go less on base models from now on. The number that makes more sense is ~220B parameters per model and 8 expert models. That is the same inference cost in total. The 13T tokens of training data seems pulled from thin air.
- mt_ 3y agoIt's Twitter, why would you think otherwise?
- s3p 3y agoI'm tired of people just saying 'oh, internet' every time there is a factual inaccuracy somewhere. Yes, we know this is a social media network. Now can we get back to discussing the topic at hand?
- DanAtC 3y agoThese words are nonsense to me. Can someone explain?
- esafak 3y agoGPT-4 is the name of a machine learning (language) model that is the basis of chatGPT. This post speculates about its internals. https://en.wikipedia.org/wiki/GPT-4 https://en.wikipedia.org/wiki/GPT-4
- DanAtC 3y agoI meant things like: - parameters - layers - "Mixture Of Experts" - tokens That's about as far as I made it
- DANmode 3y agoParameters: In the context of AI and language models, parameters refer to the internal settings or variables that an AI model uses to make predictions or generate responses. Think of them as the knobs and switches that can be adjusted to fine-tune how the AI understands and generates language. These parameters are learned during the training process, where the AI model analyzes vast amounts of data to optimize its performance. Layers: In AI, layers are like stacked building blocks within a neural network, which is the fundamental structure of many AI models. Each layer performs different computations, transforming the input data as it passes through them. Think of a layer as a specific task or filter that the AI model can utilize to understand and process information. The deeper the neural network, the more layers it has, allowing for more complex patterns and representations to be learned. "Mixture Of Experts": An "MoE" is an approach in AI that combines multiple specialized AI models, known as "experts," to work together on a task. Each expert focuses on a particular subset or aspect of the problem, leveraging their expertise to contribute to the final result. It's like having a team of experts who specialize in different areas collaborating to provide the best solution. By dividing the task and letting each expert handle their niche, the AI model can achieve better overall performance. Tokens: In the context of AI and language models, tokens are chunks of text that are used as input or output. They can be individual words, characters, or even subwords, depending on how the language model is designed. For example, in the sentence "I love cats," the tokens would be "I," "love," and "cats." Tokens help the AI model understand and process language by breaking it down into manageable units. They allow the model to learn patterns, context, and relationships between words to generate meaningful responses or predictions. See: https://platform.openai.com/tokenizer https://platform.openai.com/tokenizer
- TeMPOraL 3y ago> "Mixture Of Experts": An "MoE" is an approach in AI that combines multiple specialized AI models, known as "experts," (...) Wonder when that stopped being called just an "ensemble model", which is a term I recall from 10 years ago. Terminology churn?
- CSMastermind 3y agoPreviously posted about here: https://news.ycombinator.com/item?id=36671588 https://news.ycombinator.com/item?id=36671588 and here: https://news.ycombinator.com/item?id=36674905 https://news.ycombinator.com/item?id=36674905 With the original source being: https://www.semianalysis.com/p/gpt-4-architecture-infrastructure https://www.semianalysis.com/p/gpt-4-architecture-infrastruc... The twitter guy seems to just be paraphrasing the actual blog post? That's presumably why the tweets are now deleted. --- The fact that they're using MoE was news to me and very interesting. I'd love to know more details about how they got that to work. Variations in that implementation would explain the fluctuations in the quality of output that people have observed. I'm still waiting for the release of their vision model which is mentioned here but we still know little about, sans a few demos a few months ago.
- jph00 3y agoThe previous posts are to a twitter thread that's been taken down, and the preview of a post that requires a $1000 subscription. This post however is freely available (for now at least).
- londons_explore 3y agoAnd the tweeter of the twitter thread paid the $1000, copied the useful info to twitter, and then did a credit card chargeback.
- renlo 3y agoSeems he summarized it and didn't copy it
- londons_explore 3y agoA summary isn't allowed under US copyright law. The copyright office calls them "condensations", and they are considered derivative works. His use was likely not within US copyright law. "Effect of the use upon the potential market for or value of the copyrighted work" is one of four factors a judge should use to decide if fair use applies, and it is clear that publishing the main information from an article, information which is not available elsewhere, freely, severely degrades the market for the original.
- deleted 3y ago[deleted]
- neonate 3y agohttps://archive.ph/2RQ8X https://archive.ph/2RQ8X
- refulgentis 3y agoNo, this is fake, a light dusting of nothing on top of a meme post that was circulating in grifting communities as early as Q4 2022. It gains a little bit in every retelling, sort of impressive to see its almost blog scale.
- tills13 3y agoThat explains why an ad when I tried to click through to the tweet.
- jph00 3y agoNo, a meme post from 2022 did not in fact reference papers posted in 2023. You must be thinking of some other post, or you're just making stuff up.
- refulgentis 3y agoLike I said, it gains a little bit in every retelling. Why are you aggressively defending unsourced tripe?
- swyx 3y agowell as someone out of the loop - is there a source on the Q4 2022 version then?
- ompto 3y agoNot sure about the Q4 2022 version but there was a post [1] a month ago that also claimed something like 16 MOE that got a lot of attention, and some similar-ish rumors before too with less detail. So could be either more detail leaking over time or just a random made up post that took root a long time ago continually being retold with slightly more guesstimated detail added own each time to make the poster sound like they're in the know. Impossible to tell until the actual details are confirmed I guess. [1] https://news.ycombinator.com/item?id=36413296 https://news.ycombinator.com/item?id=36413296
- eminence32 3y ago> It is over. What does this mean?
- purplecats 3y agopresumably the speculation
- aaronbrethorst 3y agoIt's the tweet equivalent of one of those obnoxious YouTube 'reaction' thumbnails.
- echelon 3y agoIf this is true, anyone [1] can now build a GPT-4 given training data and budget. There's no magic here. [1] That's probably twenty or so orgs right now, which will blow away OpenAI's moat and margins.
- Gigachad 3y ago“It’s over” is the latest meme phrase to describe some kind of defeat.
- potatoman22 3y agoGoogle has been doing research into mixture of experts for scaling LLMs. Their GLaM model published in 2022 has 1.7 trillion parameters and 64 experts. https://icml.cc/media/icml-2022/Slides/17378.pdf https://icml.cc/media/icml-2022/Slides/17378.pdf
- behnamoh 3y agoGoogle is jokingly behind in terms of LLMs. They've done a pretty good job at incorporating vision and audio ML models into their ecosystem, but they underestimated language.
- chucknthem 3y agoHow do you know? do you have insider knowledge of this or is it just based on what they share publically?
- seanthemon 3y agoFrom what I see in that GPT 3 and 4 was a bit of a rugpull for the industry, now we're all laughing at Google because seemingly they had their hands on the rug for nearly a decade and did nothing - but from the other perspective, maybe they saw the future openai has now brought us and decided against being the pioneers
- tiffanyg 3y agoAmazingly enough, I think this is a bit of it. Some powerful enough people at Google became concerned about implications, including around "hallucinations", "poisoning", etc., and decided to put this sort of research on something of a backburner - justified, in part, by a lack of some obvious easy interfacing of this with search (scaling, hallucinations, etc.). Of course, the 'wonderful' thing about humans / "independent agents with survival drives in competitive game-theoretic type scenarios" is: if enough people / "agents" have access / opportunity, someone WILL "push the button". It's just delicious ... the same kinds of patterns over and over - "oh, we should really do something about X / nobody should have power like X, ... but, there's no stopping it, ... oh well". And, the "rules" really are subtly, many levels down, in place, to make it apparently impossible to not get trapped, one way or another. (Anyway... [Cartman voice] Screw you guys, I'm going to my other planet...)
- npsomaratna 3y agoThis is unsubstantiated. The only folks who know exactly how GPT-4 works are employed at OpenAI. The rest of us can only guess.
- YetAnotherNick 3y agoEven if I just go with Sam Altman's public comment, I would have came to similar conclusion: GPT-4 is big and it is hard to make it is faster. The secret sauce and moat lies in data though. I have heard rumour that they have paid competitive coders to write and annotate code with information like complexity for them.
- astrange 3y agoGPT4 can diagram sentences using link grammar parsing (https://www.link.cs.cmu.edu/link/ https://www.link.cs.cmu.edu/link/) which is obscure enough I really don't think they've generated data for it. So it can get pretty good without that.
- YetAnotherNick 3y agoIt's obvious they use data from github and other places. I am talking about extra 0.00..1% very high quality data they (likely)created.
- PostOnce 3y ago"Open" AI, a charity to benefit us all by pushing and publishing the frontier of scientific knowledge. Nevermind, fuckers, actually it's just to take your jobs and make a few VCs richer. We'll keep the science a secret and try to pressure the government into making it illegal for you to compete with us. https://github.com/ggerganov/llama.cpp https://github.com/ggerganov/llama.cpp https://github.com/openlm-research/open_llama https://github.com/openlm-research/open_llama https://huggingface.co/TheBloke/open-llama-7b-open-instruct-GGML https://huggingface.co/TheBloke/open-llama-7b-open-instruct-... https://huggingface.co/TheBloke/open-llama-13b-open-instruct-GGML https://huggingface.co/TheBloke/open-llama-13b-open-instruct... You can use the above without paying OpenAI. You don't even need a GPU. There are no license issues like with the facebook llama.
- BoorishBears 3y agoWhy the vitriol towards OpenAI? If Elon hadn't pulled the rug out from under them after they refused his forceful takeover*, they wouldn't have had to go to Microsoft and they'd still be open. * a takeover which he predicated on the claim that OpenAI was "doomed to fail"
- PostOnce 3y agoIt's several reasons, first it's the lies and the abuse of a charity, that wouldn't be an issue if they had began as a private company instead of robbing a charity. But secondly, even if they were a private company, it's dishonest and reprehensible to claim to congress that you want to "protect the public" when you really only give a shit about protecting your moat, I'm not happy about that either. I'm also tired of tech oligarchs general tomfuckery in all our daily lives, as I suspect many more people here are. OpenAI is just particularly egregious about it. I also think it's my civic duty to let other developers know that OpenAI does not have, by any stretch of the imagination, a stranglehold on this technology or any secret sauce. That's why they're lying and sweating in front of congress.
- victor9000 3y agoYeah, they went from v1 to regulatory capture in the span of months
- RC_ITR 3y agoFor all the 'I know every number' certainty of this post, there's some weird stuff: >(Today, the pre-training could be done with ~8,192 H100 in ~55 days for $21.5 million at $2 per H100 hour.) Why flex both system size and training time to arbitrary numbers? >For example, MoE is incredibly difficult to deal with on inference because not every part of the model is utilized on every token generation. This means parts may sit dormant when other parts are being used. When serving users, this really hurts utilization rates. Utilization of what? Memory? If you're that worried about inference utilization, then why not just fire up a non-MOE model? Here's what the post said about MQA: >Because of that only 1 head is needed and memory capacity can be significantly reduced for the KV cache This is close but wrong. You only need one Key and Value (KV) head, but you still have the same amount of query heads. My guess is that this is all a relatively knowledgeable person, using formulas laid out by the 2020 scaling paper and making a fantasy system (with the correct math), based on that. Put differently, I could probably fake my way through a similar post and be an equal level of close but definitely wrong because I'm way out of my league. That vibe makes me very suspicious.
- deleted 3y ago[deleted]
- moconnor 3y agoNo, the post is correct about MQA. A KV-cache only caches the key and value heads. The point of MQA is that your KV-cache is 1/heads smaller than usual because of this sharing. Having multiple query heads does not affect the cache size, which is the limiting factor in MHA decoding for both memory capacity and bandwidth reasons.
- RC_ITR 3y ago>Autoregressive decoder inference is a severe bottleneck for Transformer models due to the memory bandwidth overhead from loading decoder weights and all attention keys and values at every decoding step (Shazeer, 2019; Pope et al., 2022; de Jong et al., 2022). The memory bandwidth from loading keys and values can be sharply reduced through multi-query attention (Shazeer, 2019), which uses multiple query heads but single key and value heads. Emphasis mine, source here [0] [0] https://arxiv.org/pdf/2305.13245.pdf https://arxiv.org/pdf/2305.13245.pdf FWIW the original MQA paper is called One Write head is all you need. Here's the quote from that referencing multiple heads [1] >We propose a variant called multi-query attention, where the keys and values are shared across all of the different attention "heads", greatly reducing the size of these tensors and hence the memory bandwidth requirements of incremental decoding. We verify experimentally that the resulting models can indeed be much faster to decode, and incur only minor quality degradation from the baseline. [1]https://arxiv.org/pdf/1911.02150.pdf https://arxiv.org/pdf/1911.02150.pdf
- potatoman22 3y agoI wonder what the legal implications of them using SciHub and Libgen would be if that's true. I'd imagine OpenAI is big enough to make deals with publishers.
- twayt 3y agoLibgen / Scihub or not, if the model can provide details about the book other than just high level info like the summary and no explicit deal with the publisher has been made, you can make a strong argument that it is plagiarism. Even if bits and pieces of the book text are distributed across the internet and you end up picking up portions of the book, you still read the book. It is extremely sad but ChatGPT will be taken down by the end of this year and replaced by a highly neutered model next year.
- why_only_15 3y agoIf I read a book and then write a summary, is that plagiarism? What's the difference? I am legitimately not familiar with copyright law, but real lawyers seem to think it is unclear whether training on copyrighted data is illegal (in Japan it's definitely not).
- capableweb 3y agoI'm not a lawyer and obviously we won't get any definite answer unless it actually goes to court, all of this is just hand waving and guessing. But I think that unless GPT starts reciting large parts outside of the context of learning/education/research, reciting smaller snippets would fall into "fair use" and not be illegal.
- twayt 3y agoIf you recite enough small snippets, you make a large one. Especially with ChatGPT you can probe the model by asking certain questions about the material at hand to see if it has seen the entire book. Also you don’t have to be able to recite the book verbatim for it to have been in your training set. The snippets I am referring to are on the side of the training data
- abrax3141 3y agoIf it was trained on CS textbooks, they weren't very good ones. I asked it (GPT4) to write a quantum computer algorithm to square a number. It very confidently told me that to simplify the problem it would use two bits. Okay, fine. But then the algorithm it (again confidently) implemented did a left shift (which it reminded me was multiplying by 2, so it definitely intended this!) and then add the number to itself. It then wrote that in terms of QC gates. Tada! It took me a half beat to realize that rather than this being some new version of squaring a number that I somehow wasn't aware of, it's completely wrong. It only works on 00! Confronted, of course it did the usual "So sorry... I guess I don't know how to do this." dance. I don't get why anyone thinks that this thing is worth anything at all, except for cheating on creative writing tests.
- wastewastewaste 3y agoDamn so you tried once to use it for a thing and it failed? That's crazy, truly a mystery then why so many devs continue to use it daily.
- Tostino 3y agoIt sounds like you don't know how to use it effectively is all I can see from your post.
- mindwok 3y agoIt's arguably the first useful general purpose AI. Claiming it is not worth anything at all because it can't solve a problem that 99.999% of humans would not be able to solve is a pretty ridiculous definition of 'worth'.
- tillinghast 3y ago[flagged]
- dang 3y agoCould you please stop posting unsubstantive comments and flamebait? You've unfortunately been doing it repeatedly. It's not what this site is for, and destroys what it is for. If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.
- getmeinrn 3y ago>If their cost in the cloud was about $1 per A100 hour, the training costs for this run alone would be about $63 million. If someone legitimate put together a crowd funding effort, I would donate a non-insignificant amount to train an open model. Has it been tried before?
- asynchronous 3y agoNot yet, heard tale of several people having the same idea to train an open model though through either crowdfunding or some wizardry with crowdsourcing GPUs. $65 million sounds pretty high though.
- fredoliveira 3y agoConsidering an effort to buy a copy of the constitution raised almost $47M, I wouldn't be so sure. [^1] Worth noting, though, that it isn't just the computing budget that's missing here - it is also (and perhaps even more importantly) the high quality data to actually train the model. [^1]: https://en.wikipedia.org/wiki/ConstitutionDAO https://en.wikipedia.org/wiki/ConstitutionDAO
- holoduke 3y agoSome kind of SETI project, but for training a high number parameter llm would be awesome.
- drexlspivey 3y agoHow many people have A100s at home?
- TheRealPomax 3y agoThe comparison to SETI@home[1] is that you just need the same mount of total processing power, not the same supercomputer setup. People don't need to own A100s, they just need to be willing to be part of a distributed supercomputer by running a background app that downloads chunks of data, processes them, and sends the result back. The utility comes from having enough people participate (which worked quite well for SETI@home, but helping find "signals from outer space" is a little bit more interesting than "helping train an LLM") [1] https://setiathome.berkeley.edu/ https://setiathome.berkeley.edu/
- esaym 3y agoEveryone hates on crypto because of all the electricity use for mining. But how much electricity is the training of all the giant LLMs costing us?
- flangola7 3y agoLLMs actually have tangible use (granted far from bespoke yet). Crypto mining has no tangible benefit.
- Roark66 3y agoThe tweet is gone. What was in it? Also, I'm dubious about this unsubstantiated claim. The biggest past innovation (training with human feedback) actually shrunk the size of a model. Compare Bloom-366B with falcon-40B (much better). I would be mildly surprised if it turned out Gpt4 has 1.8T parameters. (even if it's a composite model as they say) The article says they use 16 experts 111B each. So the best thing to assume is probably that each of these experts is basically a fine tuned version of the same initial model for some problem domain.
- Al0neStar 3y agoMaybe 111B is the base GPT-3.5 model.
- why_only_15 3y agoAs a note the 366B in Bloom-366B refers to the number of tokens, not the number of parameters. Bloom had 176B parameters (still many more than Falcon)
- langsoul-com 3y agoWe should default to using the thread aggregators instead of using twitter links. My God Twitter threads are unreadable.
- Ozzie_osman 3y agoThere's a section at the end where there is speculation on what the entire dataset entails. My guess is a chunk of it is probably from ChatGPT data (or GPT3 data from when training on your requests was opt-out rather than opt-in).
- mmahemoff 3y agoI've been wondering how freemium services like Thread Reader still operate now that Twitter is charging prohibitive prices for API access and taking measures to prevent scraping. The cheapest API plan with read access is $100/month, which reads 10,000 tweets, so could only produce about 500 pages like this one on demand.
- errantmind 3y agoThere was a post on HN recently with a workaround these apps are using. I don't have it handy but I'm sure you can find it if you look.
- n1c 3y agoThere's probably some interesting bits of info in yesterday's Nitter thread: https://news.ycombinator.com/item?id=36665406 https://news.ycombinator.com/item?id=36665406
- xeckr 3y agoconst puppeteer = require('puppeteer'); and so on and so forth.
- TeMPOraL 3y ago> The conspiracy theory that the new GPT-4 quality had been deteriorated might be simply because they are letting the oracle model accept lower probability sequences from the speculative decoding model. In other words: the speculation was likely right, I'll propose a specific mechanism explaining it, but then still insult the people bringing it up and keep gaslighting them.
- mitchdoogle 3y agoCalling something a conspiracy theory is not an insult against anybody. It's a theory because it's unproven and it's a conspiracy because people think OpenAI purposely degraded their own service, hence conspiracy theory.
- TeMPOraL 3y agoThat's a motte-and-bailey defense. Yes, what you say is technically correct with respect to meaning of "conspiracy" and "theory" as individual words. But it's also completely false with respect to what "conspiracy theory" means in actual use - which is to group the subjects (here: people believing GPT-4 quality has been degrading over time, in spite of OpenAI strongly implying otherwise) in the same bucket as flat earthers, vaccine denialists, UFO believers, NWO fearmongers, etc. Calling the belief "that the new GPT-4 quality had been deteriorated" a "conspiracy theory" goes beyond claiming the belief itself is wrong - it's also claiming that holding this belief implies significantly compromised reasoning skills. That is, it's just a drive-by insult.
- qwertox 3y ago> This, of course, is “only” a batch size of 7.5 million tokens per expert due to not every expert seeing all tokens. > Mixture of Expert Tradeoffs: There are multiple MoE tradeoffs taken: For example, MoE is incredibly difficult to deal with on inference because not every part of the model is utilized on every token generation. Are these experts able to communicate among them in one query? How do they get selected? How do they know who to pass information to? Would I be able to influence the selection of experts by how I create my questions? For example to ensure that a question about code gets passed directly to an expert in code? I feel silly asking this question, but I honestly have no idea how to interpret this.
- l33tman 3y agoYou shouldn't take the "mixture of experts" too literally, it's yet another architecture to use internally for a gradient descent optimized graph of ops. I obviously don't know how GPT-4 do it (or if it even does it) but think of partitioning your network into a couple of very isolated sub-graphs (the "experts"), and add another learnable network between the input tokens and the experts, that learns to route tokens to 1 or more expert sub-graphs. Then the gain is that you can potentially ignore running the unused sub-graphs completely for that token, and you can distribute them on other GPUs as except for the input and output they are independent of each other. It all depends on the problem, data, and if the gradient descent optimizer can find a way to actually partition the problem usefully using the router and "experts".
- jonplackett 3y agoBard taking notes…
- rjb7731 3y agoI've previously noticed when playing with GPT-4 it can sometimes 'autocomplete' on different sections of the text its feeding back, sometimes what looks like 4 or more different sections. Might be unrelated but is this MoE in action or them streaming the response in some way?
- LiamPowell 3y agoThis is just an issue with their frontend that seems to occur when it encounters \n\n. The actual data coming in only changes at the end of the message.
- wokwokwok 3y agoThis a duplicate post of pure speculation.
- aussieguy1234 3y agoThe fact they are using MoE is interesting. There are alot of specialised open source models on HuggingFace. You just need an LLM to act as the core "brain" and a few other components. HuggingGPT works similar to this. It automatically chooses, downloads and runs the right "expert" model from HuggingFace https://arxiv.org/abs/2303.17580 https://arxiv.org/abs/2303.17580
- xeckr 3y agoIf this is true, then: 1. Training took 21 yottaflops. When was the last time you saw the yotta- prefix for anything? 2. The training cost of GPT-4 is now only 1/3 of what it was about a year ago. It is absolutely staggering how quickly the price of training an LLM is dropping, which is great news for open source. The google memo was right about the lack of a moat.
- theLiminator 3y agoThe real moat is an abundance of high quality data.
- classified 3y ago... stolen without regard for copyright and licensing.
- pas 3y agoFair use!? /s
- quickthrower2 3y agoIMO the real moat right now is expertise / smart teams and cash.
- hospitalJail 3y agoThe infrastructure/training libraries already exists. I'm sure you can get people who worked at scale that can figure out how to glue things together. Reddit, twitter, etc.. raising prices is going to make it more expensive.
- quickthrower2 3y agoIf you are right then it just becomes who wants to throw the most cash in like a giant game of poker but where you don’t know the pot odds.
- PUSH_AX 3y agoWhat is this hyper dramatic nonsense tweet about, “It’s over“? What’s over?
- sweezyjeezy 3y agoThe wait to find out what the model is I'm guessing?
- toxicFork 3y agoThe thing, dude, the thing, is over!
- astrange 3y agoIt's a meme based on quoting this tweet. https://twitter.com/jebbush/status/929541504187686912 https://twitter.com/jebbush/status/929541504187686912
- nightsd01 3y ago“Leaked” seems like a strong clickbait claim from whoever wrote this, along with the “it’s over” part….
- why_only_15 3y agoLeaked is I think an accurate term -- this (or the original post) is fairly new information leaked from openai.
- qaq 3y agoHmm “Sam Altman won't tell you that GPT-4 has 220B parameters and is 16-way mixture model with 8 sets of weights” George Hotz said this in his recent interview with Lex Fridman. It looked like Lex knew this to be true by the way he reacted.
- imtemplain 3y ago[dead]
- dmarchand90 3y agoCan anyone provide an alternative link to https://twitter.com/i/web/status/1678545170508267522 https://twitter.com/i/web/status/1678545170508267522 I haven't registered for Twitter since it started and I'd rather not now (though I probably will if it's the only way to get leaked gpt4 training details)
- _a9 3y agoWayback failed to load the subtweets but archive.is has a copy but it seems to stop after around 10 subtweets. The threader link that was posted has it all though. https://archive.is/Y72Gu https://archive.is/Y72Gu
- elzbardico 3y agoIt is a bit problematic if it is being trained on copyrighted textbooks without compensation for the authors. Even for open-source science, I think it is a bit unethical if OpenAI is using public founded research without attribution or compensation. Tax Payers paid for those NIH grants, you know...
- freedmand 3y agoWait til you see all the copyrighted and pirated data most large language models are trained on: - https://www.washingtonpost.com/technology/interactive/2023/ai-chatbot-learning/ https://www.washingtonpost.com/technology/interactive/2023/a... - https://pile.eleuther.ai/ https://pile.eleuther.ai/ (data hosted by https://the-eye.eu/ https://the-eye.eu/, where it's not too hard to find pirated, copyrighted books, e.g. https://the-eye.eu/public/Books/cdn.preterhuman.net/texts/literature/ https://the-eye.eu/public/Books/cdn.preterhuman.net/texts/li...)
- rurp 3y ago> The conspiracy theory that the new GPT-4 quality had been deteriorated might be simply because they are letting the oracle model accept lower probability sequences from the speculative decoding model. Whether or not this specific theory is true something along these lines seems like the most likely explanation for the quality degradation that many have noticed; where OpenAI's claims about not changing the model are both technically true and conpletely misleading.
- gdubs 3y agoRecently I was saying how much amazing stuff there is in retro computing. One thing that keeps coming to mind for me recently is just how visionary Thinking Machines Connection Machine supercomputer architecture was with its massive parallelism built in, with neural network applications being a key predicted use case at the time. That was so long ago! Interesting to think about in comparison to the challenges today around parallelizing 'commodity' GPUs. Scare quotes because he A100 and H100 are pretty impressive machines in and of themselves.
- StackOverlord 3y agoRelated: https://longnow.org/essays/richard-feynman-connection-machine/ https://longnow.org/essays/richard-feynman-connection-machin...
- henkdehenker 3y agoSo George Hotz was right
- dataangel 3y ago> Part of this extremely low utilization is due to an absurd number of failures requiring checkpoints that needed to be restarted from. Hahahaha, the truth of anyone who has worked with quanty types running Python code at scale on a cluster
- TheRealPomax 3y ago"*The post about GPT-4's architecture had been removed due to a copyright claim.", https://twitter.com/Yampeleg/status/1678582275561103360 https://twitter.com/Yampeleg/status/1678582275561103360
- kristianp 3y agoI wonder if any open source MOE models are being worked on. Could I run an 8x13B model on my 16GB graphics card, only loading the expert that is needed per run?