21 ms·
Cerebras-GPT: A Family of Open, Compute-Efficient, Large Language Models
- tombert 4y agoHas anyone tried this? I have 96GB of GPU memory; will that be enough to run the biggest model?
- cuuupid 4y ago13B fits nicely even in a 3090 (24gb vram)!
- spi 4y agoI have not tried, but 96GB of GPU memory is plenty, for inference there should certainly be no issue. Their biggest model has 13B parameters, you should be able to run inference (float16) already with 32GB of memory. With 96GB of memory you should also be able to fine-tune it (possibly some tricks like gradient accumulation and/or checkpointing might be needed), but you have to be ready for many days of computation...
- alchemist1e9 4y ago> but you have to be ready for many days of computation... I was thinking since we have API prices in tokens and now it looks like self hosted inference on high end GPUs for similar models. Then based on electricity prices there will be a self-hosted prices in tokens. Then how close are these already? What is the markup today from roughly the raw electricity cost that OpenAI has.
- eldenring 4y ago> Trained using the Chinchilla formula, these models provide the highest accuracy for a given compute budget. I'm confused as to why 111 million parameter models are trained with the Chinchilla formula. Why not scale up the training data? If you're training smaller models, surely optimizing performance is better than optimizing total compute. Seems like a silly misunderstanding of the Chinchilla paper, but I'm sure I'm missing something
- gamegoblin 4y agoTrue. There was a good blog post published about this a few weeks ago: https://finbarr.ca/llms-not-trained-enough/ https://finbarr.ca/llms-not-trained-enough/ Money quote for those who don't want to read the whole thing: ''' When people talk about training a Chinchilla-optimal model, this is what they mean: training a model that matches their estimates for optimality. They estimated the optimal model size for a given compute budget, and the optimal number of training tokens for a given compute budget. However, when we talk about “optimal” here, what is meant is “what is the cheapest way to obtain a given loss level, in FLOPS.” In practice though, we don’t care about the answer! This is exactly the answer you care about if you’re a researcher at DeepMind/FAIR/AWS who is training a model with the goal of reaching the new SOTA so you can publish a paper and get promoted. If you’re training a model with the goal of actually deploying it, the training cost is going to be dominated by the inference cost. This has two implications: 1) there is a strong incentive to train smaller models which fit on single GPUs 2) we’re fine trading off training time efficiency for inference time efficiency (probably to a ridiculous extent). Chinchilla implicitly assumes that the majority of the total cost of ownership (TCO) for a LLM is the training cost. In practice, this is only the case if you’re a researcher at a research lab who doesn’t support products (e.g. FAIR/Google Brain/DeepMind/MSR). For almost everyone else, the amount of resources spent on inference will dwarf the amount of resources spent during training. '''
- haldujai 4y agoWhile true I think this also misses that “for almost everyone else” you’re probably not (or at least should not) be trying to optimize zero-shot performance if you have an intended high inference use case so I don’t think Chinchilla would be all that relevant.
- whalesalad 4y agoThis “AI spring” is really snowballing with the crazy nouns and terminology. Alpaca, llama and now chinchilla??
- murkt 4y agoChinchilla actually came before alpaca and llama. Every new variation of model gets some new name, just like every library gets a new name. There were all kinds of BERTs before - DistilBert, Roberta, SciBERT, Schmobert, Schmuber, etc. Many hundreds of them, I think.
- ramesh1994 4y agoThe term "chinchilla" predates llama/alpaca. It doesn't directly map to a specific model, rather a family of compute-optimal models.
- whoisnnamdi 4y agoChinchilla actually came first!
- jhbadger 4y agoAs mentioned, chinchilla is not part of this trend, and chinchillas are rodents. Alpacas and llamas are South American camelids (animals related to camels). So if additional names are needed, I would expect them to be vicuña and guanaco, as they are also in the group.
- mometsi 4y agoI think the relevant category is "Adorable Fuzzy Critters of the Andes". See also https://en.wikipedia.org/wiki/Spectacled_bear https://en.wikipedia.org/wiki/Spectacled_bear
- johnchristopher 4y agoOT: I don't know about their scaling strategy for LLM but their scaling strategy for displaying pictures is disappointing. (it's all blurry)
- ricopags 4y agoCame here to point this out, though not as pithily :D Really, really bad mark on whoever is in charge of their web marketing. Images should never look that bad, not even in support, but definitely not in marketing. edit: so this post is more useful, 4k res using Edge browser
- Kelamir 4y agoLast time I viewed it, I believe it wasn't blurry. Perhaps to scale the traffic the images are now displayed in lower quality? But I'm not sure anymore that it wasn't initially blurry... Perhaps I'm hallucinating, like large language models. Current image displayed is https://www.cerebras.net/wp-content/uploads/2023/03/Scaling-laws-blog-comparison-uai-258x119.png https://www.cerebras.net/wp-content/uploads/2023/03/Scaling-... , will see if it changes.
- Kelamir 4y agoI can confirm, it does change. As of now, it displays one of higher quality: https://www.cerebras.net/wp-content/uploads/2023/03/Scaling-laws-blog-comparison.png https://www.cerebras.net/wp-content/uploads/2023/03/Scaling-...
- thewataccount 4y agoThey're dynamically scaled and something must be broken. If you inspect source you can find the raw images, here's a few: https://www.cerebras.net/wp-content/uploads/2023/03/Downstream-tasks-figure.png https://www.cerebras.net/wp-content/uploads/2023/03/Downstre... https://www.cerebras.net/wp-content/uploads/2023/03/Scaling-laws-blog-comparison.png https://www.cerebras.net/wp-content/uploads/2023/03/Scaling-... https://www.cerebras.net/wp-content/uploads/2023/03/Scaling-Laws-blog-fig-2.png https://www.cerebras.net/wp-content/uploads/2023/03/Scaling-... EDIT: Looks like it scores better with less training - up until it matches GPT-J/Pythia/OPT and doesn't appear to have much benefit. It maybe scores slightly better then GPT-J which is pretty "eh", I'm not sure if GPT-J level performance is really useful for anything? NeoX 20B outperforms it in everything if you don't care about the amount of training needed. Does the better performance for less training matter if that benefit only applies when it's only performing a lot worse then GPT-J? It appears to lose it's scaling benefits before the performance is interesting enough to matter?
- mometsi 4y agoSummary: This is a company that makes AI accelerator ICs. They reimplemented Chinchilla and released the model weights under a permissive license.
- bogwog 4y agoIn other words, they’re actually incentivized to help make LLMs as accessible as possible, rather than try to keep them locked up to hide them from competitors. Which makes me wonder if Nvidia is doing anything with LLMs too?
- vintermann 4y agoNVidia has certainly pushing the envelope on image generation. StyleGAN3 was really cool when it came. But it is an issue that their chips are hardly optimized for LLMs.
- meghan_rain 4y agoHow can a GPU be optimized for StyleGAN but not LLMs? Serious question.
- MacsHeadroom 4y agoRAM. GPT-3 is over 600GB, ie just the max RAM of 8xA100s, because that's all the hardware can fit. StableDiffusion plus a whole chain of imagenets can make any visual imagery imaginable in 2GB of RAM. Meanwhile 2GB of RAM barely runs a basic tiny text completion NN that can't do anything intelligent. Text requires a lot more parameters (and more memory/RAM) than images.
- brucethemoose2 4y agoThe Cerebras node's actual "RAM" (the 40GB of SRAM) is pretty modest too, but being an enormous chip with the networked storage pools is certainly a better situation than a bunch of A100s reaching out to every other A100. Honestly, all the AI ASIC makers drastically underestimated the RAM requirements of future models. Graphcore's 4GB and Tenstorrent's 8GB per IC is kinda laughable, and it takes them longer to adjust than Nvidia. And Cerebras' original pitch was "fit the entire model into SRAM!"
- JamesCoyne 4y agoSlightly off-topic: I remember seeing news about the enormous chip Cerebras was/is selling (pdf https://f.hubspotusercontent30.net/hubfs/8968533/WSE-2%20Datasheet.pdf https://f.hubspotusercontent30.net/hubfs/8968533/WSE-2%20Dat...). Has there been any indication that the LLMs released in the last few months use exotic hardware like this, or is it all "standard" hardware?
- wmf 4y agoOpenAI uses Nvidia GPUs and Google uses their TPUs.
- brucethemoose2 4y agoYou might see more training on Intel XPUs when they come out, since they have such enormous RAM pools. Maybe AMD MI300s and Intel Ponte Vecchio (both 128GB) in the shorter term, though I think they will mostly be in HPC supercomputers instead of cloud instances.
- ipsum2 4y agoEveryone except Google uses Nvidia for training. Cerebras, Gaudi, and other custom AI accelerators have unable to surpass Nvidia in performance/$ and performance/watt yet.
- chessgecko 4y agoI wonder what led to such a gap between llama 7b and Cerebras 13b. I hope they discuss it in the paper.
- gpm 4y agoIs there a benchmark comparing the two that I missed? Edit: The huggingface page has 0-shot benchmarks which you can compare against the llama paper https://huggingface.co/cerebras/Cerebras-GPT-13B https://huggingface.co/cerebras/Cerebras-GPT-13B https://arxiv.org/pdf/2302.13971.pdf https://arxiv.org/pdf/2302.13971.pdf
- freeqaz 4y agoI'm on mobile and struggled to compare these two tables properly. Would you mind posting a summary of your findings? Here are some values but I don't know what they mean. LLama 60B on the left, Cerebras 13B on the right. PiQA: 82.8 / 76.6 WinoGrade: 77.0 / 64.6 ARC-e: 78.9 / 71.4
- gpm 4y agoReally short summary: LLaMa is better, even smaller LLaMa models. Table format: Benchmark, Cerebras 13B, LLama 7B, LLama 13B, LLama 60B HellaSwag, 51.3, 76.1, 79.2, 84.2 Piqa, 76.6, 79.8, 80.1, 82.8 Wino-Grande, 64.6, 70.1, 73.0, 77.0 Arc-e, 71.4, 72.8, 74.8, 78.9 Arc-c, 36.7, 47.6, 52.7, 56.0 OpenBookQA, 28.6, 57.2, 56.4, 60.2
- ftxbro 4y agoThis gap makes sense to me. The academic point of the Cerebras paper is to show their nice empirical scaling law for compute-optimal training, whereas the academic point of the LLaMA paper was to show that you can make small models punch above their weight by training them in a way that is deliberately not compute-optimal. Of course both of those publications had other academic and marketing purposes. From the Cerebras blog post: "Trained using the Chinchilla formula, these models provide the highest accuracy for a given compute budget." From the LLaMA paper: "The focus of this work is to train a series of language models that achieve the best possible performance at various inference budgets, by training on more tokens than what is typically used."
- brucethemoose2 4y agoFYI: Cerebras's nodes are very different than your typical Nvidia training nodes: https://www.anandtech.com/show/16626/cerebras-unveils-wafer-scale-engine-two-wse2-26-trillion-transistors-100-yield https://www.anandtech.com/show/16626/cerebras-unveils-wafer-... Each individual "chip" has 40GB of SRAM vs ~76MB for the Nvidia H100, and networked pools of external RAM, SSDs and such. Thats why the training architecture is so different.
- arbuge 4y agohttps://www.cerebras.net/product-chip/ https://www.cerebras.net/product-chip/ There's a comparison picture there of one of their chips alongside a regular GPU chip. Effectively they use up the entire wafer.
- brucethemoose2 4y agoYeah, and that doesn't even do the nutty IO on these things justice. A 16x CS2 cluster like they describe is like a huge Nvidia cluster in terms of throughput, but more like a single Nvidia node structurally.
- alchemist1e9 4y agoIt’s unbelievable stuff. Does anyone know how much a single box costs? They are selling them it looks like.
- sbierwagen 4y agoCS-1 costs "$2-3 million", CS-2 costs "several" million. A single Nvidia H100 costs somewhere around $30,000 each, so a GPU server with every slot populated costs about $300,000.
- brucethemoose2 4y agoServeTheHome claims "HGX A100 platforms, when they are sold as single servers are generally in the $130K-$180K even leaving a very healthy margin for OEMs/ resellers" https://www.servethehome.com/graphcore-celebrates-a-stunning-loss-at-mlperf-training-v1-0/2/ https://www.servethehome.com/graphcore-celebrates-a-stunning... Not sure about the H100, but it seems to be more supply constrained (hence pricier) atm. Now, the real question is how many HGX nodes "equals" a single CS2 node. The math here is extremely fuzzy, as the benefit to such extreme node consolidation depends on the workload, and the CS-2 takes up less space, but the HGX cluster will have more directly accessible RAM and better turnkey support for stuff since its Nvidia.
- simonw 4y ago"Cerebras open sources seven GPT-3 models from 111 million to 13 billion parameters." I don't understand why they describe them as GPT-3 models here as opposed to calling them GPT models. Or even LLMs - but I guess that acronym isn't as widely recognized.
- wsgeorge 4y agoI think GPT-3 is used as a benchmark for performance, so saying a model is on par with GPT-3 should give you an idea of what you can get out of it. IIRC most open source models to date - including the semi-open LLaMAs - have GPT-3-like performance. Nothing gets close to GPT-3.5 and beyond.
- simonw 4y agoYou can try out some of these models on Hugging face here: https://huggingface.co/cerebras/Cerebras-GPT-1.3B https://huggingface.co/cerebras/Cerebras-GPT-1.3B That was the largest that had inference enabled - I'd really like to try this one: https://huggingface.co/cerebras/Cerebras-GPT-13B https://huggingface.co/cerebras/Cerebras-GPT-13B
- rnosov 4y agoI might be missing something but it looks to me that actually running this "open" model requires special hardware only accessible with a cloud subscription with 60 000 USD / week minimum spend[1]. Can anyone confirm if you can run it on your own hardware? If software is open but hardware is locked I don't see the point. [1] https://www.hpcwire.com/2021/09/16/cerebras-wafer-scale-engine-ai-system-is-now-available-in-the-cloud/#:~:text=With%20a%20weekly%20minimum%20buy,the%20entire%20CS%2D2%20system https://www.hpcwire.com/2021/09/16/cerebras-wafer-scale-engi.... EDIT: Ok, looks like I've missed the hugging face repo. The language they use is a bit confusing.
- bubblethink 4y agoYou can run inference on GPUs. These are just models and weights.
- simonw 4y agoThe PyTorch model files are already available to download from Hugging Face - the largest one looks to be 52GB. They should run on any hardware that can run regular PyTorch models.
- 2bitencryption 4y agoThis type of article (or press release, or whatever you want to call it) is exactly what makes the future so interesting. The cat is out of the bag, the genie is out of the bottle, the confetti has left the cannon[0]. It's tempting to see a world dominated by Google Bard, ChatGPT, Bing Search, etc. And no doubt, they will be huge players, with services that are far more powerful than anything that can be run on the edge. But. BUT. The things that we can do on the edge are incredible now. Just imagine a year from now, or two. These earth-shattering models, which seem to be upending a whole industry, will soon have equivalents that run on the edge. Without services spying on your data. Without censorship on what the model can/cannot say. Because it's all local. When was the last time this happened? There will be players who publish weights for models that are free to use. The moment that torrent magnet link is published, it's out in the wild. And smart people will package them as "one click installers" for people who aren't tech-savvy. This is already happening. So every time you're amazed by something chat-gpt4 says, remember that soon this will be in your pocket. [0] the "confetti" idiom brought to you by chat-gpt4.
- jazzkingrt 4y agoSerious question: is it typical to describe client-side computing as "on the edge"? I thought running something on the edge referred to running it in close network proximity to the user, rather than users having control and running things themselves.
- capableweb 4y agoYes, "edge computing" can refer to both computing done as close to the user as possible geographically, or even on the device itself. If someone says "I wanna do edge computing" it's not clear enough to know if they just want to have servers they control as close to the user as possible, or do the computing on the device itself. I think Apple would say "edge computing" is on the actual device while CloudFlare would say "edge computing" is on their infrastructure, but distributed to be physically closer to the end user.
- iamerroragent 4y ago
- simonw 4y agoDoes the chinchilla recipe still hold today? I got the impression that the LLaMA paper proposed a different result where throwing far more tokens at the problem had a very meaningful impact, or did I misunderstand that?
- evanmays 4y agoThere’s discussion elsewhere in this thread what chinchilla actually means. I’ll only compare it to llama. Tldr; Chinchilla isn’t wrong, it’s just useful for a different goal than the llama paper. There’s 3 hyper parameters to tweak here. Model size (parameter count), number of tokens pre trained on, and amount of compute available. End performance is in theory a function of these three hyperparameters. You can think of this as an optimization function. Chinchilla says, if you have a fixed amount of compute, here’s what size and number of tokens to train for maximum performance. A lot of times, we have a fixed model size though though, because size impact inference costs and latency. Llama operates in this territory. They choose to fix the model size instead of the amount of compute. This could explain gaps in performance between Cerebras models of size X and llama models of size X. Llama models of size X have way more compute behind them
- espadrine 4y agoI don’t think it holds for two reasons. First, it only holds for a given architecture and implementation. Obviously, a different architecture will have a different training slope. This is clear when comparing LSTM with Transformers, but is also true between transformers that use prenorm/SwiGLU/rotary-positional, and those that follow Vaswani 2017. In terms of implementation, some algorithms yield the same result with fewer operations (IO, like FlashAttention and other custom CUDA kernels, and parallelism, like PaLM, which both came after Chinchilla), which unambiguously affect the Tflops side of the Chinchilla equation. Also, faster algorithms and better parallelization will yield a given loss sooner, while less power-hunger setups will do that cheaper. Second, even in the original Chinchilla paper in figure 2, some lines are stopped early before reaching Pareto (likely because it ran out of tokens, but LLaMA makes it seem that >1 epoch training is fine).
- binarymax 4y agoHere are the zero-shot accuracy numbers posted in the Huggingface evaluations for Cerebras-GPT 13B vs. the results of LLaMa 13B in their paper: Model BoolQ PIQA SIQA HellaSwag WinoGrande ARC-e ARC-c OBQA LLaMa 13B 78.1 80.1 50.4 79.2 73 74.8 52.7 56.4 Cerebras-GPT 13B - 76.6 - 51.3 64.6 71.4 36.7 28.6
- wsgeorge 4y agoI guess it's something. It still goes to show how far open models are behind the proprietary SOTA.
- binarymax 4y agoIndeed but this is zero-shot performance. Fine-tuning for a task should get you pretty good results. I'm interested in seeing the results of an Alpaca method against this Cerebras 13B model.
- MacsHeadroom 4y ago>I'm interested in seeing the results of an Alpaca method You're talking apples to oranges. The "Alpaca method" is a dataset generation method. Nothing about Alpaca's training method is novel, interesting, or efficient. Alpaca used the same standard training method everyone else uses, A100 clusters. If you mean LoRA/PEFT training which people used to replicate Alpaca then that is also apples to oranges because LoRA/PEFT is a finetuning method not a pre-training method.
- deleted 4y ago[deleted]
- UncleEntity 4y agoOne could take the alpaca dataset and fine tune using the LoRA/PEFT method and compare to the Stanford alpaca fine tuned llama model. Presumably…
- 4y ago
- ftxbro 4y ago> Our paper, which will be available soon, will detail our training methods and performance results. Yay there will be a paper let's gooooooo!
- wg0 4y agoNoob to ML in practice. These models containing weights, all of them, do they have a standard file/binary format?
- examplary_cable 4y ago[I'm not an expert] but I believe .ckpt and .safetensors. The problem with .ckpt is that it executes arbitrary code in your machine(very unsafe). While .safetensors was made by huggingface in order to have a safe format to store the weights. I've also seen people load up the llama 7B via a .bin file.
- ivanvas 4y agoIs it currently possible to find-tune any of the foundation modules available on a few Gb of unsupervised text?
- amilios 4y agoComparing the 13B model here https://huggingface.co/cerebras/Cerebras-GPT-13B https://huggingface.co/cerebras/Cerebras-GPT-13B to LLaMA-13B https://github.com/facebookresearch/llama/blob/main/MODEL_CARD.md https://github.com/facebookresearch/llama/blob/main/MODEL_CA... you can see that in all of the reasoning tasks Cerebras-GPT lags behind. Any reason to use Cerebras instead of LLaMA? Doesn't seem like it.
- mdagostino 4y agoLLaMA is non-commercial
- potatoman22 4y agoCan the LLaMA weights be used for commercial products?
- gpm 4y agoUnclear, likely jurisdiction dependent, almost certainly not if you need to operate world wide.
- espadrine 4y agoThere are two aspects to it. The first one is whether they would actually sue. The optics would be terrible. A similar situation occurred in the 90s when the RC4 cipher’s code was leaked. Everyone used the leaked code pretending that it was a new cipher called arc4random, even though they had confirmation from people that licensed the cipher that its output was identical. Nobody was sued, and the RSA company never acknowledged it. The second one is related to the terms. The LLaMA weights themselves are licensed under terms that exclude commercial use:[0] > You will not […] use […] the Software Products (or any derivative works thereof, works incorporating the Software Products, or any data produced by the Software), […] for […] any commercial or production purposes. But the definition of derivative works is gray. AFAIK, if LLaMA is distilled, there is an unsettled argument to be had that the end result is not a LLaMA derivative, and cannot be considered copyright or license infringement, similar to how models trained on blog articles and tweets are not infringing on those authors’ copyright or licensing. The people that make the new model may be in breach of the license if they agreed to it, but maybe not the people that use that new model. Otherwise, ad absurdum, a model trained on the Internet will have content that was generated by LLaMA in its training set, so all models trained on the Internet after Feb 2023 will break the license. IANAL, but ultimately, Meta wins more by benefiting from what the community contributes on top of their work (similar to what happened with React), than by suing developers that use derivatives of their open models. [0]: https://docs.google.com/forms/d/e/1FAIpQLSfqNECQnMkycAp2jP4Z9TFX0cGR4uf7b_fBxjY_OjhJILlKGA/viewform https://docs.google.com/forms/d/e/1FAIpQLSfqNECQnMkycAp2jP4Z...
- eternalban 4y ago> It takes substantial technical expertise to train very large models on GPUs. In the recently released GPT-4 Technical Report, OpenAI credits over thirty contributors just for compute infrastructure and scaling. This is called a silver lining for some (in case you were worried about gpt taking your job). Privacy requirements alone will in the near term force major companies to run their own inference (if not training). The expertise required are nearly identical to that of running large scale distributed computational graphs. This is an interesting diveragence from what happened with web. The backends started out simple before map-reduce and before deconstructing databases and processing distributed logs. With ML, we'll jump right into the complex backends in tandem with easy-picking early stage edge applications (which we see daily on HN).
- skybrian 4y agoWhat’s in the Pile training data they used? How much source code does it include?
- sanxiyn 4y agohttps://arxiv.org/abs/2101.00027 https://arxiv.org/abs/2101.00027 is the paper and it includes 95.16 GiB from GitHub.
- antimatter15 4y agoLooking at their charts it seems like their 6.7B model is considerably worse than GPT-J which is an existing open 6B model from several years ago. I wish rather than stopping training early they would have run more data through a small model so we could have something more competitive with LLaMA 7B.
- cs-fan-101 4y agoSomeone posted this repost from the Cerebras Discord earlier, but sharing for visibility - "We chose to train these models to 20 tokens per param to fit a scaling law to the Pile data set. These models are optimal for a fixed compute budget, not necessarily "best for use". If you had a fixed parameter budget (e.g., because you wanted to fit models on certain hardware) you would train on more tokens. We do that for our customers that seek that performance and want to get LLaMA-like quality with a commercial license"
- HanClinto 4y agoSounds like we should crowd-fund the cost to train and open source one of these models with LLaMa-like quality. I'd chip in!
- brucethemoose2 4y agoTBH that seems like a good job for Cerebras. There are plenty of such efforts, but the organizer needs some kind of significance to attract a critical mass, and a AI ASIC chip designer seems like a good candidate. Then again, maybe they prefer a bunch of privately trained models over an open one since that sells more ASIC time?
- brucethemoose2 4y ago> Cerebras Discord This is really weird to hear out loud. I still think of Discord as a niche gaming chatroom, even though I know that (for instance) a wafer scale IC design company is hosting a Discord now.
- visarga 4y agoOf course this is great news, I hope these models can be fine-tuned to be like lighter versions of chatGPT. But I remember reading in the LLaMA paper that a small model can still improve when trained more than the Chinchilla budget. > For instance, although Hoffmann et al. (2022) recommends training a 10B model on 200B tokens, we find that the performance of a 7B model continues to improve even after 1T tokens. Cerebras says: > For instance, training a small model with too much data results in diminishing returns and less accuracy gains per FLOP But this is only of concern when you care about the training cost, such as when you are budget limited researcher or a company who doesn't deploy models at scale. But when you care about the total cost of deployment, then making a small model even better with lots of data is a smart move. In the end it matters more to have the most efficient model in prediction, not the most efficient model in training.
- Garcia98 4y agoI've been following open source LLMs for a while and at first glance this doesn't seem too powerful compared to other open models, Flan-Alpaca[0] is licensed under Apache 2.0, and it seems to perform much better. Although I'm not sure about the legalities about that licensing, since it's basically Flan-T5 fine-tuned using the Alpaca dataset (which is under a Non-Commercial license). Nonetheless, it's exciting to see all these open models popping up, and I hope that a LLM equivalent to Stable Diffusion comes sooner than later. [0]: https://github.com/declare-lab/flan-alpaca https://github.com/declare-lab/flan-alpaca
- alchemist1e9 4y agoSounds like you might be the right person to ask the “big” question. For a small organization or individual who is technically competent and wants to try and do self-hosted inference. What open model is showing the most promise and how does it’s results compare to the various openAI GPTs? A simple example problem would be asking for a summary of code. I’ve found openAI’s GPT 3.5 and 4 to give pretty impressive english descriptions of code. Running that locally in batch would retain privacy and even if slow could just be kept running.
- Garcia98 4y agoGoogle's Flan-T5, Flan-UL2 and derivatives, are so far the most promising open (including commercial use) models that I have tried, however they are very "general purpose" and don't perform well in specific tasks like code understanding or generation. You could fine-tune Flan-T5 with a dataset that suits your specific task and get much better results, as shown by Flan-Alpaca. Sadly, there's no open model yet that acts like a Swiss knife and gets good-enough results for multiple use cases.
- capableweb 4y agoIterating on the question, what model/weights would be the most appropriate for the specific use case of code generation right now?
- patientplatypus 4y ago[dead]
- fuzzieozzie 4y agoCerebras has an efficiency advantage at generating LLMs (assuming IP is open). This is going to be fun to be a part of.
- dukeofdoom 4y agoI wonder how decrotive our world will become as a consequence of how cheap it will become to make art using AI. I kind of want 3d marble statues and baroque art of a future reinasance everywhere. But wonder if we will turn minimalistic as a response.
- ericd 4y agoI've been wondering about the best way to print large format versions of custom Renaissance-style paintings with goofy subjects for our walls at home. I guess I have to figure out how to best upscale the output first.
- rbanffy 4y agoA tangential question: I wonder what, as chiplets become increasingly more common, will Cerebras do to keep their technological advantage of wafer-scale integration. What is the bandwidth and latency of the connections between the tiles? Is there such a thing as bandwidth per frontier length?
- mark_l_watson 4y agoEven though I usually use OpenAI's APIs, just because that is the easiest path, I do also use Hugging Face open models (via their APIs, and running locally) and I will check out Cerebras also. Alternatives are good!
- AlexanderTheGr8 4y agoIs there a regularly updated repository containing all the releases of LLMs as they happen? TBH I am tired of having to doommark (doom-bookmark) so many repositories and links...Would appreciate some collected database.
- adt 4y agoThis is close, table of LLMs as released, and I try and add repos for the 'open' models: https://lifearchitect.ai/models-table/ https://lifearchitect.ai/models-table/
- ioulaum 4y agoI wonder if they've done some Alpaca style training on it... Granted, what made Alpaca useful was that it was finetuned with GPT-3's instruction following completions as examples. And, at least officially, OpenAI's outputs can't be used to train other AI models. Otherwise, if GPT-4 outputs were used to finetune these models, they may become much more interesting.