16 ms·
OpenLLaMA: An Open Reproduction of LLaMA
- venelin_valkov 3y agoI made a YouTube video on how to run OpenLLaMa on Google Colab with Hugging Face Transformers (using a T4 GPU): https://www.youtube.com/watch?v=1NOPciKuQb8 https://www.youtube.com/watch?v=1NOPciKuQb8 Hope that helps!
- logicchains 3y agoIt's not clear from the GitHub; are there any plans to eventually train the 30 or 65 billion weight LLaMA models? The 65B model seems comparable to GPT3.5 for many things, and can run fine on a beefy desktop just on CPU (CPU ram is much cheaper than GPU ram). It'd be amazing to have an open source version.
- tarruda 3y agoI ran the 30b and 65b Q4 on a laptop with 64 gb of RAM (8/16 CPU). It worked but token/s was very low for it to be practically useful.
- bagels 3y agoHow low? I think everybody has different requirements there.
- extasia 3y agoI ran it on a modern desktop and was getting sub 1 token/s
- asah 3y agocould it parallelize across multiple PCs ?
- serialx 3y agoNo since it’s stateful in the sense that inferencing is dependent on the past generated tokens.
- brutus1213 3y agoI'm really curious how Meta, DeepMind and OpenAI make the big models work. The biggest A100 you can buy is just 80GB. And I assume the big companies use single precision floating point during training. Are they actually partitioning the big model across multiple GPU instances? If one had the hardware, how many GPUs does the biggest LLAMA take? These are systems issues and I have not read papers or blog posts on how this works. To me, this infra is very non-trivial.
- 15155 3y agoNVLink
- spi 3y agoThe "standard" machine for these things has 8x80GB = 640GB memory (p4de instances here: https://aws.amazon.com/ec2/instance-types/p4/ https://aws.amazon.com/ec2/instance-types/p4/), with _very_ fast connections between GPUs. This fits even a large model comfortably. Nowadays probably most training use half precision ("bf16", not exactly float16, but still 2 bytes per parameter). However during training you easily get a 10-20x factor between the number of parameters and the bytes of memory needed, due to additional things you have to store in memory (activations, gradients, etc.). So in practice the largest models (70-175B parameters) can't be trained even on one of these beefy machines. And even if you could, it would be awfully slow. In practice, they typically use servers with clusters of these machines, up to about 1000 GPUs in total (so around 80TB of memory, give or take a few?). This allows even the biggest models to be trained on large batches of several hundreds, or even thousands, of elements (the total memory usage is _not_ proportional to the product of number of parameters and the batch size, but it does increase as a function of both of them, a term of which being indeed the product of the two). It makes for some very tricky engineering choices to make just the right data travel across connections, trying to avoid as much as possible that you have to sync large amount of data between different machines (so "chunking" things to stay on the 640GB range) with strategies such as ZeRO being published every now and then. Plus of course the practical effort to make physical connections as fast as possible... To get an idea of how hard these things are, take a look at how long the list of names in the published paper about BLOOM language model is :-)
- tarruda 3y agoI didn't measure, but IIRC it was lower than 1 token/sec
- simion314 3y agoslow could be useful if you do not want to chat with it, and instead you could code it to do a long running job, like code review your entire project like a code analysis tool. Or summarize a lot of content.
- logicchains 3y agoThat's unfortunate. Running the 65B Q4 on an AMD Epyc with 32 1.5ghz cores and 256 GB of ram I get around 3 tokens/sec, which is useable if not ideal. I wonder if the difference is related to the RAM or the number of CPUs?
- tarruda 3y ago3 tokens/sec is a lot faster than what I experienced. Even though your CPU has a lot more cores, I think llama.cpp was not being able to make good use of more than 8 threads. When did you test this? Maybe llama.cpp had some improvements since I used it (which was at the start of the project).
- logicchains 3y agoI tested this on the latest master. Llama.cpp has had some performance improvements, although I don't know if that'd be enough to make it 3x faster.
- Ambix 3y agoIt's not about threads number, it about memory bottleneck. Sweet spot for my M1 Pro laptop is around 6 threads and 4bit model - I've managed to get 20 tokens per sec, really impressive
- azeirah 3y agoI'm following the discussions on GitHub as well as their PRs closely. The primary bottleneck for now is compute. They've recently made a big improvement to performance by introducing partial gpu acceleration if you compile with a gpu accelerated variant of BLAS. Either cublas (Nvidia) or CLBlast (slightly slower but supports almost everything: Nvidia, Apple, AMD, mobile, raspberry pi etc)
- lhl 3y agoAlthough there are multiple bottlenecks, my understanding (and why at a certain point, throwing more threads doesn't work) is that inference for dense LLMs are largely limited by memory bandwidth. Most desktop computers will have dual channel DDR4/DDR5 memory which will be hard pressed to get >60GB/s. A last-gen Epyc/Threadripper Pro should have 8 channel memory DDR4-3200 support, which should get you a theoretical max of 204.8 GB/s (benchmarking ends up more around 150GB/s in AIDA64). The latest Genoa has 12 channel DDR5-4800 support (and boosted AVX-512) and I'd imagine should perform quite well, but if you primarily want to run inference on a quantized 65B model, I think you're best bang/buck (for local hardware) would be 2 x RTX 3090s (each of those has 24GB of GDDR6X w/ just shy of 1TB/s of memory bandwidth).
- quickthrower2 3y agoIf I rent an A100 what kind of speed could I expect?
- GC_tris 3y agoWhile I do not have any A100 handy right now I have an instance running on Genesis Cloud with 4x RTX 3090. A quick, very unscientific, test using the oobabooba/text-generation-webui with some models I tried earlier gives me: * oasst-sft-7-llama-30b (spread over 4x GPU): Output generated in 28.26 seconds (5.77 tokens/s, 163 tokens, context 55, seed 1589698825) * llama-30b-4bit-128g (only using 1 GPU as it is so small): Output generated in 12.88 seconds (6.29 tokens/s, 81 tokens, context 308, seed 1374806153) * llama-65b-4bit-128g (only using 2 GPU): Output generated in 33.36 seconds (3.81 tokens/s, 127 tokens, context 94, seed 512503086) * llama (vanilla, using 4x GPU): Output generated in 5.75 seconds (4.69 tokens/s, 27 tokens, context 160, seed 1561420693) They all feel fast enough for interactive use. If you do not have an interface that streams the output (so you can see it progressing) it might feel a bit weird if you often have to wait ~30s to get the whole output chunk.
- wokwokwok 3y agoThere’s a lot of controversy about “7B is good enough and small enough for consumer hardware so it’s good enough fullstop” …but, although it is true that for a fixed compute budget that these small models can have impressive results with good training data, it is also true that smaller models (7B) appear to have an upper performance bound that is beaten easily by larger well trained models. It’s just way more expensive to train larger models. They specifically note they are training a smaller 3B model In the future. So… it seems reasonable to assume that this is a proof of concept, and that no, the Berkeley AI lab will not be fielding the cost for training a larger model. This is probably more about exploring the “can we make a cheap good-enough model?” than “here is your GPT4 replacement”.
- scotty79 3y agoDo you know of any research that tries to take large pre-trained model and make it smaller by cutting out least activated neurons and training it a bit not to loose performance?
- sebzim4500 3y agohttps://arxiv.org/pdf/2301.00774.pdf https://arxiv.org/pdf/2301.00774.pdf
- KRAKRISMOTT 3y agoThe entire field of ML distillation.
- moffkalast 3y ago> They specifically note they are training a smaller 3B model In the future. They're kidding right, there's no way that thing will be more useful than one of those flan models.
- b33j0r 3y agoAgreed. With some work, 13B runs on consumer hardware at this point. That redefines consumer to a 3090 (but hey, some depressed crypto guys are selling them. I recently got another GPU for my homelab this way). 30B is within reach, with compression techniques that seem to lose very little information of the overall network. Many argue that machine learning IS fundamentally a compression technique, but the topology of the trained network turns out to be more important. Assuming an appropriate activation function after this transformation. No… definitely not your GPT4 replacement. However this is the kind of PoC I keep following… every… 18 hours or so? Amazing.
- newswasboring 3y agoAt least for now they are focused on 7B and then 3B[1]. [1]https://github.com/openlm-research/open_llama#future-plans https://github.com/openlm-research/open_llama#future-plans
- Silverback_VII 3y agoI'm not sure whether the number of parameters serves as a reliable measure of quality. I believe that these models have a lot of redundant computation and could be a lot smaller without losing quality.
- cubefox 3y agoThe Chinchilla scaling law describes, apart from the training data size, the optimal number of parameters for a given amount of computing power for training. See https://dynomight.net/scaling/ https://dynomight.net/scaling/
- sp332 3y agoFor training, yes, but these models are optimized for inference, since inference will be run many more times than training. The original Llama models were run way past chinchilla-optimal amounts of data.
- bluecoconut 3y agoReally exciting how fast fully pre-trained new models are appearing. Here's another repo (with the same "open-llama" name) that has been available on hugging face as well for a few weeks. (different training dataset) https://github.com/s-JoL/Open-Llama https://github.com/s-JoL/Open-Llama https://huggingface.co/s-JoL/Open-Llama-V1 https://huggingface.co/s-JoL/Open-Llama-V1
- deleted 3y ago[deleted]
- LudwigNagasena 3y agoIs anyone familiar with the BOINC-style grid computing scene for ML and, specifically, LLM? Is there something interesting going on, or is it infeasible? Will things like OpenLLaMA help it?
- literalAardvark 3y agoThey seem to scale up, not out, so grids don't really work. What everyone is using are HPC grade low latency interconnects to make the cluster look as close as possible to a single big TPU.
- pmoriarty 3y ago"They seem to scale up, not out, so grids don't really work." Can someone explain what this means? I don't understand.
- balloonfencist 3y agoUp=bigger machine Out=lots of machines through network
- natmaka 3y agohttps://en.wikipedia.org/wiki/Scalability#Horizontal_or_scale_out https://en.wikipedia.org/wiki/Scalability#Horizontal_or_scal...
- moffkalast 3y agohttps://openmetal.io/docs/edu/openstack/horizontal-scaling-vs-vertical-scaling https://openmetal.io/docs/edu/openstack/horizontal-scaling-v... In a typical fully connected hidden layer, the neurons each need to compute the values of the all others in the previous layer, so you need all the data in one place. Obviously you can distribute the actual calculations which is what a GPU does, but distributing that over networked CPUs will be incredibly slow and require the whole thing to be loaded into memory on all instances. My bet is on some kind of light based or analog electric accelerator PCIE card to be the next best thing for this sort of inference, since it should be able to calculate multiple layers at once. FPGAs also work but only for fixed weights.
- bighoki2885000 3y ago[dead]
- scotty79 3y agoMotivation?
- newswasboring 3y agoSadly, licensing.
- igravious 3y agoHappily, licensing.
- newswasboring 3y agowhy the hell will you be happy about duplicate work?
- zirgs 3y agoGood luck convincing Meta to release their models with a proper licence.
- newswasboring 3y agoThat is why its sadly, licensing.
- rodoxcasta 3y agoActually, replication is very important. If no one can make new llamas, that would mean that facebook used some secret sauce in their training. Understanding publicly how to train these 'enhanced' models that shows performance of much greater models is a very strong motive. And getting hid of the NC clause of the original llamas too, of course. As of right now, there's trouble replicating the eval results of the paper, for example.
- newswasboring 3y agoYeah but that wasn't the reason, was it? They didn't do it because they wanted to replicate work, they did it because they didn't want the Meta lawyers to be big mad at them.
- quickthrower2 3y agoSo is this free as in “do what you f’ing like with it”?
- mkl 3y agoMostly, yes. It's Apache License 2.0: https://github.com/openlm-research/open_llama/blob/main/LICENSE https://github.com/openlm-research/open_llama/blob/main/LICE...
- jasonm23 3y agoForgive me for the ignorance, but can a refined training model be a specific codebase, after say training on all standard docs for the language, and 3rd party libs, and so on. I have no formal idea how this is done, but my assumption is that "something like that" should work. Please disabuse me of any silly ideas.
- charcircuit 3y agoYou can train the model on more training data after it has been released.
- yakorevivan 3y ago[dead]
- heliophobicdude 3y agoHi Jason! I have a few thoughts on this! Refined training is usually updating the weights of usually what's called a foundational model with well structured and numerous data. It's very expensive and can disrupt the usefulness of having all the generalizations baked in from training data [1]. While LLMs can generate text based on a wide range of inputs, they're not designed to retrieve specific pieces of information in the same way that a database or a search engine would. But I do think they hold a lot of promise in reasoning. Small corollary: LLMs do not know a head of time what they are generating. Secondly, they use the input from you and itself to drive the next message. This sets us up for a strategy called in-context learning [1]. We take advantage of the above corollary and prime the model with context to drive the next message. In your case, a query about some specific code base with knowledge about standard docs etc. Only there is a big problem, context sizes. Damn. 4k tokens? We can be clever about this but there is still a lot of work and research needed. We can take all that code and standard docs and create embeddings of them [2]. Embeddings are mathematical representations of words or phrases that capture some of their semantic meaning. Basically the state of a trained neural network given inputs. This will allow us to group similar words and concepts together closer in what is called a vector space. We can then do the same for our query and iterate over each pair finding the top-k or whatever most similar pairs. Many ways to find the most similar pairs but what's nice is cosine similarity search. Basically a fancy dot product of the pairs with a higher score indicating greater similarity. This will allow us to prime our model with the most "relevant" information to deal with the context limit. We can hope that the LLM would reason about the information just right and voila. So yeah basically create a fancy information retrieval system that picks the most relevant information to give your model to reason about (basically this [3]). That and while also skirting around the context limitations and not overfitting and narrowing the training information that allow them to reason (controversial). 1: "Language Models are Few-Shot Learners" Brown et al. https://arxiv.org/pdf/2005.14165.pdf https://arxiv.org/pdf/2005.14165.pdf 2: Embeddings https://arxiv.org/pdf/2201.10005.pdf https://arxiv.org/pdf/2201.10005.pdf 3: https://twitter.com/marktenenholtz/status/1651568107192983553 https://twitter.com/marktenenholtz/status/165156810719298355...
- newswasboring 3y agoHow is this model performing better than LLaMa in a lot of tasks[1] even though its trained on a fifth of the data (1 trillion vs 200 billion). [1]https://github.com/openlm-research/open_llama#evaluation https://github.com/openlm-research/open_llama#evaluation
- slekker 3y agoNobody knows :^)
- tarruda 3y agoMaybe it uses a higher quality dataset
- YetAnotherNick 3y agoThey are likely doing some interpolation for 200B or benchmarking it in wrong way. e.g. Hellaswag accuracy for llama 7b is 0.76[1], but it is written 0.56 in the repo. Even at 200B tokens, it is higher than 0.56 for llama looking at the charts. [1]: https://arxiv.org/pdf/2302.13971.pdf https://arxiv.org/pdf/2302.13971.pdf
- byefruit 3y agoThey ran lm-evaluation-harness on both this model and the original llama weights, which is the correct way to do it. Many people have been struggling to reproduce the benchmark numbers included in the original llama paper.
- vrglvrglvrgl 3y ago[dead]
- logicchains 3y agoWould be very interesting to see https://github.com/BlinkDL/RWKV-LM https://github.com/BlinkDL/RWKV-LM trained on the same data
- leobg 3y agoInteresting. Have you done anything with RWKV?
- logicchains 3y agoNope, not yet, the current 14B version is much worse than LLaMA 65B. But there are apparently plans to train a RWKV-65B by the end of the year, and if including the LLaMA training dataset results in something like LLaMA-65B but with infinite context then that'd be really amazing.
- vessenes 3y agoI evaluated RWKV recently, and it's interesting for sure. It's undertrained, and has a quirky architect, so some parts of it are different than playing with the llama ecosystem. The huge context length is super appealing, and in my tests, long prompts do seem to work and get coherent results. Where it's slow is in tokenization -- it can be very, very slow to make an initial tokenization of a prompt. I think this has to do with how the network actually functions, like there's a forward loop that feeds each token in to the network sequentially. I would guess if it had the same level of attention and work that the Llama stack is getting it would be pretty fantastic, but that's just a guess, I'm a hobbyist only.
- ianpurton 3y ago> We are currently focused on completing the training process on the entire RedPajama dataset. So that's 1.2 trillion tokens. Nice.
- quickthrower2 3y agoI am quite new to this, I would like to get it running. Would the process roughly be: 1. Get a machine with decent GPU, probably rent cloud GPU. 2. On that machine download the weights/model/vocab files from https://huggingface.co/openlm-research/open_llama_7b_preview_200bt/tree/main https://huggingface.co/openlm-research/open_llama_7b_preview... 3. Install Anaconda. Clone https://github.com/young-geng/EasyLM/ https://github.com/young-geng/EasyLM/. 4. Install EasyLM: conda env create -f scripts/gpu_environment.yml conda activate EasyLM 5. Run this command, as per https://github.com/young-geng/EasyLM/blob/main/docs/llama.md https://github.com/young-geng/EasyLM/blob/main/docs/llama.md: python -m EasyLM.models.llama.llama_serve \ --mesh_dim='1,1,-1' \ --load_llama_config='13B' \ --load_checkpoint='params::path/to/easylm/llama/checkpoint' \ Am I even close?
- jbandela1 3y agoI think llama.cpp might be easier to set up and get running. https://github.com/ggerganov/llama.cpp https://github.com/ggerganov/llama.cpp
- JLCarveth 3y agoYes, I can clone this and get into a prompt in less than 5 minutes on an M2 MBA.
- quickthrower2 3y agomight try it first. seems to be only CPU?
- themulticaster 3y agoI'd see that as a benefit of llama.cpp - it's specifically designed to be usable on consumer hardware such as laptops, without professional GPUs.
- azeirah 3y agoIt has partial gpu acceleration if you compile it with LLAMA_CUBLAS or LLAMA_CLBLAST They really have come a long way since... A few weeks ago. Using cublas with my 1080ti results in a 52% speedup compared to cpu-only. Vram usage is very minimal.
- Eduard 3y agoCan someone explain how to tell if a model doesn't require a GPU and can run on a CPU? After setting up dalai, OpenAssistant, gpt4all and a bunch of other (albeit nonworking) LLM thingies, my current hunch is: if the model somewhere has "GGML" in its name, it doesn't require a GPU.
- execveat 3y agoTechnically anything that's based on pytorch can run on CPU, you just need to tell it to do so. For example, in textgen add '--cpu' and you're done. It will be super slow though. GGML format is meant to be executed through llama.cpp, which doesn't use GPU by default. You can often find these models in a quantized form as well, which helps performance (at a cost of accuracy). Look for q4_0 for the fastest performance and lowest RAM requirements, look for 5_1 for the best quality right now (well, among quantized models). Oh yeah, textgen supports llama.cpp, and also provides API, so it looks like a clear winner. You might want to manually pull newer dependencies for torch and llama.cpp though: pip install -U --pre torch torchvision -f https://download.pytorch.org/whl/nightly/cpu/torch_nightly.html https://download.pytorch.org/whl/nightly/cpu/torch_nightly.h... pip install -U llama-cpp-python
- Taek 3y agoHow is this different from what RedPajamas is doing? Also, most people don't mind running LLaMA 7B at home so much because of enforceability, but a lot of commercial businesses would love to run a 65b parameter model if possible and can't because the license is more meaningfully prohibitive in a business context. Open versions of the larger models are a lot more meaningful to society at this point.
- execveat 3y agoRedPajama is creating a dataset. This is a permissively licensed model trained on that dataset.
- slama 3y agoRedPajama is also training both foundation and instruct-tuned models Source: https://twitter.com/togethercompute/status/1652735096150175744 https://twitter.com/togethercompute/status/16527350961501757...
- bradleyjg 3y agoI agree with this. For a lot of companies hundreds of thousands of dollars or single digit millions on fine tuning, inference, and so on is entirely feasible but using model weights with clouded legal status isn’t.
- superpope99 3y agoI'm always curious about the cost of these training runs. Some back of the envelope calculations: > Overall we reach a throughput of over 1900 tokens / second / TPU-v4 chip in our training run 1 trillion / 1900 = 526315789 chip seconds ~= 150000 chip hours. Assuming "on-demand" pricing [1] that's about $500,000 training cost. [1] https://cloud.google.com/tpu/pricing https://cloud.google.com/tpu/pricing
- p1esk 3y agoAt these levels of spending the actual cost is heavily negotiated and is usually far below the advertised on-demand pricing. Considering I could negotiate A100 for under a dollar/hr - 8 months ago, when they were in high demand, I wouldn't be surprised if the cost was close to 100k for this training run.
- jerrygenser 3y agoThey haven't trained a 1 trillion token model yet. They have only done 200bn so far
- execveat 3y agoNobody in their right mind is using GCE for training. Take a look at real prices: https://vast.ai/ https://vast.ai/
- bravura 3y agoThese nodes typically have slow downstream, and thus are hard to use when training requires pulling a huge dataset.
- simonw 3y agoI got the impression that kind of thing (buying time on GPUs hosted in people's homes) isn't useful for training large models, because model training requires extremely high bandwidth connections between the GPUs such that you effectively need them in the same rack.
- 3y ago
- deleted 3y ago[deleted]
- jjice 3y agoDoes anyone have any resources they recommend for just understanding the base terminology of models like this? I always see the terms "weights", "tokens", "model", etc. I feel like I understand what these mean, but I have no idea what I need to care about them for in open models like this? If I were to download an open model to run on my machine, would I download the weights? I'm just ignorant in the ML space I guess but not sure where to start.
- zoogeny 3y agoAndrej Karpathy's Zero to Hero video series [1] is a good middle ground. It isn't super low-level but it also isn't super high-level. I think seeing how the pieces actually fit together in a working project is valuable to get a real understanding. After going through this series I can say I basically understand weights, tokens, back-propagation, layers, embeddings, etc. 1. https://karpathy.ai/zero-to-hero.html https://karpathy.ai/zero-to-hero.html
- data_maan 3y agoWhen was this published? Is this an older tutorial by Karpathy? Just curious, didn't see any date...
- CamperBob2 3y agoI'm working my way through that series now. He really is a good teacher -- I keep waiting for the inevitable "Next, draw the rest of the fucking owl" moment, but so far he does seem to be sticking to his commitment to a from-scratch approach.
- martythemaniak 3y agoHas anyone successfully used embeddings with anything other than OpenAI's APIs? I've seen lots of debates on using embeddings vs fine-tuning for things like chatbots on private data, but is there a reason why you can't use both? IE, fine-tune LLaMA on your data, then run the same embeddings approach on top of your own fine-tuned model?
- diimdeep 3y agoTo use with llama.cpp on CPU and 8GB RAM git clone https://github.com/ggerganov/llama.cpp && cd llama.cpp && cmake -B build && cmake --build build python3 -m pip install -r requirements.txt cd models && git clone https://huggingface.co/openlm-research/open_llama_7b_preview_200bt/ && cd - python3 convert-pth-to-ggml.py models/open_llama_7b_preview_200bt/open_llama_7b_preview_200bt_transformers_weights 1 ./build/bin/quantize models/open_llama_7b_preview_200bt/open_llama_7b_preview_200bt_transformers_weights/ggml-model-f16.bin models/open_llama_7b_preview_200bt_q5_0.ggml q5_0 ./build/bin/main -m models/open_llama_7b_preview_200bt_q5_0.ggml --ignore-eos -n 1280 -p "Building a website can be done in 10 simple steps:" --mlock
- gigel82 3y agoYou the real MVP! Though I'm getting this error on an Intel macbook (Monterey); it works fine on a Windows11 box: python3 convert-pth-to-ggml.py models/open_llama_7b_preview_200bt/open_llama_7b_preview_200bt_transformers_weights 1 Loading model file models/open_llama_7b_preview_200bt/open_llama_7b_preview_200bt_transformers_weights/pytorch_model-00001-of-00002.bin Traceback (most recent call last): File "/l/llama.cpp/convert-pth-to-ggml.py", line 11, in <module> convert.main(['--outtype', 'f16' if args.ftype == 1 else 'f32', '--', args.dir_model]) File "/l/llama.cpp/convert.py", line 1129, in main model_plus = load_some_model(args.model) File "/l/llama.cpp/convert.py", line 1055, in load_some_model models_plus.append(lazy_load_file(path)) File "/l/llama.cpp/convert.py", line 857, in lazy_load_file raise ValueError(f"unknown format: {path}") ValueError: unknown format: models/open_llama_7b_preview_200bt/open_llama_7b_preview_200bt_transformers_weights/pytorch_model-00001-of-00002.bin
- kdtsh 3y agoI get the same error on an M series MacBook (Ventura). However from the repo README.md it looks like make should work instead of cmake, I’ll give that a try.
- sebastianhoitz 3y agoI had the same issue and then noticed that I need git lfs - otherwise just cloning the repo will not download the weights.
- version_five 3y agoHas anyone actually used this? I poked around and it's so poorly documented that I don't see how one can readily, short of trying to go through the code, understand how to do a minimal run.
- gigel82 3y agoI've used it with llama.cpp; results are not great, but not entirely terrible (I'd say somewhere between GPT-2 and GPT-3). Still, totally free and open source is great and I'm looking forward to more development from them (and others building on top like an RLHF / alpaca / chat kind of thing).
- version_five 3y agoThanks for answering! In my skim of the thread I only saw people mention trying it with llama.cpp. I tried to get his EasyML framework going but could not figure out the parameters I needed. Definitely agree it's great to see real open source models being built.