8 ms·
Refact Code LLM: 1.6B LLM for code that reaches 32% HumanEval
- kateklink 3y agoWe’ve finished training a new code model Refact LLM which took us about a month. The main use-case is for blazing-fast code completion with fill-in-the-middle, additionally, the model could reply to chat prompts. It has much better performance than all of the code models of similar size, and almost reaches the same HumanEval as Starcoder being 10x smaller in size. With the small size, it can work with most modern GPUs requiring just 3GB Ram. You can try self-hosting it in Refact https://github.com/smallcloudai/refact/ https://github.com/smallcloudai/refact/ and get a local fast copilot alternative with decent suggestions. Weights and model card https://huggingface.co/smallcloudai/Refact-1_6B-fim https://huggingface.co/smallcloudai/Refact-1_6B-fim. We would love to hear your feedback!
- sparrow0519 3y agohi, i try to fine tune refact model using evolve code alpaca, but the loss is always bigger than 2, i try some different params but it doesn't work, can you give me some advice?
- drcongo 3y agoIs it possible to run it as an LSP so that it can be used in editors other than VSCode and JetBrains? (sorry if this question is completely mad, my understanding of how these things work is extremely limited)
- OlegKlimov1337 3y agoYes, it's coming up in a couple of weeks.
- drcongo 3y agoGreat, thanks. I'll keep an eye out.
- riku_iki 3y ago> almost reaches the same HumanEval how can you tell that HumanEval is not leaked to your training data in some form?
- mityamitya 3y agoHi! We ran LSH filtering over datasets to remove all code that can be similar to HumanEval samples.
- riku_iki 3y agoso, we have to trust your procedure..
- JegernOUTT 3y agoIt can be checked if the model predicts canonical solutions from humaneval. I understand it is not ideal, but at least you can check it yourself There are a bunch of other benchmarks too, check out the page https://huggingface.co/smallcloudai/Refact-1_6B-fim https://huggingface.co/smallcloudai/Refact-1_6B-fim Also, feel free to run any new benchmarks
- diminish 3y agoDoes ctransformer (https://github.com/marella/ctransformers#supported-models https://github.com/marella/ctransformers#supported-models) support running refact? I see that model type "gpt_refact" in https://huggingface.co/smallcloudai/Refact-1_6B-fim/blob/main/config.json https://huggingface.co/smallcloudai/Refact-1_6B-fim/blob/mai...
- deleted 3y ago[deleted]
- ALittleLight 3y agoHow does it compare to Copilot? A metric I'd like to see is % of proposed completions accepted by a human user. If you had an extension that 50% of the time proposed a Copilot extension and 50% of the time proposed a Refact extension (blind to the user) then you could come up with a metric like this.
- Havoc 3y agoThat’s an impressive result The open rail license seems to reference some sort of limitations on safety and unethical use but I can’t see where in the repo that’s spelled out precisely what the authors have in mind?
- deleted 3y ago[deleted]
- iFire 3y agoLICENSE bigscience-openrail-m https://huggingface.co/smallcloudai/Refact-1_6B-fim/blob/main/README.md https://huggingface.co/smallcloudai/Refact-1_6B-fim/blob/mai...
- deleted 3y ago[deleted]
- notsahil 3y agoModel Stats - Architecture: LLAMA-like model with multi-query attention - Objectives Fill-in-the-Middle, Chat - Tokens context: 4096 - Pretraining tokens: 1.2T - Finetuning tokens: 40B - Precision: bfloat16 - GPUs 64 NVidia A5000 - Training time 28 days
- deleted 3y ago[deleted]
- _xnmw 3y agoFor the sake of not giving Microsoft and a few other tech giants immense power over the world, I really do hope the cost and efficiency of LLMs improve dramatically, until we can get GPT-4-equivalent models trained on a few graphics cards and running offline on an iPhone. Really rooting for these kinds of projects until someone makes the breakthrough.
- taywrobel 3y agoYou may be interested in what we’re working on at Symbolica AI. We’re using formal logic in the form of abstract rewrite systems over a causal graph to perform geometric deep learning. In theory it should be able to learn the same topological structure of data that neural networks do, but using entirely discrete operations and without the random walk inherent to stochastic gradient descent. Current experiments are really promising, and assuming the growth curve continues as we scale up you should be able to train a GPT-4 scale LLM in a few weeks on commodity hardware (we are using a desktop with 4 4090’s currently), and be able to do both inference and continual fine tuning/online learning on device.
- pawelduda 3y agoSounds cool, but what are the drawbacks?
- k__ 3y agoIt doesn't exist at scale yet.
- taywrobel 3y agoBiggest drawback is that since the structure is all discrete, it is inherently weak at modeling statistical distributions. For example, it'll likely never best a neural network at stock market prediction or medical data extrapolation. However, for things that are discrete and/or causal in nature, we expect it to outperform deep learning by a wide margin. We're focused on language to start, but want to eventually target planning and controls problems as well, such as self-driving and robotics. Another drawback is that the algorithm as it stands today is based on a subgraph isomorphism search, which is hard. Not hard as in tricky to get right like Paxos or other complex algorithms; like NP-Hard, so very difficult to scale. We have some fantastic Ph.Ds working with us who focus on optimization of subgraph isomorphism search, and category theorists working to formalize what constraints we can relax without effecting the learning mechanism of the rewrite system, so we're confident that it's achievable, but the time horizon is unknown currently.
- howon92 3y agoCongrats on your achievement! I'm curious about your end goal. Do you aim to beat GitHub Copilot's performance and convince devs to use Refact for code completion instead of GitHub Copilot? I want to understand the motivation behind these different code-completion models that are not solely for academic research.
- kateklink 3y agowe want to help developers who need either on-premise or permissive code assistant, copilot has neither of this. We also wanted to lower the barriers for self-hosting, so that the model is available on most GPUs with just 3GB Ram. Plus making the code completions fast and efficient (understanding entire context, not just the previous tokens).
- OlegKlimov1337 3y agoYou can use it in practice, that was the goal of that particular model! It's fast, runs on your own hardware if you want it to.
- glutamate 3y agoLicense text: https://drive.google.com/file/d/16NqKiAkzyZ55NClubCIFup8pT2jnyVIo/view https://drive.google.com/file/d/16NqKiAkzyZ55NClubCIFup8pT2j... [PDF] See last page for restrictions
- Havoc 3y agoThanks. That looks pretty relaxed on terms
- lordofgibbons 3y ago> In any way that violates any applicable national, federal, state, local or international law or regulation; Darn! Foiled again! I was planning on breaking some federal laws, but the license says that I can't ;( \s Open-RAIL license has the be the worst license in existence claiming to be "open". > You shall undertake reasonable efforts to use the latest version of the Model. Plea to folks releasing models: Please stop using this user-hostile and deranged license
- umutisik 3y agoThe title is misleading This model is not "SOTA for the size", there are smaller models that do 10-18% better in absolute score. The text says it's SOTA "among similar models" where they probably compare with other models with permissive licensing.
- mrob 3y ago"Permissive" usually refers to Free Software or Open Source licenses without copyleft requirements. OpenRAIL is a proprietary license because it imposes usage restrictions, contrary to both the Free Software and Open Source definitions.
- OlegKlimov1337 3y agoAFAIK There is only one model that do better, it’s phi-1 and it’s python only, and it does not support fill-in-the-middle so you can't really use it.
- umutisik 3y agoPhi-1-small also scores higher with 350M parameters. It helps to be specific about what the comparison is against when claiming SOTA.
- ldjkfkdsjnv 3y agoI dont trust any benchmarks for any LLM thats not coming from FB, Google, OpenAI, Anthropic, or Microsoft. These models are so dynamic, the simple benchmark numbers never tell the whole story of the quality of the model. Take for instance, a recent posting by anyscale, claiming their fine tuning of Llama 2 was competitive with OpenAI's model. The reality being their fined tuned model is basically worthless, and was competitive along a single metric/very narrow commoditized task. Its a great way to get clicks by posting these metrics though
- SparkyMcUnicorn 3y agoThe community has fine-tuned some really good llama models (much better than llama-chat), but I get what you're saying. I've been testing the best performing models on the huggingface leaderboard lately. Some of them are really impressive, and others are so bad that I second guess the prompt format or if the benchmarked model is actually the same one I'm testing.
- breadsniffer01 3y agoWhich models were really bad?
- SparkyMcUnicorn 3y agoI was keeping track of the good ones, and don't have many notes on the bad ones. I do remember testing "LoKuS" last week and it was quite terrible (sometimes gave completely off-topic answers). It scored as one of the highest 13B models on the leaderboard (~65 average), but appears to be removed now.
- nomel 3y agoThis is the goal of humaneval, correct?
- breadsniffer01 3y agoThey could have easily benchmarked with the Spider SQL test set but they didn’t. I have a feeling that the more robust models might be the ones that don’t perform best on benchmarks right away.
- mholubowski 3y agoHey, I have a genuine question: What is the point of a new model that isn’t better than the best possible model (example: OpenAI GPT-4)? What’s the point in having a smaller model? Who cares? —- This is a real, genuine question that I don’t have a clear answer to. Excuse my ignorance, plz enlighten your boi.
- deleted 3y ago[deleted]
- yieldcrv 3y ago1) people can run a 1.6B model for free on consumer hardware 2) any model that's run on computational resources you are owning or leasing will have more privacy than an explicit cloud offering. running completely on your own local hardware will be private. this means you don't have to think twice about asking the LLM about the proprietary code or information you are working on. 3) smaller models gain the performance improvements from all the other improvements in interpreters and quantizing, allowing for even more consumer friendly offline use 4) oh yeah, offline use. could expand use cases to having LLM's baked into operating systems directly, including leading phones 5) showing what's possible, pushing towards the benchmarks of the best possible model while using less computational resources. this also makes the hosts of the best possible model realize that they could either A) be using less computational resources and increasing the bandwidth for their users B) further improve their own model because of competition. Basically if ChatGPT 4 was using similar improvements in technology across all areas of reasoning/whatever, there never would have been a rate limit on ChatGPT 4. 6) more demand for other computational resources. Nvidia is backordered till maybe Q2 2024 right now. If people realize AMD or even their ARM chips can offer same performance with the right combination of hardware and software, It alleviates pressure on other ventures that want computation power.
- SparkyMcUnicorn 3y agoYou can use it 100% locally, and it doesn't cost anything.
- yunwal 3y agoGPT4 is expensive to run, even more expensive to finetune, and for all practical purposes can’t be run offline (because the model is too big to run outside of a huge data center). Evaluation latency is also an issue for many usecases, and you have to share your query with openai, so you can’t run sensitive queries. The output is also controlled/censored by OpenAI. Here’s a few usecases that I wouldn’t want to use OpenAI/GPT for - Advanced autocomplete for texting and private communications - Querying sensitive document databases like emails - Traveling in low connectivity areas - Politically incorrect usecases (generating erotic content for example) List kinda goes on and on
- _lvbh 3y agoSay I want to fine tune a Golang specific model. How much $ and effort would I have to put in? Would using this as a base help in any way compared to starting from llama?
- OlegKlimov1337 3y agoMaybe it makes sense to start from llama-code, not llama :D I think golang specific model will not be that different from a multi-language model. But it definitely will work better after fine tuning on your code. Check out refact self hosting docker in a couple of days, finetune will be there soon. It will take you 1 GPU and almost no money )
- vikp 3y agoThis post is misleading, in a way that is hard to do accidentally. - They compare the performance of this model to the worst 7B code llama model. The base code llama 7B python model scores 38.4% on humaneval, versus the non-python model, which only scores 33%. - They compare their instruct tuned model to non-instruct-tuned models. Instruction tuning can add 20% or more to humaneval performance. For example, WizardLM 7B scores 55% on humaneval [1], and I've trained a 7B model that scores 62% [2]. - For another example of instruction tuning, Stablecode instruct tuned benchmarks at 26%, not the 20% they cite for the base model [3] - Starcoder, when prompted properly, scores 40% on humaneval [4] - They do not report their base model performance (as far as I can tell) This is interesting work, and a good contribution, but it's important to compare similar models. [1] https://github.com/nlpxucan/WizardLM https://github.com/nlpxucan/WizardLM [2] https://huggingface.co/vikp/llama_coder https://huggingface.co/vikp/llama_coder [3] https://stability.ai/blog/stablecode-llm-generative-ai-coding https://stability.ai/blog/stablecode-llm-generative-ai-codin... [4] https://github.com/huggingface/blog/blob/main/starcoder.md https://github.com/huggingface/blog/blob/main/starcoder.md
- JegernOUTT 3y agoHi, thank you for your attention! > They compare the performance of this model to the worst 7B code llama model. The base code llama 7B python model scores 38.4% on humaneval, versus the non-python model, which only scores 33%. We are comparing multilingual models, and we are not focused on python-finetuned versions > They compare their instruct tuned model to non-instruct-tuned models. Instruction tuning can add 20% or more to humaneval performance. For example, WizardLM 7B scores 55% on humaneval [1], and I've trained a 7B model that scores 62% [2]. > For another example of instruction tuning, Stablecode instruct tuned benchmarks at 26%, not the 20% they cite for the base model [3] We have two separate comparisons (see https://huggingface.co/smallcloudai/Refact-1_6B-fim https://huggingface.co/smallcloudai/Refact-1_6B-fim) for completion-based models and instruction-following-based models with different humaneval formats. But we are considering our model as a completion (FIM) one in the first place and we were using 85% non-instruction following data to make the final model. The chat functionality is really limited for such small models > Starcoder, when prompted properly, scores 40% on humaneval Yep, that is right. But worth mentioning, the starcoder model showed 40% while being extra finetuned exclusively on python > They do not report their base model performance (as far as I can tell) Our base model gets around 20-23% humaneval. But it is not the case since the model was trained using 50% non-code data (considering the model's size it was really hard to keep the model converging)
- brucethemoose2 3y agoOne misleading thing is the notion that you need a 1-2B model to run on commodity hardware. This is not really true. Llama 7B runs with Vulkan/llama.cpp on ~8GB smartphones and ~12GB laptops. That ease is going to get much better over time, as lower RAM hardware starts dropping out of the market and the Vulkan implementations get more widespread. For users trying to run LLMs on 8GB or less machines, the AI Horde approach of distributed models seems much more practical anyway.
- palmer_fox 3y agoPerhaps the wrong thread to ask this question... Is it not possible to load a model on something like an NVMe M.2 drive instead of RAM? It's slower of course, but only 5-10x if I understand correctly.
- kirill5pol 3y agoYes but they’re slow enough on normal hardware for that 5-10x to be painful…
- mirekrusin 3y agoCan you RAID them?
- brucethemoose2 3y agoTechnically yes? But its way beyond the point where its going to help LLMs. CPU RAM is already "too slow" in machines big enough for multiple NVMe SSDs.
- naillo 3y ago7b runs on my 4gb vram machine (8gb memory). I.e. quantization helps a lot too
- btown 3y agoAh, but have no fear - as lower RAM hardware starts dropping out of the market, the RAM usage of Microsoft Teams will increase to compensate! (Not even /s - while the developers of LLM applications may have 64GB RAM in their laptops or desktops, the less-technical early adopters of LLMs running locally are likely to be power users with lower-powered laptops, much more stringent RAM limits, and numerous line-of-business applications and browser tabs contending for that RAM. Causing those applications to be swapped onto disk will almost certainly result in a degraded overall experience that could easily be blamed on the LLM application itself.)
- zcesur 3y agotangentially related: refact recently shared 4 bounties worth $9,000 to help improve their tech! https://algora.io/org/smallcloudai/bounties https://algora.io/org/smallcloudai/bounties disclaimer: i'm a cofounder of algora, the platform enabling these bounties
- deleted 3y ago[deleted]
- palmer_fox 3y agoAll these LLMs are pretty general if I understand correctly. Are there any efforts to create specialized models (other than for coding)? Or, what would be even better, "extract" certain areas from existing LLMs as a way to specialize them? With the goal to drastically reduce model size to be able to run on less powerful devices. E.g. a model specializing in chemistry doesn't need to include data on world's history or to be able to write poetry.
- hnhg 3y agoI am not an expert but it still has to learn human language/grammar/whathaveyou, and that is where scale seems to matter. Fine-tuning on a subset of knowledge after that is typically how the domain-specialisation is achieved, by my understanding.
- charcircuit 3y agoDomain specialization is done by continuing the full training process. Fine tuning is more for changing the style of the output than adding new knowledge.
- palmer_fox 3y agoWhat if the initial training already contains all necessary data for a particular specialization? What would be the benefit of continuing the training process?
- viraptor 3y agoImagine someone tells you about how someone committed a crime and asks you to summarise. Now imagine the same question is asked to a lawyer. Even if you both knew the same facts, the response would be very different in style, highlighted points, mentioned references, etc. The domain specific fine tuning does exactly that. Sure, sometimes you can get very close by changing the prompt to include "respond like a lawyer in situation X with following extra rules", but not always and the fine-tuning gives better results and shorter prompt.
- holoduke 3y agoWhats the difference between 1% and 99% of HumanEval? What does it tell really?
- kateklink 3y agofor pass@1 HumanEval tells how well the model solves a task from a set, given only one chance to solve it. It's not the perfect metric, there're other like DS-1000, MBPP (we have included them on HuggingFace model card). HumanEval is good for benchmarking with other models as it gives a fast idea how powerful the model is.
- swyx 3y ago> given only one chance to solve it my understanding is that there are 2 usages of the pass@{number} syntax. the HumanEval/Codex paper interprets the {number} as number of attempts[0]. however language modelers seem to use it to denote the number of few shot example demonstrations given in the context. these are starkly different and i wish the syntax wasnt overloaded --- [0] https://arxiv.org/pdf/2107.03374.pdf https://arxiv.org/pdf/2107.03374.pdf > Kulal et al. (2019) evaluate functional correctness using the pass@k metric, where k code samples are generated per problem, a problem is considered solved if any sample passes the unit tests, and the total fraction of problems solved is reported.
- Manjuuu 3y agoAnother model that we'll soon forget it ever existed.
- smcleod 3y agoJust trying out the official container image for self-hosting along side the VSCode extension - I've got to say I'm really impressed with the scaffolding especially for an early stage project. The web interface for the LLM server is especially nice and clean compared to many of the others I've tried - and it "just works". Very interested to see how this evolves.