7 ms·
Some links: - Repo: https://github.com/replit/ReplitLM/tree/main/replit-code-v1-3b https://github.com/replit/ReplitLM/tree/main/replit-code-v1-... - HuggingFa
by amasad 3y ago
Some links:
- Repo: https://github.com/replit/ReplitLM/tree/main/replit-code-v1-3b https://github.com/replit/ReplitLM/tree/main/replit-code-v1-...
- HuggingFace: https://huggingface.co/replit/replit-code-v1-3b https://huggingface.co/replit/replit-code-v1-3b
- Demo: https://huggingface.co/spaces/replit/replit-code-v1-3b-demo https://huggingface.co/spaces/replit/replit-code-v1-3b-demo
- Early benchmark results: https://twitter.com/amasad/status/1651019556423598081 https://twitter.com/amasad/status/1651019556423598081
A lot about this project was surprising. We knew it was going to be good, but didn't expect to be this good -- especially surprising was the finetuned performance boost, and the fact that the model is decent at language tasks and reasoning (in some cases much better than much larger general-purpose models).
It feels like there is a lot more to do with this model, and I have a suspicion you can even make a half-decent chatbot (at least one focused on code) by finetuning it on conversation (and/or instruction) datasets.
Will follow up with a more comprehensive technical report and the UL2R version (fill-in-the-middle support).
- pera 3y agoHi there, I have two question: 1 - Why did you choose Markdown? It seems an odd choice for training a model like this. 2 - Have you tried to train only one single PL and then benchmark it against this more general version?
- amasad 3y ago1- We trained on languages that are most popular on Replit. Markdown is important because you need some amount of natural language in the data, and it will act as a sort of "natural language label" for code. 2- I like how portable it is being a single small model doing a lot of languages. Single code models are an approach that models like Salesforce/Codegen did that, but I believe we beat (or get very close) to their mono models on benchmarks.
- fuzzythinker 3y agoHave you thought of finding or creating something like this [0]? I created this as the basis for my origami folding descriptive language. I tried to find something similar, requirements being both well structured and English-like but couldn't find any, so I created it. The origami folding app will hopefully be out in 2 weeks, so you can see how it's used. [0] https://github.com/fuzzthink/mation-spec https://github.com/fuzzthink/mation-spec
- runnerup 3y agoThey trained on https://huggingface.co/datasets/bigcode/the-stack-dedup https://huggingface.co/datasets/bigcode/the-stack-dedup which is a massive curated dataset accumulated from GitHub. Details are here: https://www.bigcode-project.org/docs/about/the-stack/ https://www.bigcode-project.org/docs/about/the-stack/ Many of the most-represented "languages" on GitHub are actually things like JSON, XML, HTML, CSV, text, markdown, YAML, and SVG. More details from them here: https://blog.replit.com/llm-training https://blog.replit.com/llm-training
- gbasin 3y agoVery exciting, thanks for sharing all this
- letitgo12345 3y agoDoesn't the Stack contain HumanEval? So you're basically comparing numbers on the pretraining data.
- amasad 3y agoCan't find it now but pretty sure BigCode said somewhere they explicitly looked for it and removed it. Also subjective measure does match up to the benchmark. Our finetuned model performed +50% on HumanEval and then when using it felt at least that much improved.
- godelski 3y agoYou can view the prompts, solutions, and checks here[0]. See my sibling comment (to yours) where I quote the Human Eval paper and do some more analysis. But I think if you look at [0] you'll see that these aren't really unique problems and are likely to have large repetitions in the dataset. I should add to that comment to include the dataset[1] (too late to edit) where they mention that they just scrape all of GitHub (Jan 1 2015 - Mar 31 2022). They do exact and near de-duplicate but near de-duplication is messy. > We implement near-deduplication in our pre-processing pipeline on top of exact deduplication. We first split the files into words/tokens based on non-alphanumeric characters and remove files with fewer than 10 tokens. Next, we compute the MinHash with 256 permutations of all documents, and use Locality Sensitive Hashing to find clusters of duplicates. We further reduce these clusters by ensuring that each file in the original cluster is similar to at least one other file in the reduced cluster. We consider two files similar when their Jaccard similarity exceeds 0.85. Near-duplicates are still difficult to measure. So we should expect duplication, and it should be proportional to the number of samples we have (even if the same variance, but I'd wager higher variance with larger duplications). [0] https://github.com/openai/code-align-evals-data/tree/97446d992c3785d6605f1500b2c9b95d042e7b9c/human_eval https://github.com/openai/code-align-evals-data/tree/97446d9... [1] https://arxiv.org/abs/2211.15533 https://arxiv.org/abs/2211.15533
- godelski 3y agoMy favorite line from the HumanEval paper[0] > It is important for these tasks to be hand-written, since our models are trained on a large fraction of GitHub, which already contains solutions to problems from a variety of sources. So to answer your question, yes, the evaluation dataset is spoiled. You can find such unique and never before seen docstrings like > For a given list of input numbers calculate the Mean Absolute Deviation around the mean of this dataset. Mean Absolute Deviation is the absolute difference between each element and a centerpoint (mean in this case)[1] And here's a repo I found that is 8 years old[2]. But how about a more recent one that is even closer?[3] There's plenty more examples[4] (does anyone know how actually limit the date to prior to 2021? `pushed:<2021` doesn't work nor does using the `created` keyword. Date searching doesn't seem to work well). In essence, we can still use this evaluation method to determine how good our model is at doing fuzzy searching. Which, mind you, is still a useful thing. But I would be careful in concluding that this means the model is good at generalizing arbitrary descriptions of code or novel pieces of code. That said, one may be able to argue that not many lines of code are actually that novel. Still, we need to be careful about our conclusions and understand the limitations of our metrics (something I am currently deeply troubled by) [0] https://arxiv.org/abs/2107.03374 https://arxiv.org/abs/2107.03374 [1] https://github.com/openai/code-align-evals-data/blob/97446d992c3785d6605f1500b2c9b95d042e7b9c/human_eval/floats_mean_absolute_deviation.py#L6 https://github.com/openai/code-align-evals-data/blob/97446d9... [2] https://github.com/bertomartin/stat4701/blob/ec2b64f629cbbf6267169302265f73a98edef67d/stck.py#L175 https://github.com/bertomartin/stat4701/blob/ec2b64f629cbbf6... [3] https://github.com/danielwatson6/hate-speech-project/blob/64a2eecce5218373ef5c449eeb1dfb397532eda5/scripts/wordnet_mf.py#L61 https://github.com/danielwatson6/hate-speech-project/blob/64... [4] https://github.com/search?q=abs%28x+-+mean%29+for+language%3APython&type=code https://github.com/search?q=abs%28x+-+mean%29+for+language%3...
- newhouseb 3y agoFirst - thank you for open sourcing this! It's a real gift to the community to have a model intended for "commercial use" that's actually licensed as such. I'd be very interested to hear about the choice/evaluation of the ALiBi approach for positional embedding (perhaps in the technical report). My intuition suggests that while this allows for better generalizability for longer sequence lengths, it penalizes scenarios where an LLM might need to check for things like a function signature far away from where the next token is generated. My initial testing of this model tracks with this intuition but that's by no means a rigorous evaluation.
- ofirpress 3y ago(I wrote ALiBi) You can read the paper here https://arxiv.org/abs/2108.12409 https://arxiv.org/abs/2108.12409 While intuitively it does seem like ALiBi would make it hard for the model to attend to things that are far away, in many scenarios we've tested with different models trained on different datasets, ALiBi always performs better than sinusoidal, rotary, and other embedding types, even when we're not using it to extrapolate to longer sequence lengths. These findings have been confirmed by others, including by the BLOOM open source LM project.
- newhouseb 3y agoSmall world! Thanks for the link (which I've now skimmed beyond the abstract). What wasn't obvious to me from the abstract is that different attention heads have different penalty strengths, so if some prediction task requires long range dependencies you might expect one of the less-penalized heads to end up specializing. I wonder what would happen if the penalty for one head is zero? (The paper suggests this might've been tried and just made things worse, but unclear) I must admit that this is a wonderfully elegant (and interpretable) way to do this... much more intuitive (to me at least, a wannabe practitioner) than all of the trig-based embeddings.
- ofirpress 3y ago> so if some prediction task requires long range dependencies you might expect one of the less-penalized heads to end up specializing Exactly. You have heads that focus on content nearby and ones that focus on stuff that is far away. > I wonder what would happen if the penalty for one head is zero? (The paper suggests this might've been tried and just made things worse, but unclear) Yup, this is something we tried. Making one of the heads zero doesn't improve or degrade performance. >I must admit that this is a wonderfully elegant (and interpretable) way to do this... much more intuitive (to me at least, a wannabe practitioner) than all of the trig-based embeddings. Thanks so much!!
- sputknick 3y agoWhat does "fine tuning" mean in this context? Does it mean you fine-tuned it on a specific code repository, or collection of code repositories and then had it do work in those repositories?
- amasad 3y agoBroadly finetuning is any post pretraining training. Most of the time it is used in the context of fitting a more narrow task. In our case, it was the same training objective as the pretraining but meant to be more representative of what Replit users like to code. However, we were surprised by how well it boosted overall performance. Best guess: it's a) novel data and b) the model could take even more training!!
- spenczar5 3y agoHow feasible and effective would it be to fine-tune a model against an organization's private source code, resulting in an "internal" model that knows how to work with that org's stuff? Could you, say, fine-tune the model every week with the latest merges? Every hour?
- pyth0 3y agoFinetuning is a relatively quick process. Training the base model is the expensive part (can take weeks and huge amounts of compute), whereas finetuning usually is only on the last few layers and can be done with much less resources. You could definitely have a "nightly" finetune model that is retrained every day or so.
- rattray 3y agoInteresting - how would that work for a company that wanted to run their own codex model, on-prem, trained on their own code? Perhaps also trained on their dependencies?
- naderkhalil 3y agoFinetuning a smaller model leading to better performance seems like a significant finding that'll lead to a lot of companies fine-tuning their own internal "ChatGPT"s
- spenczar5 3y agoHow is this code licensed? I didn't see a license in the repo. It looks interesting!
- dgacmu 3y agoThe README indicates: The base model checkpoint is licensed under the Creative Commons license (CC BY-SA-4.0). Under the license, you must give credit to Replit, provide a link to the license, and indicate if changes were made. You may do so in any reasonable manner, but not in any way that suggests that Replit endorses you or your use.
- kir-gadjello 3y agoImpressive model, thank you for releasing it under a business-friendly license! Have you considered using Google's sparse "scaling transformer" architecture as the base? Even at 3B scale it can generate 3-4x more tokens per FLOP while being competitive at perplexity with a dense transformer. I think OpenAI uses a variant of it in their ChatGPT-3.5-Turbo product. Here is the paper https://arxiv.org/abs/2111.12763 https://arxiv.org/abs/2111.12763 and the implementation https://github.com/google/trax/blob/master/trax/models/research/terraformer.py https://github.com/google/trax/blob/master/trax/models/resea... if you are interested. Hope you get to look into this!
- b33j0r 3y agoThank you for releasing the weights along with the announcement. The posts that made great headlines, but “weights are on their way!” Like why did we even get excited? This? Great work.
- chaxor 3y agoI don't think it's a business friendly license?
- kir-gadjello 3y agoIt allows for modifications and commercial use: https://creativecommons.org/licenses/by-sa/4.0/ https://creativecommons.org/licenses/by-sa/4.0/ >You are free to: >Share — copy and redistribute the material in any medium or format >Adapt — remix, transform, and build upon the material >for any purpose, even commercially. Compare this to the latest release from StabilityAI lab DeepFloyd, "IF", which in addition to various restrictive clauses strictly prohibits commercial use: https://github.com/deep-floyd/IF/blob/develop/LICENSE-MODEL https://github.com/deep-floyd/IF/blob/develop/LICENSE-MODEL Repl.it's release is as open as it gets these days, in my book.
- LukeShu 3y agoIt's a copyleft license; and lots of folks on HN seem to think that copyleft, while being open, isn't business friendly.
- curiousgal 3y agoDid any interns help in developing this? If so are you planning on intimidating them as usual? :) Reference: How Replit used legal threats to kill my open-source project https://intuitiveexplanations.com/tech/replit/ https://intuitiveexplanations.com/tech/replit/
- robertlagrant 3y agoWow. That's extremely poor behaviour if the account is accurate.
- curiousgal 3y agoOh it is. 4000+ upvotes on HN. https://news.ycombinator.com/item?id=27424195 https://news.ycombinator.com/item?id=27424195