9 ms·
IBM Granite: A Family of Open Foundation Models for Code Intelligence
- koryk 2y agoI am seeing at least one granite model on ollama, wonder when they will all show up!
- victor9000 2y agoDoes anyone know of other open models available for code intelligence?
- continuational 2y agoIs there an online demo of this somewhere?
- hustwindmaple1 2y agoI wonder why companies like IBM are jumping on the LLM bandwagon and training/releasing models that have no chance of competing with Llama/Mistral? To me it just looks like a complete waste of $$ because nobody will use them in any serious scenarios
- paxys 2y agoIBM made $60 billion in revenue last year. Where do you think it all came from? The same companies/governments that buy their overpriced crap are going to buy these new LLMs as well.
- breezeTrowel 2y agoThese are open weight models released under an Apache 2.0 license. There's nothing to buy.
- kkielhofner 2y agoIBM is a sales and services org. Their customers aren’t going to build their own RAG and agent frameworks, vector DBs, data ingest pipelines, finetunes, high scale inference serving solutions, etc, etc. There’s an incredible amount of stuff to buy.
- hustwindmaple1 2y agoRight, but they can just use Llama/Mistral for free, instead of their inferior models, which I'm sure take quite a bit of resources to train in the first place.
- kkielhofner 2y agoYes but using someone else's models doesn't make them an "AI company".
- paxys 2y agoWho is going to host them?
- Brajeshwar 2y agoWhen they pitch potential clients for their services, their slides on LLM, AI, ML, etc., must be their own. Whether they use it or not for the services does not matter. These are like the side projects that service companies release to help them close their clients.
- adt 2y agohttps://lifearchitect.ai/models-table/ https://lifearchitect.ai/models-table/
- khana 2y ago[dead]
- reacharavindh 2y agohttps://i.kym-cdn.com/photos/images/original/001/138/631/b7a.png https://i.kym-cdn.com/photos/images/original/001/138/631/b7a...
- holografix 2y agoIs this a segway for IBM to release Terraform specific LLMs so I never have to write that hot garbage ever again? Sign me up IBM!
- nwsm 2y agoHere's a similar existing product- https://www.ibm.com/products/watsonx-code-assistant-ansible-lightspeed https://www.ibm.com/products/watsonx-code-assistant-ansible-...
- snapcaster 2y agojust a heads up it's segue in the context you're using it
- holografix 2y agothank you internet stranger!
- sbierwagen 2y ago3B, 8B, 20B and 34B parameter model weights available here: https://huggingface.co/collections/ibm-granite/granite-code-models-6624c5cec322e4c148c8b330 https://huggingface.co/collections/ibm-granite/granite-code-...
- dur-randir 2y agoBased on their own numbers, 8B seems decent, but 34B not worth it compared to general-purpose trained models even on specific tasks. Which is an interesting result.
- dusanh 2y agoI'm a complete newb when it comes to AI, and I am getting pretty ashamed of it too. How do I take a model like this and use it in my day to day? Can I somehow use in, say, VSCode? How do I point it at my code base, and use it to help me write new code?
- everforward 2y agoYou run most of these models in something that wraps them in an HTTP API. I use Ollama, which I think is the most popular but I’m not in a great position to judge. My impression is that it handles running models on CPU better. So you’d basically install Ollama, download one of the versions of this model off HuggingFace, create a Modelfile since this isn’t in the default Ollama repo, and then Ollama can answer prompts with the model. Modelfiles are very simple, based on Dockerfiles. It takes like 15 seconds to make one if you aren’t messing with the various parameters. Once it’s in Ollama, just get one of the various GPT plugins for VSCode and give it the Ollama URL (http://localhost:11434 http://localhost:11434 by default). I use continue.dev but there are many. Continue takes over the tab autocomplete with the LLM, and has a chat window on the right where you can use keyboard shortcuts to copy code into the prompt and ask it to edit/generate code or ask questions about existing code.
- dusanh 2y agoThank you so much! That sounds surprisingly straightforward. I expected a lot more fiddling to get going. Where would I start if I wanted to use a model programmatically ? Like let's say I am building a chat bot. I have a large data set of replies I want the model to mimic, and I'd want to do this in Python. Of course, I'd probably use a different model than Granite.
- everforward 2y agoThis is stretching my own knowledge, so if someone else knowledgeable wants to take a stab here I would appreciate a response as well! Before doing that, I would start basic. Pull llama3 and see what it does with your prompts. You may be surprised how much is already in there and just not need to involve your own data at all. If that doesn’t work, check HuggingFace to see if someone has already made a model/fine tune/LoRA for what you’re trying to do. There are many, eg I found a Magic The Gathering rules model the other day. If those fails, or you just want to play with your own data, you’ll need to figure out what “mimic” means. If the model does okay with generating content but the content is factually wrong or missing background, you may be able to just do RAG (retrieval augmented generation). Basically running your documents through an AI that converts them to embeddings (some kind of vector, I don’t understand how they work). Then when you run a query, you can search for related embeddings and pass them to the model so that it “knows” the content that was in the document. This is the easiest; open-webui (the Ollama web chat interface) has some RAG support. Danswer is open source and built from the ground up to do RAG, and has built in support for ingesting from Slack, Drive, etc, etc. OpenAI also has embedding as a service. A step up from that is making a LoRA. To my novice eyes, LoRA’s are basically a diff of the models parameters or weights. So rather than training a whole new model, you just add deltas to an existing one. These let you “teach” the model something while preserving the base generation capabilities of the underlying model. Ie you won’t have to worry about making sure you feed it enough data that it can speak English properly, because it gets that from the base model, you only have to give it enough data to speak about whatever you’re training it on. If that doesn’t make any sense, go check CivitAI for Stable Diffusion (image model) LoRAs. The effects are way more obvious on image AIs. Anyways, LoRAs are trained so you’re into training there. I think HuggingFace has tools that make this easy, but I don’t know enough to say anything with confidence. The last option, which you almost certainly don’t want, is to train a new base model like llama3. You’re starting from 0 there; you have no existing model so you will have to teach it everything. It will take a ton of data, it will take forever to train, and it will likely be much worse than even randomly clicking models on HuggingFace. Meta has spent who knows how much on Llama and it still hallucinates. If you end up training, you’ll probably end up doing it in the cloud unless you have tons of VRAM doing nothing. Prices are pretty reasonable, I think A100s are around $2/hr. I don’t know how to gauge how long it needs to train, but I believe it’s related to the amount of data you’re training on. I believe it’s pretty reasonable for LoRAs though, I’m guesstimating in the $20-ish range? Edit: oh, and I’m not affiliated in any way, but I found out last night that Fireworks’ new function calling model is free while it’s in beta, which is a neat/fun thing to play with. https://fireworks.ai/blog/firefunction-v1-gpt-4-level-function-calling https://fireworks.ai/blog/firefunction-v1-gpt-4-level-functi... it’s also open weights if you want to run it locally, but it’s a 40B model so I can’t on my 3060
- throwaway290 2y agoAs usual, license/copyright violation: > Our process to prepare code pretraining data involves several stages. First, we collect a combination of publicly available datasets (e.g., GitHub Code Clean, Starcoder data), public code repositories, and issues from GitHub