4 ms·
GPT-4o has seen a lot of examples of what people write about coding on the web, what code exists, and what tasks people want to do with that. But that general d
by johntb86 2y ago
GPT-4o has seen a lot of examples of what people write about coding on the web, what code exists, and what tasks people want to do with that. But that general data doesn't include full process of coding - get a bug report, look at a codebase beforehand, make changes to the codebase, test them and see exactly what bugs they have, and iterate. That process is what SWE-bench tests.
It's possible OpenAI did some coding fine-tuning themselves; Meta's Llama 3 paper [0], section 4.3.1 mentions what sort of work is needed. However, anything OpenAI did is based on their own tooling and set of assumptions - e.g. how is the existing code input into the LLM, what set of actions can the LLM take (e.g. look up documentation), what language is the output code being written in, etc. Cosine's LLM framework may do things differently and have different features, so you'd need to fine-tune the LLM to take maximum advantage of the framework.
It's like dropping the LLM down in front of Vim when it had only ever used or even heard of notepad (or even emacs); there needs to be some training to make it work well with the new tools it has.
[0]: https://arxiv.org/pdf/2407.21783 https://arxiv.org/pdf/2407.21783