4 ms·
> The most promising model, Meta’s open source model Llama2-70B This is an old model that was not dominant even when released. This study must be fairly old o
by brrrrrm 2y ago
> The most promising model, Meta’s open source model Llama2-70B
This is an old model that was not dominant even when released. This study must be fairly old or I question the qualification of the group running it.
- staticman2 2y agoThey tested it in February 2024. There are presumably privacy reasons limiting them to open models. It probably was the most promising at the time for their use case.
- Etheryte 2y agoIt was released roughly a year ago, I wouldn't really say this is an issue. While there has been plenty of progress in the meanwhile, actually conducting research takes time and effort, you can't expect everything to be done on the very bleeding edge at all times.
- brrrrrm 2y ago> actually conducting research 5 people evaluating 45 responses? This doesn't take a year to do. The issue is that this study is was poorly funded and slow - model development has far outpaced the results and there's likely little to no longevity in the outcome.
- a2128 2y agoAccording to the linked PDF which includes a report from AWS professional services, the PoC was originally devised in September 2023 and conducted in January/February of 2024, before the release of Llama 3. They tested Llama2-70B, Mistral-7b and MistralLite. They didn't evaluate proprietary models such as GPT4 probably because they would've wanted to be able to deploy it within Australia or with an Australian company
- brrrrrm 2y agoperhaps the title should be "Llama2-70B worse than humans ..."