3 ms·
What are those “tasks in your pipeline”? We dedicated an entire team of 6 for two months on evaluating LLMs and while Claude 2 was the final choice, we found L
by 19h 3y ago
What are those “tasks in your pipeline”?
We dedicated an entire team of 6 for two months on evaluating LLMs and while Claude 2 was the final choice, we found Llama 2 70b to be absolutely great (for summaries and structured data generation from free text).
We chose Claude because running Llama on an A100 or H100 comes with a baseline cost that doesn’t go to zero when you don’t need it (you could spawn a new instance but right now GPUs are so rare everywhere except for expensive cloud providers that it’s possible you don’t get one).
That said, we found the smaller Llama models to be so hilariously bad we have an internal slack channel where Llama 7b writes jokes and the funny part isn’t the jokes but how utterly stupid and random they are.