3 ms·
If only LLM benchmarks could benchmark it in the first day! Still no Artificial Analysis benchmark yet. Or benchmark for Laguna S 2.1 or Meituan models or lots
by jychang 2mo ago
If only LLM benchmarks could benchmark it in the first day!
Still no Artificial Analysis benchmark yet. Or benchmark for Laguna S 2.1 or Meituan models or lots of other models.
- eli 2mo agoAccording to a few tasks from my little personal coding benchmark it's very good at coding and kinda bad at web design. (Also excellent at "draw me a picture" one-shot prompts, for whatever that's worth) On a sneaky one that involved parsing MIME headers and dealing with character encodings it did better than Kimi K3 at Max and for 38% lower cost. Interestingly it seems noticeably better than the qwen3.8-max-preview model they offered just a few weeks ago.