Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
m-dot-reviews
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
m-dot-reviews
3mo ago
I've been plugging this perhaps too many times now, but I am trying to bootstrap a user-sourced corpus of exactly "what model is good at task X". So, not benchmarks, but high-level tasks. There's a bit of a ordering prob
2.
▲
by
m-dot-reviews
4mo ago
Oops, thanks for telling me that. I think the issue should be fixed now.
3.
▲
by
m-dot-reviews
4mo ago
So, this may not be precisely what you're looking for but it may come close. I've put together a simple site for sharing ratings/opinions on models on a task-specific granularity. https://model.reviews/ The i
4.
▲
by
m-dot-reviews
4mo ago
For anyone who's interested, I've put together a simple site for sharing ratings/opinions on models at a task-specific granularity. https://model.reviews/ The idea is that benchmark score comparisons are usef
5.
▲
by
m-dot-reviews
4mo ago
I'm starting a repository of LLM reviews [1] with the goal of creating a catalog that is more task-oriented and less marketing-y than corporate blogs or benchmark leaderboards. You seem to have a lot of experience across a bunch of dif
6.
▲
by
m-dot-reviews
4mo ago
Anecdotally, yes there is definitely a difference. Even e.g. Haiku (cheapest Anthropic model) vs gpt-oss-120b had a big difference in quality and syntax issues when I was testing them for DSL generation. Granted, that's a little differ
7.
▲
by
m-dot-reviews
4mo ago
I looked for a forum like this a few months ago during my own model research, and didn't find one. So, here's the "catalog of clankers," a task-structured review site for LLMs. The idea is that if you're looking for