Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ed-is-ai
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
ed-is-ai
28d ago
Pretty nice pelicans!
2.
▲
by
ed-is-ai
1mo ago
ok first thoughts - this is a cool, but ludicrously expensive toothbrush! Anyone who buys one is just flaunt their wealth when $2 gets you a manual version
3.
▲
by
ed-is-ai
1mo ago
The big issue I have with Fable is this. From the Anthropic email announcing Fable 5.1. So basically they're giving us a Ferrari, which will point blank refuse to do certain stuff - forcing us to go out in our Mustang. Their choice, n
4.
▲
by
ed-is-ai
1mo ago
The point I'm making is that most models are good enough for most tasks, so choose on speed/cost. Definitely benchmark on your own tasks. GLM-5.3 is the winner on mine, on yours maybe not. I am not trying to be a universal benchm
5.
▲
by
ed-is-ai
1mo ago
Noted. Feedback received
6.
▲
by
ed-is-ai
1mo ago
Indeed; I find it really interesting there is so much disbelief and outright hostility to my post, in these comments. I am an AI Realists rather than AI hype-artist, say it as it is. For me, the evidence is clear that for our own use cases
7.
▲
by
ed-is-ai
1mo ago
I had to google astroturfing. I liked the due diligence... I am a hacker news infant, so all my history is about this. Give it a proper look. The results all there, and code if you want to run the evals (or make your own - it's very e
8.
▲
by
ed-is-ai
1mo ago
Impolite
9.
▲
by
ed-is-ai
1mo ago
It's good to check the facts. Time will tell, I'm not the only one saying these things. But you look at the raw data and test to your hearts content if you care enough - all results and the code use to run it is above. And you c
10.
▲
by
ed-is-ai
1mo ago
Thanks for point out - I updated the benchmark today to include. It is every bit as good as everyone says it is. The intelligence / $ is something else...
11.
▲
by
ed-is-ai
1mo ago
That's pretty cool. All available already in OpenCode if you're a developer, see for yourself!
12.
▲
by
ed-is-ai
1mo ago
Thanks for the feedback. It was limited to 7 to keep things balanced. On the premise most people don't just code. The 7 were an example of my realworld use cases - the point of this is to encourage people to run their own benchmarks
13.
▲
by
ed-is-ai
1mo ago
https://github.com/ed-is-ai/featherbench
14.
▲
by
ed-is-ai
1mo ago
pcwelder - https://github.com/ed-is-ai/featherbench all the data is here if you care to look But thanks for the feedback
15.
▲
by
ed-is-ai
1mo ago
Hmmm, as the author: the feedback is interesting indeed and something to take onboard... Ultimately I hope the result provided value by stimulating debate. To me it feels like lots of people rejecting the idea that a Chinese open weights m
16.
▲
by
ed-is-ai
1mo ago
Read the benchmark - it's on the basis of being 'good enough' for everyday tasks. Which is what fits most applications right. This is not a benchmark for testing them against an Einstein.
17.
▲
GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
(reinvently.co.uk)
240 points
by
ed-is-ai
1mo ago
|
124 comments
18.
▲
by
ed-is-ai
2mo ago
It looks like you're mostly 1-shotting apps and reviewing the results based on # interventions etc Like it!!
19.
▲
by
ed-is-ai
2mo ago
neat! I will check it out. I'm not really using local LLMs as dont have a big enough machine for these to run with a decent tps
20.
▲
by
ed-is-ai
2mo ago
Author here. This began as a weekend project because I didn't trust the public leaderboards — as what I was experiencing seemed to differ from the professional benchmarkers. I really wanted to see how the latest LLMs worked on 'my
21.
▲
Building the Ed-O-Meter: Notes on Writing My Own LLM Benchmark
(reinvently.co.uk)
1 points
by
ed-is-ai
2mo ago
|
4 comments