Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
XCSme
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
181.
▲
by
XCSme
3mo ago
Only supporting "max" reasoning is weird, their parameters are quite inflexible atm: Important limits: reasoning_effort currently supports only max; K3 always has thinking mode enabled. max_completion_tokens defaul
182.
▲
by
XCSme
3mo ago
No blog post? Benchmarks?
183.
▲
by
XCSme
3mo ago
Manually select a rectangle on a video frame, then do basic computer vision to detect notes, or even a simple image processing algorithm to find the lines and notes.
184.
▲
by
XCSme
3mo ago
I was excited to use it, as I really wanted something like this, then I realized it needs AI/Claude (?) That sounds like it can get quite costly. Probably there are ways to do it without AI, I would rather manually annotate the tab are
185.
▲
by
XCSme
3mo ago
Improving my free/hobby AI Benchmarks website: https://aibenchy.com And, as always, working on my main self-hosted analytics platform: https://www.uxwizz.com/
186.
▲
by
XCSme
3mo ago
With subscriptions, you want to have ways to increase the subscription amount and retain people, which usually leads to adding features no one asks for and bloating the product, trying to upsell users.
187.
▲
by
XCSme
3mo ago
And in my tests, that point of "overthinking" depends on the problem's complexity, so it's not necessarily that using "xhigh" is always bad or good.
188.
▲
by
XCSme
3mo ago
One example where the order seems correct, is this SVG generation test: https://aibenchy.com/showcase/?q=Gemini+3.5%2Cgpt+5.6%2C+5.3... You can see that most Gemini 3.5 generations are more correct than 5.6 Sol (the ne
189.
▲
by
XCSme
3mo ago
It's because the benchmark is not coding-only. Gemini models tend to have most knowledge for most domains, and are one of the most intelligent overall. You can check other benchmarks too, on specific categories, those models still beat
190.
▲
by
XCSme
3mo ago
In my tests, in almost all cases, using Sol on (low) reasoning is the best option intelligence/price-wise. Luna is good too, for classification tasks or any pre-processing task that is not critical
191.
▲
by
XCSme
3mo ago
Also for most, there doesn't seem to be a big difference between (medium) and (high).
192.
▲
by
XCSme
3mo ago
Yeah, for some reason the (low) versions do really well, like they think directly of the solution instead of going around all the edge-cases and getting lost in one of them.
193.
▲
by
XCSme
3mo ago
Here's all 3 (medium), and GPT-5.5 It GPT-5.6 doesn't seem to be a lot smarter than 5.5, but it is faster, cheaper, more efficient and more consistent: https://aibenchy.com/compare/openai-gpt-5-6-sol-medium&#x
194.
▲
by
XCSme
3mo ago
Not always, in some cases, changing to a higher reasoning makes the AI doubt itself too much, and skip over the correct answer by overcomplicating the problem and polluting the context. It would be nice to see on which categories of problem
195.
▲
by
XCSme
3mo ago
GPT-5.6 is a really good model, and quite cheap. I can finally replace GPT-5.3-Codex for my Tool Calling in n8n. Here's my benchmark results for GPT-5.6: https://aibenchy.com/?q=gpt-5.6 (the high reasoning variants are
196.
▲
by
XCSme
3mo ago
What would the advantage be?
197.
▲
by
XCSme
3mo ago
Hamsters are also getting better, but still quite off compared to SOTA models: https://aibenchy.com/showcase/?q=grok
198.
▲
by
XCSme
3mo ago
I am trying to benchmark it now, but: - It doesn't seem available in EU (?) - Using a VPN seems to sort of fix it, but it's way slower than I expected, when everyone was praising it, it feels like the speed is slowly ram
199.
▲
by
XCSme
3mo ago
I often use Gemini as my "chat" app to ask questions, etc. I stopped using ChatGPT because of they're weird login system, where it keeps switching to my Workspace Codex account, which doesn't actually have the free/
200.
▲
by
XCSme
3mo ago
I spent 30mins debugging why my Github Pages were serving old versions... It has been down for at least 2 hours, with actions not being executed. I can not finalize my deployment because of this outage, so now I have to delay my table-tenni
201.
▲
by
XCSme
3mo ago
It is on top for many benchmarks, only not the coding/agentic ones. Still one of the most intelligent models overall, most likely to get any question you ask correctly (without tools).
202.
▲
by
XCSme
3mo ago
What's interesting, is that Sonnet 5 is actually worse[0] than 4.6 without reasoning. It makes some sense, as models are trained more and more with reasoning, than without. [0]: https://aibenchy.com/compare/anthrop
203.
▲
by
XCSme
3mo ago
Well, it is a Sonnet model, it is indeed better[0] than Sonnet 4.6 (smarter, faster, cheaper), but I don't see why would you use it as opposed to Opus 4.8 low or GLM-5.2... [0]: https://aibenchy.com/compare/anthrop
204.
▲
by
XCSme
3mo ago
As always, note: faster than GLM-5.2 doesn't mean too much, as GLM-5.2 is served by different providers, so the inference speed can vary drastically between providers or over time.
205.
▲
by
XCSme
3mo ago
I just tested it on my benchmarks[0], it's GLM-5.2 level, at 2x cost, but also 2x faster. Weak spots (categories it fails): - Trivia — 0/3 - basically not much built-in knowledge - Combined tool-calling tasks — score 45&
206.
▲
by
XCSme
3mo ago
Considering the cloud version, all three models compared in the article (Qwen 3.6 35BA3b, 3.6 27B and DeepSeek V4 Flash), have very similar performance[0], BUT on cloud, for some reason DeepSeek V4 Flash is 10-20x cheaper than the Qwen mode
207.
▲
by
XCSme
3mo ago
Another example was Gemini 3.1 flash lite, which on high was basically just burning tokens, costing like 30x more, while giving worse answers: https://aibenchy.com/compare/google-gemini-3-1-flash-lite-hi...
208.
▲
by
XCSme
3mo ago
You'd be surprised, some models on high do worse than on medium, because they start overthinking and doubting themselves, polluting the context with too much information, etc. It depends a lot on the task and harness too (using plans a
209.
▲
by
XCSme
3mo ago
Note that being open-weights, "slower" is relative, as it depends on who's serving the model. This can drastically change over time too.
210.
▲
by
XCSme
3mo ago
Does a bit worse than Opus 4.8 in my tests[0], but it's 5x cheaper and 3x slower. [0]: https://aibenchy.com/compare/anthropic-claude-opus-4-8-mediu...
More ›