Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
XCSme
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
151.
▲
by
XCSme
3mo ago
Is a solution to protect against this to go almost fully offline? To have your own local models, local software, local everything and just allow very strict and well-protected pathways into external network traffic. I am thinking of having
152.
▲
by
XCSme
3mo ago
I agree. I (15+ years of coding) feel like I'm 100x more productive now using LLMs to code. But I don't think a beginner would have the same experience. The AI still makes A LOT of stupid mistakes and decisions, but I catch them e
153.
▲
by
XCSme
3mo ago
Not only performance optimizations, but also UI/UX. If you ask to implement a custom dropdown that does something, it often comes with good spacing, aria-accessible tags, keyboard accessibility, etc. A junior dev wouldn't think of
154.
▲
by
XCSme
3mo ago
In my experience, when adding new features with LLMs, most of those optimizations come automatically. Good models now already follow best practices when implementing, better than junior devs. I wrote a bit about this, I call it "AI sla
155.
▲
by
XCSme
3mo ago
I think sometimes optimizations are more about using the right data structure/ideas/libraries. In my case, I recently optimized for https://uxwizz.com the session playback: before, it was saving the entire recording in
156.
▲
by
XCSme
3mo ago
Congrats, I love performance optimizations! Hardware nowadays is so powerful, but our code so inefficient... I think most libraries/apps could easily be 10x-100x faster if we really try to optimize them. The good thing is, that now wit
157.
▲
by
XCSme
3mo ago
Not for iOS, but for web apps, if you like self-hosted analytics that has more than basic stats, I'm really proud of what I've built with https://uxwizz.com
158.
▲
by
XCSme
3mo ago
The linked website shows this for me: 451: Unavailable due to legal reasons We recognize you are attempting to access this website from a country belonging to the European Economic Area (EEA) including the EU which enforces the General Data
159.
▲
by
XCSme
3mo ago
Thanks for the feedback, really good points! There is some short info about the methodology here: https://aibenchy.com/methodology/ > GPT-5.6 Sol on Low beats Fable Medium by 10% > Fable is number 20 Fable loses
160.
▲
by
XCSme
3mo ago
They heavily boosted the amount of users in the past months (I myself got a business account because they had $1k free credits, and a personal one because of all the resets they give). Now they are trying to quickly boost revenue with ads.
161.
▲
by
XCSme
3mo ago
And two years before it's just silently plugging the product. And it will be considered "increased brand exposure", not ads...
162.
▲
by
XCSme
3mo ago
Interesting, that was expected to happen at some point. I didn't see any mention of prices, click rates, conversion rates compared to other type of ads, like search ads. I assume it's going to be bidding based too?
163.
▲
by
XCSme
3mo ago
I laughed, she laughed, the toaster laughed...
164.
▲
by
XCSme
3mo ago
I don't like to divulge tests, but one of them is a chess puzzle. > would be great if you could add GPT 5.6 Sol XHigh and Fable 5 High as well I would like too, but I avoided them for several reasons: 1) Cost - this is a hobby proje
165.
▲
by
XCSme
3mo ago
Oh, and I've also added weights to different categories, so Coding and Tool usage categories influence the score more. This done both to better account for how most people are being used, and also to reduce Gemini's dominance in g
166.
▲
by
XCSme
3mo ago
I have created various questions/tests and put the models through the same tests. I record whether the answers are correct, and the generation stats (costs, latencies, tokens used, etc.). I have no idea why the Gemini models do so well
167.
▲
by
XCSme
3mo ago
Here, my comparison of 3.6 Flash vs Sol vs Luna vs Terra: https://aibenchy.com/compare/google-gemini-3-6-flash-medium/...
168.
▲
by
XCSme
3mo ago
In my tests, 3.6 Flash is NOT more token efficient, so it actually ends up costing more than 3.5 Flash, even with the output price reduction. EDIT: It less less verbose in final output though, but it reasons more. I assume the optimization
169.
▲
by
XCSme
3mo ago
tl;dr: 3.6 flash is a bit smarter than 3.5 flash, but also a bit more expensive. My results [0] put Gemini 3.6 Flash at the top. 3.6 Flash high has same $1.5 input price as 3.5 Flash, but output is cheaper from $9.0 to $7.5. Google said 3.6
170.
▲
by
XCSme
3mo ago
I was expecting 3.6 Pro. It's been so long since the last Pro model...
171.
▲
by
XCSme
3mo ago
Woaah, zooming all the way out is really cool
172.
▲
by
XCSme
3mo ago
Gsplat + dlss
173.
▲
by
XCSme
3mo ago
I was not referring to the input/output price, but the cost of doing a specific tasks, in practice it is ~10x cheaper than GLM-5.2 for example, to accomplish the same task (for the tasks it can do). I have been happily using DeepSeek V
174.
▲
by
XCSme
3mo ago
I use it through OpenRouter via Kilo Code VS Code extension. You can check the logs in OpenRouter and see which providers it used and how many tokens you used.
175.
▲
by
XCSme
3mo ago
I am not even sure if Fable is as smart as they say, I can't get it to answer almost any question, it always refuses for "cyber-security" concerns...
176.
▲
by
XCSme
3mo ago
DeepSeek V4 pricing is insane, 10x-30x cheaper to use than most other models, and it usually is good enough for most tasks.
177.
▲
by
XCSme
3mo ago
Benchmarks look ok, but they don't mention anything about the issue with the model being extremely slow and verbose. That being said, it's awesome to have such an open-source model, even if now it's unusable mostly locally, w
178.
▲
by
XCSme
3mo ago
Just saw the logs, coding demos failed due to the 5 minute/task timeout. I have increased it and retesting it now. EDIT: With 10 minutes timeout, the CSS task completed, but the SVG generation task still timed out. Trying again with 30
179.
▲
by
XCSme
3mo ago
I finished benchmarking[0] it, but it was not fun, it only supports (max) reasoning and the model is quite slow. Apart from a few requests timing out, it also has some issues with tool calling/response format schemas (Moonshot rejected
180.
▲
by
XCSme
3mo ago
I am trying to benchmark it, but it only supports (max) reasoning, and even for simple questions, it takes forever to answer/times out :(
More ›