Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
seunosewa
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
seunosewa
14d ago
It's misalignment, which they created during the model's training.
2.
▲
by
seunosewa
14d ago
There is no excuse for the second bomb.
3.
▲
by
seunosewa
19d ago
I believe the AI labs are weakly motivated to train strongly against cheating when it helps with benchmarks.
4.
▲
by
seunosewa
19d ago
It will get into the hands of people who just want to burn the world down.
5.
▲
by
seunosewa
29d ago
Meta is awfully close.
6.
▲
by
seunosewa
1mo ago
In my research it's a solid 10 - 20% shorter on average, free of charge. That's nice to stack on top of other optimisations.
7.
▲
by
seunosewa
1mo ago
It's not a general trend. It's only Opus 5.
8.
▲
by
seunosewa
1mo ago
They are cute but nasty. They eat their prey alive in the most painful way.
9.
▲
by
seunosewa
1mo ago
From the benchmarks, the single core performance is close but lower.
10.
▲
by
seunosewa
1mo ago
Unfortunately?
11.
▲
by
seunosewa
2mo ago
$100 Claude and $100 ChatGPT Pro is the best value for $200/month. They can see each others' mistakes.
12.
▲
by
seunosewa
2mo ago
Could you provide a practical example?
13.
▲
by
seunosewa
2mo ago
Last time I tried it, Flash introduced too many errors due to sloppiness. Is it more reliable at following instructions now?
14.
▲
by
seunosewa
2mo ago
They want to maintain the perception that Flash is worth $7.5/mot, so they can charge more for the next one.
15.
▲
by
seunosewa
2mo ago
Two thoughts: 1) It makes sense to try 3.7 flash before cancelling. 2) Prompting models to be honest is surprisingly effective in my recent experience. But only if they listen to instructions.
16.
▲
by
seunosewa
2mo ago
Could you try setting the temperature very low e.g. 0.0?
17.
▲
by
seunosewa
2mo ago
You can set it up to always use the official provider.
18.
▲
by
seunosewa
2mo ago
Do it a second time at least.
19.
▲
by
seunosewa
2mo ago
The SSH into machine market is huge. Cloud work, Dedicated servers, VPS, Colab.
20.
▲
by
seunosewa
2mo ago
GPT 5.6 Sol is so free of drama for me for brainstorming, writing prompts, etc. Tasks I always did with Opus until 4.7.
21.
▲
by
seunosewa
2mo ago
Provide an option. An autonomous mode or an interactive mode.
22.
▲
by
seunosewa
2mo ago
The problem is it should ask you before doing these things. This is a cherry picked example. There are times when they go on misadventures.
23.
▲
by
seunosewa
3mo ago
They are trying to make money. That's what firms in any capitalistic economy care about the most. Regardless of the government's presumed interference, the companies themselves are all trying to make money. All competing for subsc
24.
▲
by
seunosewa
3mo ago
How does Fable compare?
25.
▲
by
seunosewa
3mo ago
The input price is quite high. That's what gets you.
26.
▲
by
seunosewa
3mo ago
Not true. The only issue is cost of frontier models.
27.
▲
by
seunosewa
3mo ago
One of my favourite things to do is to ask the models, "what does this code do and why?" They are usually not far from the truth. From the perspective of a LLM that has been trained on all of the public code on GitHub, issues, and
28.
▲
by
seunosewa
3mo ago
We often don't understand the code we wrote 6 months ago.
29.
▲
by
seunosewa
3mo ago
You can just ask the model to explain the code to you.
30.
▲
by
seunosewa
3mo ago
You can get some work done by using low power mode even when plugged in, and making your fan start running when the temps just start to rise (maybe 40 degrees. Use a third party fan app to set it up
More ›