Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
hagen8
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
hagen8
2mo ago
Mountains close to leon?
2.
▲
by
hagen8
2mo ago
Cached input tokens are what drives most costs.
3.
▲
by
hagen8
2mo ago
Wrong. They are commonly used by millions.
4.
▲
by
hagen8
2mo ago
Check out academic papers about: 1. Hierarchical skills, workflow, skill learning 2. Meta Harness, self-learning harnesses 3. Trace/trajectory representation 4. Common agentic benchmarks But first more basic things like 5. Blog posts f
5.
▲
by
hagen8
2mo ago
Check out https://agents-last-exam.org/ there is still room for improvements!
6.
▲
by
hagen8
2mo ago
There are certain physical limits. Calculations need to be done. Either less calculations are necessary for the intelligence, or u accept less intelligence. But there is a limit in what u can do with specific hardware.
7.
▲
by
hagen8
2mo ago
Most importantly, after entering the elevator. First press the close button and then the floor. That way u, safe the time of pressing a button as the door is already closing.
8.
▲
by
hagen8
2mo ago
Where are the sources for that?
9.
▲
by
hagen8
2mo ago
This is the claude code frontend-skill.
10.
▲
by
hagen8
2mo ago
Just switch the model, its not that much effort tbh. And u can also get a cheaper model than 2.5 lite for the same intelligence
11.
▲
by
hagen8
2mo ago
This will soon happen with theoretical physics, computer science, and everything which can be verified cheaply. Then, we will have long running projects augmented by agents for 2-4 years while AI companies are collecting data of human workf
12.
▲
by
hagen8
3mo ago
Inference costs will go down massively once they use the upcoming GPUs. I estimated that a model like GLM5.2 will be around 0.03USD/M output tokens in 2 years when the Feynman GPUs will be available in 2028. And this did not even consi
13.
▲
by
hagen8
3mo ago
Some ppl don't like to hear it. But I would assume that token costs when using an inference provider are cheaper than electricity of using locally. If we just take into account output token generation for simplicity. With 5tps u get 18
14.
▲
by
hagen8
3mo ago
In my opinion Opus is waaayy better in agentic orchestration. It feels like it can natively deal with multiple subagents whereas gpt needs to be taught extensively.
15.
▲
by
hagen8
3mo ago
This is way to complex... Why don't just use some harness which manages all that and give u a good UI?
16.
▲
by
hagen8
7mo ago
Well, the question is what is contributing to the usage. Because as the context grows, the amount of input tokens are increasing. A model call with 800K token as input is 8 times more expensive than a model call with 100K tokens as input. E
17.
▲
by
hagen8
7mo ago
Did u use the API or subscription?
18.
▲
by
hagen8
7mo ago
They will sooner or later change that policy or get very slow in keeping up.
19.
▲
by
hagen8
7mo ago
But does it use the same agent harness? Because the harness determines the behavior a lot.