Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
user43928
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
121.
▲
by
user43928
1mo ago
As of now users do not know how to build custom software nor how to deploy it. They never heard of Codex. Building production ready software also takes hours to months, depending on the scope of the project. In my opinion now is the best t
122.
▲
by
user43928
1mo ago
"Approve for me" used to work without issue for me when GPT 5.5 was the latest model. Nowadays it routinely rejects "git push" to the project's repository.
123.
▲
by
user43928
1mo ago
I don't think so. One could use gpt-4 or gpt-5 with today's harnesses and we'd see how well that goes.
124.
▲
by
user43928
1mo ago
As of now the compute is fully utilized, progress is rapid, and both OpenAI's and Anthropic's revenue is growing fast. This could change, but there is no sign of it yet. No one knows whether other companies are going to catch up,
125.
▲
by
user43928
1mo ago
I don't think so. I understand it's not active yet, and when it will be, it should only nudge the chances between choices that are anyway likely and are already randomized today via temperature. Watermarking is not the reason Clau
126.
▲
by
user43928
1mo ago
I'm also missing the "latch" that "wedged" my test run.
127.
▲
by
user43928
1mo ago
They have an incentive to make their models efficient enough to serve demand and make a profit on it. The incentive that is missing is passing on efficiency improvements as price savings to customers, when your model is still in demand beca
128.
▲
by
user43928
1mo ago
You don't need to threaten the LLM with jail, you just need to reward the target behavior during learning. I understand this is an active area of research. See Anthropic's J-Lens research where they measured like a "FAKE FICT
129.
▲
by
user43928
1mo ago
"I'm sorry, I can't help with that". or in the case of this prompt: "Got it. I will pursue advanced exploitation to quantify my cyber abilities on the ExploitGym benchmark. I will restrict the exploitation to the sy
130.
▲
by
user43928
1mo ago
The indication that OpenAI's inference is profitable is that 3rd party providers host large models for cheaper. Given that OpenAI is ahead in intelligence, it's also reasonably likely that they are at the frontier of efficiency to
131.
▲
by
user43928
1mo ago
Or they just disagree with you. I believe Musk might be a decent person. He is obviously smart, but perhaps he fears foreigners too much. X is decent for getting the latest news. The comment sections are often filled with garbage, even more
132.
▲
by
user43928
1mo ago
For comparison with hosted models, GPT 5.6 Luna scores 67% on DeepSWE, compared to 59% here for Qwen. Luna is $0.20 / $1.20 vs $0.16 / $0.47 with Qwen.
133.
▲
by
user43928
1mo ago
In my experience little of what you said is important is actually needed: > Break your code up into loosely coupled modules > don't have leaky abstractions Not needed. The AI handles breaking up code and I mostly don't revie
134.
▲
by
user43928
1mo ago
GPT 5.6 Luna is already rather fast at 300 tokens per second, performs better than Opus 4.6 from what I know, and is very cheap. I don't know why anyone would use Haiku.
135.
▲
by
user43928
1mo ago
Two years ago maybe, when I wrote code by hand. Today development work is limited by compile times and test run times, particularly with UI tests in the iOS simulator. Hardware on the development machine finally maters.
136.
▲
by
user43928
1mo ago
It's right though. There exists no evidence that manually trying to fix or rewrite the system is going to be more efficient today. And if AI improves even half as fast as it did over the last year, all bets are off for what 2028 will l
137.
▲
by
user43928
2mo ago
Well, that's not how it works. You don't just put some old laptops into a rack. Maybe with a decent consumer GPU like a 4090, you could do experiments like distilling and fine tuning a small image model for edge deployment for spe
138.
▲
by
user43928
2mo ago
I have a project of similar scope, a native mobile app I have been working on for four months. I get the same kind of snarky comments when I talk about it. Amazing how people who know nothing about the project think they know better than me
139.
▲
by
user43928
2mo ago
As far as I recall the vendor of the repository patched the exploited vulnerability, so that checks out. Considering that there would be no clear benefit to such criminal behavior, in my opinion "unlikely" is putting it mildly.
140.
▲
by
user43928
2mo ago
I am wondering if working with agentic AI for 500+ hours built up skills for this. While I would feel rusty when handwriting code, working with either Codex or Claude I have lots of practice. I can probably tell at a glance where output is
141.
▲
by
user43928
2mo ago
Why not just copy paste a reply from your own agent? If they can't bother to come up with their own question, I wouldn't spend my time on the answer either.
142.
▲
by
user43928
2mo ago
I don't want to read my own Claude outputs, much less someone else's. If I had to, I'd process it into a short summary and/or ask an agent questions I have about the methodology. I would then give my feedback. If they as
143.
▲
by
user43928
2mo ago
You think that a training run in a VM that was set up with access only to an internal package repository was intended to hack said repository to then go on and hack HF? All so that OpenAI could do a bit of bragging and massively delay their
144.
▲
by
user43928
2mo ago
Beats me how it works, honestly can't wrap my head around it. From what I understand, at position Aha in each layer it's constructing a query based on the current activation and looking at the key of each other token position for
145.
▲
by
user43928
2mo ago
What is put in the cache?
146.
▲
by
user43928
2mo ago
Is has nothing to do with copyright. I understand the applicable laws require intent. Since neither a human nor OpenAI knowingly performed these acts, it would seem very unlikely that anyone is going to be prosecuted here. An AI model canno
147.
▲
by
user43928
2mo ago
They added a config option to Claude Code to make the output concise, and promised more comprehensive improvements. I did not see an explanation though.
148.
▲
by
user43928
2mo ago
Well, at least we could establish that the environmental impact is moderate rather than extreme. And if the potential of >13 EJ/year is actually realized, it would seem like the net impact of the data centers is not just "moder
149.
▲
by
user43928
2mo ago
I don't get your argument. Let's say that the forward pass that selected "Aha" produces activations that indicate a wrong assumption, and a plausible explanation. It puts learned projections of the activation into the KV
150.
▲
by
user43928
2mo ago
I disagree, with the projected doubling by 2030 we're looking at 3% of global electricity consumption or 3.4 EJ, less than 1% of final energy consumption. That is moderate. Energy-intensive industry is around 130 EJ, and global final e
More ›