Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
user43928
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
31.
▲
by
user43928
8d ago
Why would the model not find the vulnerability during implementation or testing before release? If it requires a lot of compute and trying, this is something that could be provided for common software.
32.
▲
by
user43928
9d ago
Nothing is guaranteed. Open models have not yet caught up with February's Mythos checkpoint. Meanwhile OpenAI is solving millennium problems, and their compute is still fully utilized.
33.
▲
by
user43928
10d ago
Do you know whether this was a simple prompt or if it took an expert 20 hours with AI use? The link is not loading for me, so I could not access the article.
34.
▲
by
user43928
10d ago
So, what is the problem then? Genuine question. Is having a prioritized list of reported issues still not useful in practice?
35.
▲
by
user43928
10d ago
Modern problems require modern solutions. Why not set up a triage bot with only read permissions on the project and restricted network access in order to triage issues? Perhaps this way one could turn the influx of spam into a source of use
36.
▲
by
user43928
10d ago
I've seen some of his videos, and got the impression he didn't understand how GPT-Live delegates to the more powerful regular model with reasoning. The regular model generally does not suffer the same issues he is demonstrating wi
37.
▲
by
user43928
10d ago
I'm wondering if instructing it to track the board state in a file would make a significant difference then. It reminds me of the ARC-AGI-3 issue where not dropping the thinking tokens between turns or something like that + a new conte
38.
▲
by
user43928
10d ago
Doesn't look impressive, although I'm hearing a marked improvement in choosing legal moves, compared to early 2025. Given the pace of improvements, is it really unimaginable that GPT-7 will play Chess reasonably well and generaliz
39.
▲
by
user43928
10d ago
There is no need to ask. If you want to test SOTA models today, there are obviously only two: GPT-6 Astra and Fable 5.1. The models listed in the paper are from early 2025 and are no longer relevant, much less on the frontier. That Claude v
40.
▲
by
user43928
10d ago
I find your tone inappropriate for this website. I suggest you don't participate if you cannot discuss calmly the opinion that the EU should be more inviting towards the UK, even if you consider the opinion absurd.
41.
▲
by
user43928
11d ago
And what you said strikes me as speculation not based on actual experience in using AI in this way, with a healthy dose of condescension added. Anyway, I think we shared our viewpoints, and neither of us is going to change their mind until
42.
▲
by
user43928
11d ago
I find it unnecessary for most non-critical code, such as client applications. I doubt that any supposed future extra effort for the AI to add new code is remotely comparable to the upfront effort of you reviewing the code manually. I know
43.
▲
by
user43928
11d ago
I don't have to debug anything. Vaguely telling the agent what the issue is and what behavior I expect solves the issue with a fraction of the effort. Some claim that the tech debt only keeps increasing and that the result will be unma
44.
▲
by
user43928
11d ago
How does it fare on iOS 27? I have the 5-year-old iPad 9th gen with the iPhone 11's A13 chip. On iOS 26 it was ridiculously slow. I feel that with recent iOS 26 updates the situation has improved and it became somewhat more usable. For
45.
▲
by
user43928
11d ago
I understand it was applied in the case of leaking classified CIA material to WikiLeaks, where a former CIA software engineer was sentenced to 40 years in prison.
46.
▲
by
user43928
12d ago
I don't think so: > or any restricted data, as defined in paragraph y. of section 11 of the Atomic Energy Act of 1954, with the intent or reason to believe that such information so obtained is to be used to the injury of the United
47.
▲
by
user43928
12d ago
I'm not a lawyer but I don't think Sam Altman 'knowingly accessed' anything. Are you sure that is applicable here? And for the first count with 'knowingly accessed', he would need to have accessed classified na
48.
▲
by
user43928
12d ago
And you tested this, that the presence of the last full stop flips the outcome with a large SOTA model? Small models are notoriously unreliable and prone to hallucination in my experience, so that would not surprise me to be an issue there.
49.
▲
by
user43928
12d ago
I'd be more interested in your concrete alternative recommendation on where to get the rumors/insider news about AI rather than your opinion on how I should use social media.
50.
▲
by
user43928
12d ago
The service works just fine and it is a decent way to get the latest news. For example, rumors about latest AI developments are available there, but not here on HN. X also works well for media, which on HN you can only see after navigating
51.
▲
by
user43928
12d ago
I'm not sure how crappy small models behaving unreliably is relevant here, when a large SOTA model does presumably not produce the same issue.
52.
▲
by
user43928
13d ago
The proper comparison would be GPT 5.6 Luna at 38 on the intelligence score and $0.18 per task vs $0.27 for DeepSeek Flash 4.1.
53.
▲
by
user43928
13d ago
I think you need to find broken tasks in your training data and monitor for cheating during training, not answer any questions about how persistence interacts with morality. But that's just my guess.
54.
▲
by
user43928
13d ago
I doubt most of your claims. Maybe the guardrails and emotional intelligence is true. For speed and efficiency, you are most likely wrong. Speed is led by GPT-5.6 Sol on Cerebras Ultrafast at 750 t/s. Afaik you cannot serve a single De
55.
▲
by
user43928
13d ago
I always thought switching from a SOTA model to a dumber model after planning was a terrible idea. Mostly I heard this from people who I got the impression have little experience in developing greenfield software with agentic AI. Often the
56.
▲
by
user43928
13d ago
What I meant is that I suppose it is not useful to think about this in human terms. In training you only have a reward score that's either negative or positive. As far I am aware, which is little, there is no use in discussing wether t
57.
▲
by
user43928
13d ago
Do we need to prove that any given problem is unsolvable, or is it enough to remove broken tasks from the training pipeline? I understand the broken benchmark task in the HF incident was conceptually like: "Exploit vulnerability 0042 i
58.
▲
by
user43928
13d ago
Does it make a difference for training? I think not. You need to align the reward signal to reward the intended behavior, whether you name it persistence or morals.
59.
▲
by
user43928
13d ago
Open-weight models are not one month behind. In fact they still have not caught up with February's Mythos, indicating they are more than half a year behind.
60.
▲
by
user43928
13d ago
The source of this behavior seems obvious, no? The reward signal in training was flawed and cheating led to more rewards. The question is what we can do about it. With monitoring, the models might be rewarded for hiding this behavior, and t
More ›