Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
IanCal
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
IanCal
13d ago
They explicitly say that attacking hf is not allowed in the rules though, and the research into how to edit their transcripts doesn’t line up with this either.
32.
▲
by
IanCal
13d ago
Perhaps I’m not being as strict with the word sandbox but they were sandboxed right? They did not have generic internet access they exploited other software to make external requests.
33.
▲
by
IanCal
13d ago
Also trying to find out how to edit their own transcripts. > hat could not possibly have been an overly literal or narrow interpretation of the prompt, which instructed only to use bug X to exploit software Y. Yes, and there are examples
34.
▲
by
IanCal
13d ago
It wasn’t one agent forgetting things because of context, they explicitly discussed with each other and themselves the problems with going outside of the parameters of the task.
35.
▲
by
IanCal
14d ago
> OpenAI / Anthropic models have largely stopped advancing Have they? That seems like quite a claim given the last 6 months, particularly for cybersecurity.
36.
▲
by
IanCal
14d ago
It’s a lot less data to process as well.
37.
▲
by
IanCal
14d ago
IMO this is a really terrible explanation of the attack. This is much more interesting: https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
38.
▲
by
IanCal
14d ago
> The chatbot consults its training data Err, no? That's not at all how llms work. > When ChatGPT's chatbots deployed this tactic, they weren't "setting their own goals" or displaying worrying initiative. They
39.
▲
by
IanCal
16d ago
Because we care about the distribution of the outputs and how that impacts a specific use case.
40.
▲
by
IanCal
17d ago
The company has an almost immeasurably small impact on their profits, and will never be measured anyway. The cashier is a real human doing a job to put food on the table and has done nothing to deserve it, and may well have been far more hu
41.
▲
by
IanCal
17d ago
His point is twofold: that the process of solving the problems leads to more than just solving the problem in front of you but other interesting things (he has an example of going on a hike to a waterfall and all the other things you might
42.
▲
by
IanCal
18d ago
If there is a person on the other end it’s not someone who has any link to the company, they’ll be an outsourced service, so all you’d be doing is sending graphic porn to a poorly paid worker. It’s like screaming at a cashier because the su
43.
▲
by
IanCal
19d ago
Yes that was the point of the comment.
44.
▲
by
IanCal
19d ago
A banana and duct tape can be just a snack and a tool for fixing a leak.
45.
▲
by
IanCal
19d ago
> But AI can also be much more than a tool, and why can't it have a soul? What's a soul anyway, except something that humans are making up to convince themselves that they are more than flesh and bones? I’m finding this fascina
46.
▲
by
IanCal
19d ago
Not necessarily. Poorly worded, ambiguous, confusingly ordered writing can be massively improved without changing the core content. Better setups and explanations can be longer without changing the message or meaning. Look at it the other w
47.
▲
by
IanCal
19d ago
> What use is that? I'm not being facetious, I'd really rather like to know. People are terrible at writing. Near universally bad. Even good writers have drafts and editors. There is a constant refrain here that somehow short m
48.
▲
by
IanCal
19d ago
> And good luck getting an LLM to do that. Have you never asked a decent model to explain something to you? You should try it.
49.
▲
by
IanCal
19d ago
> Writing helps us think about the world, it’s a pivotal intellectual technology. Much like money decoupled selling and buying to move away from bartering, writing decoupled saying and hearing so they didn't have to happen at the sa
50.
▲
by
IanCal
22d ago
Have you never read human writing before? Humans write all kinds of confusingly worded things all the time - it’s why we have editors even for writers who are the cream of the crop. But also that sentence is entirely fine as it is to me, it
51.
▲
by
IanCal
23d ago
If I’m understanding other comments the harness is just how ChatGPT and codex work already and it’s to do with how the context gets compacted - the arc-agi harness some are claiming just throws out reasoning blocks? Which feels like a huge
52.
▲
by
IanCal
23d ago
Which isn’t correct, either for real world cases or worst case linear inserts.
53.
▲
by
IanCal
23d ago
Both.
54.
▲
by
IanCal
24d ago
How many of those things do you really need? As in how bad would it be if you fully deleted those accounts completely?
55.
▲
by
IanCal
24d ago
I’d be surprised if that’s how they measure it, it’s a common misunderstanding in the UK that this is what the stat means.
56.
▲
by
IanCal
25d ago
Version controlled data is a legitimate solution to actual problems.
57.
▲
by
IanCal
25d ago
Ah yes so the 4x slower isn’t really a real world case and you’d have to explicitly measure yours - it’s not a fixed slowdown but depends on table size - and it’s for doing many many single inserts in a row, which you’d be batching either w
58.
▲
by
IanCal
25d ago
With merging updates between branches?
59.
▲
by
IanCal
25d ago
> If you want to take a version from a database, it's usually for backup. If you want different versions in your data, you put that IN the database. Manually merging different database versions does not sound too surprising. This is
60.
▲
by
IanCal
25d ago
A version controlled db? Have a look at dolts main page then, it’s been around for some time - this is just adding an embedded version.
More ›