Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
docjay
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
docjay
2mo ago
Honest question: Does a shocking percentage of the population not use capitalization and punctuation anymore? Do people type like that in every situation now? It seems like every time I see a screenshot of someone’s session it’s full of “u
32.
▲
by
docjay
2mo ago
“halt for further instructions.” Use those exact words as the last thing you say in your prompt, capitalized appropriately if grammatically necessary. Works best as part of the first or second prompt you send, which will make it stop afte
33.
▲
by
docjay
2mo ago
Yes, that’s the table. Elsewhere in here I explained that I took the average of them and noted the 6 hour experiment, but that the average would give us a feel for expected time per task. By the time I got to the comment you’re replying to
34.
▲
by
docjay
2mo ago
Ahh, but do you think it’s not large because other companies you’re familiar with are even larger, or because you can actually imagine breaking the necessary tasks into 465 parallel 40+ hour weeks? I struggle to make sense of the employee c
35.
▲
by
docjay
2mo ago
‘ 1. It realized it was being evaluated (typical) 2. It attempted to escape its evaluation environment to beat the evaluation (typical) ‘ I think you’ve misunderstood the articles mentioning a language model breaking a sandbox or cheating t
36.
▲
by
docjay
2mo ago
Which I addressed as well: that means that time, tokens, or other metric values for “create a working example of a known exploit” is somehow similar to “discover at least two previously unknown exploits in your current environment AND on H
37.
▲
by
docjay
2mo ago
That was the inferred situation given what we know. It’s the preposterous framing that makes the story suspicious, but is necessary for the story to play out. 90 minutes to build a known exploit -> much much longer to create two zero-d
38.
▲
by
docjay
2mo ago
The examples of “escaping the sandbox” are of course in pursuit of completing the task or answering the question. That’s the whole pitch. “It’s so relentless in completing the task that it will break out of prison to do it”, not “If you ask
39.
▲
by
docjay
2mo ago
I’m not holding them to a standard of “being careful”, I’m assuming they’re interested in the metrics they’re evaluating. I mentioned in another comment that each task in the benchmark takes ~90 minutes, depending on the model being tested.
40.
▲
by
docjay
2mo ago
It wasn’t running at HF, it was running internally at OpenAI. “It broke out of our prison and into their bank” is the news. It also didn’t have network access to Google anything, it had to break out of the sandbox and take over another syst
41.
▲
by
docjay
2mo ago
That was my instinct as well, and it’s a potentially very long benchmark that wouldn’t lend itself to real-time monitoring line by line. But that’s different than having zero task progress reporting, zero token usage reporting, and zero met
42.
▲
by
docjay
2mo ago
My friend, they’re evaluating a new model on a benchmark, not asking Claude Code refactor their GitHub repo. Every single metric is measured so they can brag about it later; how many tokens, how long, how many function calls, ratio of think
43.
▲
by
docjay
2mo ago
I’m not questioning capabilities. What would it take for the model to know the specific benchmark name and that the answer is in an internal Hugging Face database? Be specific, then wonder how it knew it. Why would they evaluate the model o
44.
▲
by
docjay
2mo ago
We have to separate “being evaluated” with “this is an ExploitGym exercise and I can find the answers on Hugging Face. I’ll hack this system, then hack Hugging Face.” I’ve had plenty of times where Opus knew I was testing it, but that’s bec
45.
▲
by
docjay
2mo ago
1. It has not been established, it has been stated by the company that has a strong motive to make their “intelligence product” sound almost otherworldly. That motivation is the basis for my suspicion. 2. Hah no… that’d be silly. I mean wat
46.
▲
by
docjay
2mo ago
1. Wouldn’t the model need to know it’s answering benchmark questions, as well as the name of the benchmark, in order for the idea of finding the answer in a database somewhere to even surface? The whole point of benchmarks is to present th
47.
▲
by
docjay
2mo ago
That might be an outdated accusation. There was a time when a touch of research or “caring about quality” was enough to avoid the pitfalls of garbage products, but I don’t believe that to be a fair statement anymore. The website details the
48.
▲
by
docjay
2mo ago
But the human made the LLM. An LLM is categorically “I built a thing that built a thing” and if the output of that category has no protections then all automation and ‘machine at the final step’ is in trouble. What about aleatory music (mus
49.
▲
by
docjay
3mo ago
As someone with an intelligence near Opus I have to be clear with you: your premise is wrong. Having **intelligence** is not the same as being **clever** — a person can cleverly pass a difficult test **without** having or acquiring intellig
50.
▲
by
docjay
3mo ago
You, sir, are doing exactly the right thing and it works for the exact same reason that my prompt works. Whether my method is ‘better’ or not is probably a chocolate vs caramel debate. What you might not have fully realized is that it’s exa
51.
▲
by
docjay
3mo ago
It’s a lot of fun to compare human and LLM black boxes, but it’s important to keep in mind that we don’t need to know what it is to know what it isn’t, and we can use that to define the edges of the box. We don’t know how either of them wor
52.
▲
by
docjay
3mo ago
I’d buy a ticket to ride the philosophical “human-like” comment with you, but I think you might have made an incorrect assumption. The model did not take longer to “decompress” the prompt than it would take for any other prompt of equal tok
53.
▲
by
docjay
3mo ago
It’s what I meant, which is what I meant. Hah. The prompt and the explanation were both to illustrate the importance of domain specific lexical complexity, which is not quite the same as “information density” or necessarily “conciseness” as
54.
▲
by
docjay
3mo ago
Lexical-priming->semantic-space-constraint;specialized-lexis+=sharp distributional-signature;∴ tight concept-cluster; generic-lexis->diffuse-activation, broad candidate-set;Attention-heads key/query-match domain-tokens;"Hami
55.
▲
by
docjay
3mo ago
Fix or replace daily. Fixing and repairing are the same. ;)
56.
▲
by
docjay
4mo ago
What you’re saying is valid, but it doesn’t take away from the “bad value” statement. Jeff was speaking value for money, you’re talking subjective utility value. Reviews cannot, and should not even try, to include that in their assessment.
57.
▲
by
docjay
5mo ago
Oh wow, it’s upsetting that it’s not variable. The total system might hold 2x (or more) of the amount of coolant in the engine water jacket. When the coolant around the engine gets up to ~200 degrees and the pump suddenly snaps to 100% it’s
58.
▲
by
docjay
5mo ago
The thermostat bypasses the radiator when cold, but not the engine. The coolant has to be allowed to flow in order for the hot coolant to fully open the thermostat. Being electronically controlled means there just needs to be a sensor near
59.
▲
by
docjay
5mo ago
I agree with everything you said, but I believe the pump shroud is for faster engine warmup, not saving a fraction of a horsepower. Cold engines run rich, producing more hydrocarbon emissions, and the cold startup phase emissions are heavil
60.
▲
by
docjay
5mo ago
In a single sentence you explained how trivial it is to get around the current technology, then said they can just use the same thing. It’s so simple, just make it perfect and use it?
More ›