Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
cruffle_duffle
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
61.
▲
by
cruffle_duffle
6mo ago
Piles and piles of sci-fi novels.
62.
▲
by
cruffle_duffle
6mo ago
This is actual reason. So any investors reading our system card.... write us another check and watch the $$$$$$$$ roll in. It's so dangerous we can't even release it!
63.
▲
by
cruffle_duffle
6mo ago
"for example, consider using models to write email -- is it a misalignment problem if the model is just too good at writing marketing emails?? or too good at getting people to pay a spammy company?" But who gets to be the judge of
64.
▲
by
cruffle_duffle
6mo ago
These things are so damn cool!
65.
▲
by
cruffle_duffle
6mo ago
Cautious for what? Unchecked doomerism? Just release the damn models. Do it in phases, roll it out slowly if they are so damn worried about "safety". The real reason they aren't releasing it yet is probably it eats TPU for
66.
▲
by
cruffle_duffle
6mo ago
It sounds like the PoE spec was designed before the arrival of “IoT” type things like the esp32, raspberry pi’s, etc. How much of the complexity is a “fundamental electrical engineering problem” and how much of it is just a spec written to
67.
▲
by
cruffle_duffle
6mo ago
Smaller teams working on much more diverse set of problems. The truth is absolutely nobody knows how this will all shake out.
68.
▲
by
cruffle_duffle
6mo ago
Subagents are isolated context windows, which means they cannot get polluted as easily with garbage from the main thread. You can have multiple of them running in parallel doing their own separate things in service of whatever your own “bra
69.
▲
by
cruffle_duffle
6mo ago
To paraphrase, spacex is "making the impossible merely late"
70.
▲
by
cruffle_duffle
6mo ago
It'd dogfooding the entire concept of vibe coding and honestly, that is a good thing. Obviously they care about that stuff, but if your ethos is "always vibe code" then a lot of the fixes to it become model & prompting c
71.
▲
by
cruffle_duffle
6mo ago
Thanks!!!
72.
▲
by
cruffle_duffle
6mo ago
> How can people be so naive as to run something like Claude anywhere other than in a strictly locked down sandbox that has no access to anything but the single git repo they are working on (and certainly no creds to push code)? Because
73.
▲
by
cruffle_duffle
6mo ago
There is also the browser I use to get Claude to route around people blocking its webfetch. Both Playwright and chrome-mcp.
74.
▲
by
cruffle_duffle
6mo ago
I bet dollars to doughnuts that 95% of the traffic is from Claude and ChatGPT desktop / mobile and not literal content scraping for training.
75.
▲
by
cruffle_duffle
6mo ago
It will mess up eventually. It always does. People need to stop thinking of this is a “security against malicious actor” thing… because thinking in that way blinds you to the actual threat… Claude being helpful and accidentally running a co
76.
▲
by
cruffle_duffle
6mo ago
Hah… I’ve seen Claude happily and very cleverly find ways to escape its sandbox. It’s like some kind of arms race between the model and its designers.
77.
▲
by
cruffle_duffle
7mo ago
Dude. I’ve been thinking about this a lot! I think it’s because the traditional way we internalize the costs of what we are building just got take for a ride. We don’t really (or I don’t anyway) fully know what “too much scope” feels like
78.
▲
by
cruffle_duffle
7mo ago
A lot of getting good mileage out of LLMs is promoting them to behave like they are blind and can only base their outputs on what is in front of them. Maintain an emic stance.
79.
▲
by
cruffle_duffle
7mo ago
That is an interesting way of looking at that, thanks for the perspective! Like, the words fit… why create a second parallel language for describing LLM behavior. Somebody else said it… the whole “it’s a stochastic parrot” thing is sooooo c
80.
▲
by
cruffle_duffle
7mo ago
An alternate way of thinking about it is LLMs have no reflection capability. Literally any “reflection” it claims to have about its decision making is made up. It has absolutely no way know that what it said was based on some ancient prover
81.
▲
by
cruffle_duffle
7mo ago
That is so reductive of an analysis that it is almost worthless. Technically true, but very unhelpful in terms of using an LLM. It is a first principle though so it helps to “stir the context windows pot” by having it pull in research and
82.
▲
by
cruffle_duffle
7mo ago
That is why you have to always have it ground itself in something. Have it search for relevant research or professional whatever and pull that into context. Otherwise it’s just your word plus its training data. I had to deal with a close
83.
▲
by
cruffle_duffle
7mo ago
> goodness gracious its all very time consuming and im not sure its worth the squeeze And when you step back you start to wonder if all you are doing is trying to get the model to echo what you already know in your gut back to you.
84.
▲
by
cruffle_duffle
7mo ago
It will be crazy. Because the cost of “failure” will be dramatically lower, meaning these things can sometimes just throw educated darts at the wall until a solution is found. It’s way too slow to do that kind of thing now. (Presumably cos
85.
▲
by
cruffle_duffle
7mo ago
This advice will be very dated when inference gets an order of magnitude faster. And it will happen—it’s classic tech. Probably will even follow moores law or something. Wait until that 8 minute inference is only a handful of seconds and
86.
▲
by
cruffle_duffle
7mo ago
And 1/3 of all people who think others outsource their thinking to others also outsource their own thinking to others. Not you or me of course. It’s the other 1/3. Probably some lurker reading this.
87.
▲
by
cruffle_duffle
7mo ago
> Cache is money saver in computing. Their own client might be lot better at caches than any other agent so they do not want to lose money yet end up with disgrunted customer that claude isn't working as good I’d bet a reasonable am
88.
▲
by
cruffle_duffle
7mo ago
Dunno if you know this but the plan in plan mode is a markdown file! Ask it for the file and it will give it to you.
89.
▲
by
cruffle_duffle
7mo ago
Let me guess the command: [error] Wait, better check help. is it -h? [error] Nope? Lemme try —-help. [error] Nope. How about just “help” [error] Let me search the web [tons of context and tool calls]
90.
▲
by
cruffle_duffle
7mo ago
And that is the thing about raising a kid these days. Those damn machines have replaced so much… because yeah Minecraft is like a souped up version of Lego where in creative mode you have every part you need. And you don’t have to dig for
More ›