Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
SwtCyber
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
SwtCyber
2mo ago
I think the key difference here is that you're using the LLM to create exercises, not to replace the learning material
32.
▲
by
SwtCyber
2mo ago
The test I'd like to see: take a problem set or task you couldn't solve beforehand, learn the topic this way, then try to solve it without the LLM in the loop
33.
▲
by
SwtCyber
2mo ago
One concern though: "100% accurate and free of hallucinations" is doing a lot of work. A second LLM pass can catch some mistakes, but it can also confidently agree with the first one
34.
▲
by
SwtCyber
2mo ago
That’s a direct result of how they’re fine tuned. RLHF and other mechanisms reward quick, locally correct answers that solve the user’s immediate problem. A reward for a solution like "take a two-day pause, rip out half the modules, an
35.
▲
by
SwtCyber
2mo ago
I think the best part about this benchmark is that they finally stopped pretending you can evaluate a complex engineering task with $ 5 worth of inference. If a task takes a human weeks, you need to give the agent a comparable search space
36.
▲
by
SwtCyber
2mo ago
People still cook in houses with microwaves and wash dishes despite owning dishwashers
37.
▲
by
SwtCyber
2mo ago
The future arrived quietly
38.
▲
by
SwtCyber
2mo ago
Bradbury wrote rockets the way other writers wrote sunsets
39.
▲
by
SwtCyber
3mo ago
That is an amazing sleep indicator: once the rabbit starts discussing thermodynamics, dad has left the building
40.
▲
by
SwtCyber
3mo ago
The reading task can stay largely automatic until both streams try to use the same speech-production machinery at once
41.
▲
by
SwtCyber
3mo ago
It really does feel like reading and counting can occupy separate lanes, while writing and counting are both trying to use the same internal narrator
42.
▲
by
SwtCyber
3mo ago
This is a nice example of using interference as a window into representation
43.
▲
by
SwtCyber
3mo ago
This seems potentially useful for attention-steered hearing aids. A system that waits for complete disengagement from the old speaker may react too slowly
44.
▲
by
SwtCyber
3mo ago
The thing is this isn't a schema generation or Typescript bug at all. This is just how openai's function calling works under the hood. Their weights were fine-tuned for tool use to output the most complete data structures possible
45.
▲
by
SwtCyber
3mo ago
I would rather read an article with actual production experience migrating an agent, even if it is written in this style, than a perfectly crafted long read from another evangelist that has nothing but high level fluff and general phrases a
46.
▲
by
SwtCyber
3mo ago
Its ironic that under an article with a ton of deep infrastructure insights half the comments are crying about the "forced writing style". What does it matter if claude helped the author clean up the text when inside is a ready-to
47.
▲
by
SwtCyber
3mo ago
Its funny to see how researchers bypass Githubs praised guardrails with a simple word like "Additionally". It just proves that any attempt to build hard security boundaries inside an llm context window is bound to fail. The model
48.
▲
by
SwtCyber
3mo ago
The interesting bit is that they didn't solve division by rebuilding the whole cytoskeleton, they sort of sidestepped it
49.
▲
by
SwtCyber
3mo ago
"Not alive, but doing a suspicious number of alive-looking things" is a pretty good summary of why this is cool
50.
▲
by
SwtCyber
3mo ago
I think they arent even trying to build an AI detector. This is more of a social signal like "dont send us an automatically generated flood of changes"
51.
▲
by
SwtCyber
3mo ago
And it just keeps looping like that until the context window bursts. In practice the model is great at writing new code, but when you feed it its own six month old spaghetti code with a floating bug in the state machine it just starts hallu
52.
▲
by
SwtCyber
3mo ago
AI accidentally found one of the most expensive resources in the industry: the free time of people who maintain open source in the evenings after their day job
53.
▲
by
SwtCyber
3mo ago
That's just the basics. To craft a prompt for a complex architectural task, you need to know the solution at least on an abstraction level. If you don't have the right system design in your head, no llm is gonna conjure it out of
54.
▲
by
SwtCyber
3mo ago
Funny how almost every wave of automation starts the same way: "we're gonna cut headcount," but ends with "we're just shifting roles" AI is pretty good at scaling existing knowledge, but if the actual knowledge
55.
▲
by
SwtCyber
4mo ago
I think this depends a lot on how the message is phrased and what kind of action we're talking about
56.
▲
by
SwtCyber
4mo ago
The "ask for no" approach works best where the boundaries of ownership are clear. Without that, it becomes much riskier
57.
▲
by
SwtCyber
4mo ago
There's a difference between keeping someone informed and making them reown the problem
58.
▲
by
SwtCyber
4mo ago
The phrasing is not just a communication trick, it changes who owns the decision
59.
▲
by
SwtCyber
4mo ago
Good code is absent code LLMs by nature work like autocomplete on steroids. They're always trying to write more than necessary to please the prompt. Seniority now is measured by the ability to break down a task so that the agent doesn&
60.
▲
by
SwtCyber
4mo ago
It's good to see the hype around "programmers are no longer needed" giving way to a more realistic view. Generating lines of code was never the hardest part of engineering. What's much harder is understanding exactly wha
More ›