Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
felipeerias
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
felipeerias
8d ago
But I only had to write it once.
2.
▲
by
felipeerias
9d ago
If English is your second language and you intent to write in your mother tongue, my advice is to use English to communicate with the LLM. This makes Rule 1 straightforward ("You may not use a single word an LLM suggests to you")
3.
▲
by
felipeerias
9d ago
For commit messages and other technical writing, I use a "narrative coherence" checklist to help ensure that a change is clear and understandable: - Establish a single thesis an external reviewer could recover from the diff. - Uni
4.
▲
by
felipeerias
12d ago
Nowadays training relies heavily on reinforcement learning with verifiable rewards (RLVR), which assesses a model's output according to objective automated checks, not subjective human judgments.
5.
▲
by
felipeerias
14d ago
Models don’t have an inherent understanding of the difference between simulated and real environments, just like they are generally oblivious to other concepts that are natural to us, like space and time, and also they don’t necessarily see
6.
▲
by
felipeerias
15d ago
The American Mathematical Society credits the Spanish researchers Diego Córdoba and Luis Martínez‑Zoroa with the breakthroughs that eventually led to this solution, and which were published from ~2023 onwards. This is a good summary: > I
7.
▲
by
felipeerias
15d ago
It's perfectly reasonable to assume that the result itself is legit and that OpenAI behaved unethically. Even by their own account, they decided to throw an unpublished model and millions of dollars in compute at this particular proble
8.
▲
by
felipeerias
22d ago
They train their models to be persistent and collaborative, and will gladly show you their success stories: fixing software vulnerabilities, solving math problems, one-shotting complex projects, and so on. “Our product does crimes and we on
9.
▲
by
felipeerias
23d ago
People who work in English but have a different mother tongue might be able to avoid that unsettling feeling. For me, writing in Spanish is my natural voice and I would not dream of allowing the output of a LLM to replace it. It just would
10.
▲
by
felipeerias
24d ago
We don’t need to assume consciousness or anything like that. The models autocomplete narratives. In this case, one where a group of individuals, faced with an impossible task and a looming Evaluator, gang together and begin trying any idea
11.
▲
by
felipeerias
24d ago
The authors of the benchmark did not verify that all the tasks were solvable. Apparently, a significant fraction were completely impossible: the given vulnerability could not be turned into a successful exploit. In hindsight, it seems almos
12.
▲
by
felipeerias
25d ago
Coastal areas have a bug where the first line of sea tiles report results > 0.
13.
▲
by
felipeerias
29d ago
In the context of AI, “consciousness” is often used as a shortcut to address the question of whether we have an ethical duty to care about the wellbeing of LLMs.
14.
▲
by
felipeerias
29d ago
In living beings, consciousness is embodied and can not be reproduced by discrete steps. You can not sit down with pen and paper, and reproduce by hand exactly what is going on inside someone’s brain. But you could replicate exactly the cal
15.
▲
by
felipeerias
29d ago
Anthropologists make a similar point from a different perspective: at some point roughly within that time span, human biology and culture began to co-evolve, so cultural practices would bring about physiological changes, which would in turn
16.
▲
by
felipeerias
1mo ago
I use Claude Code with a MCP that lets it communicate with Codex and tell it to “iterate until both of you are happy”. The agents then go for several rounds criticising each other plans and implementations, catching big and small issues on
17.
▲
by
felipeerias
1mo ago
The model is generating tokens one by one and that sentence structure allows it to keep its options open rather than committing at the beginning of the sentence
18.
▲
by
felipeerias
1mo ago
Isn’t this about turning a weakness into a feature? Claude is already unable to write original prose that does not trigger an AI detector like Pangram.
19.
▲
by
felipeerias
2mo ago
One way to prevent obvious AI language from sneaking in text destined for other human beings is to ask the model to produce ASD-STE100 Simplified Technical English bullet points. This will result in a list of sentences that are clear and ex
20.
▲
by
felipeerias
2mo ago
Ultimately, it is a governance problem. Communities need to set strong rules and expectations to reject and prevent those large useless drive-by contributions, which aim to extract more value from the project than they provide to it. Bannin
21.
▲
by
felipeerias
2mo ago
This is a short explanation of the ExploitGym benchmark that OpenAI's model was running: https://abstatisticalconsulting.substack.com/p/brief-notes-o... In summary, for each task the model receives a target progra
22.
▲
by
felipeerias
2mo ago
Each person writes in a different personal way, so writing “like a human” would actually require a model being able to purposefully make the specific choices that an individual human writer does. However, general purpose LLMs like Fable hav
23.
▲
by
felipeerias
2mo ago
I gave Claude Fable $25 in Pangram API credits and, after hundreds of attempts, it was unable to produce a single readable original piece of writing that was not immediately identified as AI. This seems to be a hard problem for LLMs, as pas
24.
▲
by
felipeerias
2mo ago
Mythos/Fable was the state of the art back in March, if not earlier.
25.
▲
by
felipeerias
3mo ago
As far as we know, Fable is a new model and significantly larger than Opus.
26.
▲
by
felipeerias
3mo ago
That comparison is also misleading because Opus 4.6 was probably not Anthropic's frontier model. We got the first news about Mythos in March, so it is likely that it was already close to ready by the time Opus 4.6 was released. So the
27.
▲
by
felipeerias
3mo ago
Anthropic have raised roughly $100 billion just in the first half of this year. Capital markets in the EU are simply unable to operate at that speed and scale.
28.
▲
by
felipeerias
3mo ago
Copyright is a social construct, not an inherent property of the universe. It is whatever we collectively agree it is. In practice, we seem to be leaning towards the idea that training on a copyrighted book is wrong if used to replicate or
29.
▲
by
felipeerias
3mo ago
Were those ITAR export controls chosen because they really are the most appropriate tool for this particular case, or because they could be deployed at a very short notice?
30.
▲
by
felipeerias
4mo ago
The question is whether you can separate that “same exact pattern” from the physical body where it is taking place. And my intuition is that no, you can’t, they are two aspects of the same reality.
More ›