Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
s-macke
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
91.
▲
by
s-macke
1y ago
Not sure if this counts as an implementation, but this is my Forth compiler, written in the high-level language Go [0]. This is also a challenge, as Forth usually requires machine language, call stack access, and is naturally written in ass
92.
▲
Reinforcement Learning Finetunes Small Subnetworks in Large Language Models
(arxiv.org)
3 points
by
s-macke
1y ago
|
0 comments
93.
▲
The simplest, fastest repository for training/finetuning small-sized VLMs
(github.com)
2 points
by
s-macke
1y ago
|
0 comments
94.
▲
by
s-macke
1y ago
This seems to be a formatting error. For such a huge author list, you usually write only the first name and then "et al." for "others".
95.
▲
by
s-macke
1y ago
There is a FAQ in the link.
96.
▲
Gemini 2.5 Pro won Pokémon Blue in 106k moves
(twitch.tv)
2 points
by
s-macke
1y ago
|
3 comments
97.
▲
SWE-Smith: Scaling Data for Software Engineering Agents
(arxiv.org)
1 points
by
s-macke
1y ago
|
0 comments
98.
▲
by
s-macke
1y ago
This emulator does basically the same but is much more speed optimized. It uses the OpenRISC architecture and even has networking. For what do you want to use such an emulator? [0] https://github.com/s-macke/jor1k
99.
▲
Sam Altman: we added one million users in the last hour
(twitter.com)
4 points
by
s-macke
2y ago
|
0 comments
100.
▲
Measuring AI Ability to Complete Long Tasks
(arxiv.org)
1 points
by
s-macke
2y ago
|
0 comments
101.
▲
by
s-macke
2y ago
Yes, permanently. Sonnet 3.7 is already number one in the ranking. Grok3 has no API yet.
102.
▲
by
s-macke
2y ago
Simple Bench goes in this direction: https://simple-bench.com/
103.
▲
by
s-macke
2y ago
o3-mini was announced for today, and OpenAI typically publishes in the morning hours (PT). Many people were eagerly waiting. The publication was imminent. I kept checking both Twitter and Hacker News for updates. Just add ten more people li
104.
▲
FrontierMath Was Funded by OpenAI
(lesswrong.com)
22 points
by
s-macke
2y ago
|
4 comments
105.
▲
Grokking at the Edge of Numerical Stability
(arxiv.org)
1 points
by
s-macke
2y ago
|
0 comments
106.
▲
by
s-macke
2y ago
> Notably, no self-reflection training data or prompt was included, suggesting that advanced System 2 reasoning can foster intrinsic self-reflection. They suggest, that self-reflection is an emergent phenomena of reasoning. Impressive. C
107.
▲
LLMs struggle with perception, not reasoning, in ARC-AGI
(anokas.substack.com)
5 points
by
s-macke
2y ago
|
0 comments
108.
▲
by
s-macke
2y ago
The term "agent" is quite broad. In my definition, an LLM becomes an agent when it utilizes the tool usage option. ChatGPT is a good example: you ask for an image, and you receive one; you ask for a web search, and the chatbot pro
109.
▲
by
s-macke
2y ago
For me it is 9:05 by Adam Cadre [0]. Short, linear, easy but with a great twist. [0] https://en.wikipedia.org/wiki/9:05
110.
▲
by
s-macke
2y ago
Here is the larger discussion about the Alice in Wonderland Paper on Hacker News. https://news.ycombinator.com/item?id=40585039
111.
▲
by
s-macke
2y ago
Tried it with N=2 and M=1 (brother singular) with the gpt-4o model and CoT. 1. 50% success without "full" terminology. 2. 5% success with "full" terminology. So, the improvement in clarity has exactly the opposite effect
112.
▲
by
s-macke
2y ago
I have found multiple definitions in literature of what you describe. 1. Fast thinking vs. slow thinking. 2. Intuitive thinking vs. symbolic thinking. 3. Interpolated thinking (in terms of pattern matching or curve fitting) vs. generalizati
113.
▲
by
s-macke
2y ago
Well, my perspective on this is as follows: The recurrent or transformer models are Turing complete, or at least close to being Turing complete (apologies, I’m not sure of the precise terminology here). As a result, they can at least simula
114.
▲
by
s-macke
2y ago
I am not a native English speaker. Can you reformulate the problem for me, so that every alternative interpretation is excluded?
115.
▲
by
s-macke
2y ago
We don't know. The paper and the problem was very prominent at that time. Some developers at Anthropic or OpenAI might have included that in some way. Either as test or as a task to improve the CoT via Reinforcement Learning.
116.
▲
by
s-macke
2y ago
> perhaps humans also mostly reason using previous examples rather than thinking from scratch. We do, but we can generalize better. When you exchange "hospital" with "medical centre" or change the sentence structure a
117.
▲
by
s-macke
2y ago
And here lies the exact issue. Single tests don’t provide any meaningful insights. You need to perform this test at least twenty times in separate chat windows or via the API to obtain meaningful statistics. For the "Alice in Wonderlan
118.
▲
by
s-macke
2y ago
These results are very similar to the "Alice in Wonderland" problem [1, 2], which was already discussed a few months ago. However the authors of the other paper are much more critical and call it a "Complete Reasoning Breakdo
119.
▲
by
s-macke
2y ago
When I first read about AlphaChip yesterday, my first question was how it compares to other optimization algorithms such as genetic algorithms or simulated annealing. Thank you for confirming that my questions are valid.
120.
▲
A Llama 70B finetune that has reflection baked into it's weights
(huggingface.co)
6 points
by
s-macke
2y ago
|
0 comments
More ›