Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
fesens
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
HWE Bench: A new unbounded Benchmark for LLMs (GPT 5.5 is on top)
(hwebench.com)
6 points
by
fesens
5mo ago
|
3 comments
2.
▲
by
fesens
5mo ago
Current benchmarks have ceilings, usually 100%. This benchmark aims to be a long lasting, high correlation with the ability to solve real world problems and follow complex instructions, and unbounded (meaning it can always go higher).
3.
▲
by
fesens
5mo ago
Awesome! Let me know if it works for your propose. If not, raise a issue on github and lets work together.
4.
▲
by
fesens
5mo ago
His claims are indeed correct; Yes, you got my point tks!; AND the loop produced architecture gains that are not exclusive to the GoWin FPGA (CoreMark/Mhz is higher than VexRiscV)
5.
▲
by
fesens
5mo ago
Tks! Did you apply it to hardware design or to another field?
6.
▲
by
fesens
5mo ago
Whats your take on it? "If I read it, what point should I pay attention to?", I guess is what I'm trying to say
7.
▲
by
fesens
5mo ago
"Not at all, completely novel idea. What made you thought of such thing?" hahaha
8.
▲
by
fesens
5mo ago
Yeah, you are totally right. Its a work in progress, and the post was written by an LLM - Im trying to improve on it (dash pun intended). Regarding the benchmark overfitting, absolutely, it's pretty much overfitted. This CPU will only
9.
▲
by
fesens
5mo ago
I genuinely laughed reading the first words. Yeah, its hard to be novel
10.
▲
by
fesens
5mo ago
For sure! The hypothesis generation gotta be improved. Your take on the "least likely" is interesting. In the beginning of the repo I was having problems with "hypothesis convergence", your idea may be a nice way to intr
11.
▲
by
fesens
5mo ago
Is slop verifiable? If so we can throw it in the loop... The point is that this loop can be pointed at any verifiable work. Yeah you are seeing it raw, the verifier is the principle you talked about. Yes it was fully AI generated, It will b
12.
▲
by
fesens
5mo ago
Nice references! tks
13.
▲
by
fesens
5mo ago
Absolutely, today it's FPGAs.. Tomorrow can be whole companies
14.
▲
by
fesens
5mo ago
"I would LOVE somebody to bounce AI off of reversing the architecture and bitstreams for the stupid-ass closed-source FPGAs." The only reason I'm using Gowin is because it has a slightly more mature opensource tooling. Maybe
15.
▲
by
fesens
5mo ago
The frontier is the verifier not in the sense of this project, but to every project. If we have a good verifier for a task, any task, this type of loop can be applied to it. Today LLMs are good enough to tackle FPGA projects, but what this
16.
▲
by
fesens
5mo ago
Ive been receiving rate limits even with full quotas... I guess compute isn't growing as fast as demand
17.
▲
by
fesens
5mo ago
I like the way its looking. Maybe because of the familiarity with the Jupyter Notebook
18.
▲
Show HN: Auto-Architecture: Karpathy's Loop, pointed at a CPU
(github.com)
241 points
by
fesens
5mo ago
|
76 comments
19.
▲
by
fesens
5mo ago
Recently I've been noticing that claude "stalls" on a prompt and take minutes to finish, with no apparent work being done. Maybe they are rate limiting, even when you are well under the quotas.
20.
▲
by
fesens
6mo ago
The moment that not releasing a model becomes more financially beneficial than releasing it, that's when AI becomes truly dangerous. Between doomerism and marketing, the release or rather lack of, Claude Mythos is the first model that
21.
▲
Container GUI – Native macOS GUI for Apple's Container CLI
(github.com)
3 points
by
fesens
7mo ago
|
1 comments
22.
▲
by
fesens
7mo ago
Apple recently released https://github.com/apple/container , a CLI tool for running Linux containers natively on macOS with Apple Silicon. It's fast and lightweight, but it's terminal-only. I built Container G
23.
▲
Opus 4.5 designed a political Magic the Gathering set
(github.com)
1 points
by
fesens
8mo ago
|
1 comments
24.
▲
by
fesens
8mo ago
Its based on the latest political and economic events. I think it may be better than many recent sets.
25.
▲
Lovable for Enterprise Software
(usevento.com)
6 points
by
fesens
8mo ago
|
6 comments
26.
▲
by
fesens
8mo ago
We are building lovable for enterprise software, and we are genuinely impressed with what it can do in a single prompt. Authentication and persistence comes out of the box.
27.
▲
GQLite – A Tiny Embedded Graph Database in C
(github.com)
4 points
by
fesens
1y ago
|
1 comments
28.
▲
by
fesens
1y ago
This was mostly vibe coded with grok and o3 in two hours.
29.
▲
by
fesens
2y ago
The main advantage of using a new and constant token for reasoning is that, while we would pay the full price during training, in the inference phase, we could do most, if not all, the "reasoning" in one shot, without having to fe
30.
▲
by
fesens
2y ago
Reasoning 1 vs. 3 is the number of reasoning tokens between each "text" token. The 1 reasoning token is exactly what you see in the picture explanation in the article. The generalization comes from making the network predict a <
More ›