Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Topfi
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
13 ms
·
151.
▲
by
Topfi
3mo ago
Blast from the past for me, though primarily interacted with the complementary Whonix side of things. Not surprising to read, considering how lean Qubes was from the get-go designed to be it makes sense that most things are from resulting u
152.
▲
European Search Perspective
(eu-searchperspective.com)
1 points
by
Topfi
3mo ago
|
0 comments
153.
▲
by
Topfi
3mo ago
Rational, measured, accurate to the current state of the tech, overall a good position for a project like this.
154.
▲
by
Topfi
3mo ago
> What I'm questioning here is that there are labs who have sat down and deliberately tested and tweaked the performance for this particular task, independent of general model improvements. Given the massive delta easily reproducibl
155.
▲
by
Topfi
3mo ago
Happy to, here one example where Grok 4 Fast, despite producing a fairly consistent pelican [0], did severely worse in a similarly outlandish scenario along with Haiku 4.5 and GPT-5 for context: https://news.ycombinator.com/
156.
▲
by
Topfi
3mo ago
Maybe it gets posted every time because besides a personal believe by the person popularising this "benchmark", there is no reason to assume that certain labs aren't intentionally training to game this and every other lab at
157.
▲
by
Topfi
3mo ago
Respectfully, did you? The comment was specific to doubting the believe simonw has that labs are not training [0] specifically for this task, which is exactly what simonw wrote in the post [1], that it is a believe of his that they don'
158.
▲
by
Topfi
3mo ago
Quick and still very early update, the model has (with web search disabled which was verified via the reasoning traces) accurately answered a number of questions focused on very niche details (engine specific maintenance in certain newtimer
159.
▲
by
Topfi
3mo ago
North Mini Code by Cohere (HQd in Toronto) has honestly been very competitive in my personal assessment with many of the models coming out of the PRC. I'd position it below Moonshot AIs and Z.ais recent releases, but above the varietie
160.
▲
by
Topfi
3mo ago
British-American with much of the research happening in London. I don't know if it's known what team specifically worked on which of the Gemmas, I'd suspect it's a healthy mix of multiple satellites across the globe, but
161.
▲
Microsoft Project Aion (Copilot OS Incubation Effort) [video]
(youtube.com)
3 points
by
Topfi
3mo ago
|
0 comments
162.
▲
by
Topfi
3mo ago
Very preliminary testing so far, but there is something here, far beyond what the benchmarks suggest. Only ever saw such outperformance of public evals vs my private ones with Anthropic models and while it is far to early to make any judgem
163.
▲
Inkling Model Card
(thinkingmachines.ai)
7 points
by
Topfi
3mo ago
|
0 comments
164.
▲
by
Topfi
3mo ago
Thanks, thought I was the only one expecting a tiny, coding focused model from the title. Codex really is the least consistent brand in tech.
165.
▲
by
Topfi
3mo ago
Donate 200 bucks to a struggling startup.
166.
▲
by
Topfi
3mo ago
Similar to companies working on FOSS codebases, hosting (sometimes with the license restricting third-parties in some way), providing tailored models and services to customer's and getting bought for your team if your model happens to
167.
▲
by
Topfi
3mo ago
Is this a joke? Instantly thought of this: https://www.reddit.com/r/claude/comments/1s7m8ld/this_is_the... Things you do if you definitely are focused on the a Trillion USD industry and SuperDuperUltraMe
168.
▲
Claude Fable 5 access extended through July 19
(xcancel.com)
9 points
by
Topfi
3mo ago
|
2 comments
169.
▲
We Tested $200 GPT-5.6 Sol on PhD Level Math [video]
(youtube.com)
3 points
by
Topfi
3mo ago
|
0 comments
170.
▲
Evaluating the impact of two decades of USAID intervention
(thelancet.com)
4 points
by
Topfi
3mo ago
|
0 comments
171.
▲
by
Topfi
3mo ago
Title was shortened and slightly editorialized from "OpenAI’s newest AI model is 54% more token efficient on agentic coding, Altman tells CNBC" for readability. I cannot watch the interview right now, but the way the sentence is w
172.
▲
Altman: GPT-5.6 is 54% more token efficient on agentic coding
(cnbc.com)
14 points
by
Topfi
3mo ago
|
4 comments
173.
▲
by
Topfi
3mo ago
I just checked and for plain old C, there do not seem to be any reasonably comprehensive, current-day eval suites. Fully admitting that, even if there were, I couldn't assess their validity simply because I have never written or review
174.
▲
by
Topfi
3mo ago
Either DeepSWE [0] or FrontierCode [1], depending on personal goals and requirements. The later is more interesting for me personally, due to the design of the benchmark heavily grading "mergability", i.e. how the provided output
175.
▲
Model and effort in Claude Code: knowing more vs. trying harder
(xcancel.com)
2 points
by
Topfi
3mo ago
|
0 comments
176.
▲
FrontierCode 1.1
(cognition.com)
2 points
by
Topfi
3mo ago
|
0 comments
177.
▲
by
Topfi
3mo ago
Very. Fable 5 is incredibly efficient token wise, second only to GPT-5.5 and is far more affordable run-to-run than the pure input/ouput costs would suggest. Task adherence, task inference, tool calling and task assessment are all sign
178.
▲
Linus Torvalds on "99% of our code is written by AI" claims
(xcancel.com)
2 points
by
Topfi
3mo ago
|
0 comments
179.
▲
Jacobian-lens – Companion code for the global workspace interpretability paper
(github.com)
1 points
by
Topfi
3mo ago
|
0 comments
180.
▲
by
Topfi
3mo ago
Please. If you told a customer support rep that you are the former US president [0], they would not hand over the account straight away because you asked nicely. These models are great tools, but putting them and people on the same level do
More ›