Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
versteegen
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
61.
▲
by
versteegen
6mo ago
They're definitely not subsidizing API pricing, can't believe how prevalent that fallacy is on HN of all places. The question is how profitable Claude Code is. Your example 2 is real and major but your example 1 is ridiculous, alm
62.
▲
by
versteegen
6mo ago
Even more detail in the DW article: """ Fortunately, the boy was very precise and showed me exactly where he found it on a map. Then we went into our findings registration and found that this agricultural site was actually a
63.
▲
by
versteegen
7mo ago
Where do you see that? I only skimmed the prompts but don't see any aspects of any of the games explained in there. There are a few hints which are legitimate prior knowledge about games in general, though some looks too inflexible to
64.
▲
by
versteegen
7mo ago
The dataset miscomparison is a big problem. The prompt is super specific to ARC-AGI-3, which is perfectly fine to do, but skimming it I saw nothing that appears specific to the 25 games in the dataset. Especially considering they've on
65.
▲
by
versteegen
7mo ago
...Their agent is called "Agentica ARC-AGI-3 agent for Opus 4.6 (120k) High". Yes, it's unfair to compare results for the 25 (easier) public games against scores for the 55 semi-private games (scores for which are taken from
66.
▲
by
versteegen
7mo ago
> An AI that can only perform at the average human level is useless unless it can be trained for the job like humans can. Yes, if you want skilled labour. But that's not at all what ARC-AGI attempts to test for: it's testing fo
67.
▲
by
versteegen
7mo ago
> Aren't they losing money on the retail API pricing, too? No, they aren't, and probably neither is anyone else offering API pricing. And Anthropic's API margins may be higher than anyone else. For example, DeepSeek releas
68.
▲
by
versteegen
7mo ago
Ed Zitron made that claim (in particular here: [1]). In the same article he admits he not a programmer, and had to ask someone else to try out Claude Code and ccusage for him. He doesn't have any understanding of how LLMs or caching wo
69.
▲
by
versteegen
7mo ago
I'm surprised "faulty PSU" is not on GP's list of common problems. Almost every unstable computer I've ever experienced has been due to either a dying PSU (not an under-specced one) or dying power conversion capacit
70.
▲
by
versteegen
7mo ago
AFAICT, Claude was not asked to prove its algorithm works for all odd n, but was instead told to move on to even n.
71.
▲
by
versteegen
7mo ago
> Gemini 2.5 Pro's reasoning traces (before they nerfed them) were a good example. The deep technical analysis, and then the human-friendly version in the final output. But I found their reasoning more readable than the final output
72.
▲
by
versteegen
7mo ago
Yes, .wow also.
73.
▲
by
versteegen
7mo ago
That misses my point: the evidence is the extensive argumentation provided for why it reduces risk. To quote Karnofsky: > I wish people simply evaluated whether the changes seem good on the merits, without starting from a strong presumpt
74.
▲
by
versteegen
7mo ago
Yes. And I did port my GUI layer to CimGui.jl. The rest of it is pretty intertwined with Makie, didn't do that yet. The Makie version does look better than ImGui though.
75.
▲
by
versteegen
7mo ago
I recently used Makie to create an interactive tool for inspecting nodes of a search graph (dragging, hiding, expanding edges, custom graph layout), with floating windows of data and buttons. Yes, it's great for interactive plots (you
76.
▲
by
versteegen
8mo ago
> They pragmatically changed their views of safety just recently, so those values for which they would burn at the stake are very fluid. Yes it was a pragmatic change, no it was not a change in their values. The commentary here on HN abo
77.
▲
by
versteegen
8mo ago
Sure, depending on the particular product, having control and direct local access to the data would be desirable or deal breaker. But for this Clickup integration that's not so important to us (we can duplicate where necessary), while
78.
▲
by
versteegen
8mo ago
I'd agree that this effect is probably mainly due to architectural parameters such as the number and dimensions of heads, and hidden dimension. But not so much the model size (number of parameters) or less training. Saw something about
79.
▲
by
versteegen
8mo ago
Agreed, and here's a real example from a tiny startup: Clickup's web app is too damn slow and bloated with features and UI, so we created emacs modes to access and edit Clickup workspaces (lists, kanban boards, docs, etc) via the
80.
▲
by
versteegen
8mo ago
> The only valid ARC AGI results are from tests done by the ARC AGI non-profit using an unreleased private set. I believe lab-conducted ARC AGI tests must be on public sets and taken on a 'scout's honor' basis that the lab
81.
▲
by
versteegen
9mo ago
Paid/API LLM inference is profitable, though. For example, DeepSeek R1 had "a cost profit margin of 545%" [1] (ignoring free users and using a placeholder $2/hour figure H800 GPU, which seems ballpark of real to me due t
82.
▲
by
versteegen
9mo ago
It's excellent that you're working on loneliness! Somehow. What is it your startup actually does?
83.
▲
by
versteegen
9mo ago
Yes, unfortunate that people keep perpetuating that misquote. What he actually said was "we are not far from the world—I think we’ll be there in three to six months—where AI is writing 90 percent of the code." https://w
84.
▲
by
versteegen
9mo ago
That's pretty good, I'm jealous! The last time I reinstalled my OS (Slackware) from scratch was 2009, but I run into serious problems every couple of years when upgrading it to 'Slackware64-current' pre-release, because
85.
▲
by
versteegen
10mo ago
Then what do you say to 6.14" × 2.61" × 0.11 mm = 102 cm³
86.
▲
by
versteegen
10mo ago
Yes, it's a great deal especially because you get access to such a wide range of models, including some free ones, and they only rate limit for a couple minutes at a time, not 5 hours. And if you go over the monthly limit you can just
87.
▲
by
versteegen
10mo ago
I have no idea at all whether the GCP "Service Specific Terms" [1] apply to Gemini CLI, but they do apply to Gemini used via Github Copilot [2] (the $10/mo plan is good value for money and definitely doesn't use your da
88.
▲
by
versteegen
10mo ago
That's an interesting observation. I'd suggest modelling the LLM's behaviour in that situation as selecting between different simple strategies, each of which has its own transition function. Some of the strategies will be fa
89.
▲
by
versteegen
10mo ago
Since it took me some minutes to find the description of the task, here it is: We conducted experiments on three different models, including GPT-5 Nano, Claude-4, and Gemini-2.5-flash. Each model was prompted to gener- ate a new word based
90.
▲
by
versteegen
10mo ago
One interesting and ironic part of the article is that one of the mentioned optics research groups has been submitting a lot of patents on EUV sources. Are we meant to be mad about it?
More ›