Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
nsingh2
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
nsingh2
4mo ago
Yea these plots are too noisy and dense. Especially that second one, lines all over the place.
32.
▲
by
nsingh2
4mo ago
From what my own experiences are, and what's on their checkout page, $100 is 5x base usage and $200 is 20x. If $100 was 10x, then I personally would drop down. They want people to go to the highest tier.
33.
▲
by
nsingh2
4mo ago
I'm really getting sick of reading about safeguards and what I'm not allowed to do on every model release.
34.
▲
by
nsingh2
4mo ago
To follow up on this, I had it solve a nasty ODE problem that I saw in the recent Mathematica 15 release post: Solve the following first-order ODE for f(x): ((-1 - 2*x)*f(x)*tan(1 + x - exp(-61 - 2*x)*f(x)/x) + exp(61 +
35.
▲
by
nsingh2
4mo ago
Lots of confusion about what this model is actually focused on. It is a cheap specialist for closed-world, verifiable reasoning tasks like math, self-contained coding problems, and similar. "Closed-world" means the needed informat
36.
▲
by
nsingh2
4mo ago
The lack of tool use will hinder it a lot I think, since bug hunting requires collecting context across a code base and stitching it together. It might be good in a more narrow sense, i.e "is there a bug in this block of code" and
37.
▲
by
nsingh2
4mo ago
This model doesn't support tool calling, was not part of its training. It's focused on Python (and I think C++) competitive programming and mathematics tasks, i.e. tasks with verifiable rewards. So if you have a task that fits tha
38.
▲
by
nsingh2
4mo ago
No one really goes into an interview speaking about politics, or propaganda, or unionizing. So there isn't much signal in telling people to not do that, most already bend over backwards to look like an ideal candidate. Telling people t
39.
▲
by
nsingh2
4mo ago
How does this have anything to do with a sluggish job market? Do you think people are struggling to find jobs because they are too political? One can be a servile dog, and still struggle to land a role, it's not particularly correlated
40.
▲
by
nsingh2
4mo ago
Could you provide some details, if possible, like what model & thinking effort, what kinds of tasks? I used to swap between Claude Code and Codex often, and these days use Codex more because of the usage limits. Wondering if I should go
41.
▲
by
nsingh2
4mo ago
Yup, I meant to write quantization there.
42.
▲
by
nsingh2
4mo ago
This isn't about training on the output tokens from Anthropic models, it's just about using their models to build things like pretraining pipelines, etc. Even if you train on your own data. From the phrasing, it might as well be t
43.
▲
by
nsingh2
4mo ago
It's such an obviously bad policy, it's mind-boggling that they thought this was a good idea. It just breeds paranoia and mistrust, especially when people are already a bit paranoid about silent model quantification for cost cutti
44.
▲
by
nsingh2
5mo ago
Works for me, Firefox 151.0 on Linux
45.
▲
by
nsingh2
5mo ago
Proof search isn't new, but I don't think that captures the value of LLMs. They act as a learned proposal mechanism on top of hard search. Things like suggesting relevant lemmas, tactics, turning intent into formal steps, and rank
46.
▲
by
nsingh2
5mo ago
I've been using GPT-5.4, and more recently 5.5, with Codex CLI + Ghidra MCP for reverse engineering a game without many issues. Injecting code is where it usually balks at, but I'm just trying to discover and parse structures from
47.
▲
by
nsingh2
6mo ago
> a business that puts employees first and profits for owners last can often have a shit ton of profits for owners. Owners can make 100x that shit ton if they put profits for owners first, so why wouldn’t they do that instead? Out of the
48.
▲
by
nsingh2
6mo ago
> The solution, if there is one, has to come from innovation from the private economy. Why? The problems of offshoring, consolidation, automation, you described came from private sector incentives (not to mention debt driven consumption,
49.
▲
by
nsingh2
6mo ago
These plots are terrible. Why is categorical data connected across categories with lines? Why not just use bar plots? Like in the "Web Vulns in OSS" plot, white box data for Opus 4.7 is not available, but the absurd linear interpo
50.
▲
by
nsingh2
6mo ago
It's a combination of factors. There was rate-limiting implemented by Anthropic, where the 5hr usage limit would be burned through faster at peak hours, I was personally bitten by this multiple times before one guy from Anthropic annou
51.
▲
by
nsingh2
6mo ago
> https://www.nature.com/articles/s41586-025-09797-z That title reeks of the paper equivalent of clickbait. The paper is about subjective well-being and mental health in the psychological sense. Broader well-being i
52.
▲
by
nsingh2
6mo ago
This kind of sentiment, on its own, is hollow. Just more "violence bad", until the next round. There is growing anger and discontentment in a large part of the population, driven by inequality of wealth and power. Hopelessness and
53.
▲
by
nsingh2
6mo ago
It's going to be expensive to serve (also not generally available), considering they said it's the largest model they've ever trained. I suspect it's going to be used to train/distill lighter models. The exciting pa
54.
▲
by
nsingh2
6mo ago
> They were years ahead. Considering how fast competitors caught up to them, I'm not convinced that OpenAI was years ahead. LLMs and transformers were known technology, it's just that OpenAI accidentally productized it before o
55.
▲
by
nsingh2
6mo ago
Seems like HN is doing something to combat this, considering how many [dead] comments I see in every post (which you can enable by setting `showdead` in your user profile). I've only recently enabled it so I don't know how frequen
56.
▲
by
nsingh2
6mo ago
Really uncharitable take. I did stupid things at 14, and had more unrestricted internet access too. > absent parent more concerned with his business than his son I don't know how you came to this conclusion from the post.
57.
▲
by
nsingh2
6mo ago
This a big exaggeration. Codex is probably one of the top two LLM programming tools, along with Claude Code. GPT-5.4 models are strong, unlike the initial GPT-5 ones, which were comparatively bad, and can hold up against Opus 4.6. In my exp
58.
▲
by
nsingh2
7mo ago
Gemini CLI has been broken for the past 2-3 days, with no response from Google. Really embarrassing for a multi-trillion dollar company. At this point Codex is the only reliable CLI app, out of the big three. https://www.reddit.c
59.
▲
by
nsingh2
7mo ago
This morning I hit 100% 5hr usage on a task that took ~10% in the past. Looks like they are still testing the limits, but it seems over-tuned to me. Also not great that they communicate this now, since people have been complaining about sud
60.
▲
by
nsingh2
7mo ago
>> More free time? > Yes! Time we can reclaim from the mundane chores of life to do with as we choose! How could you not want that? We already had a huge productivity boom these past decades, but wages flat-lined and the vast major
More ›