Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
GodelNumbering
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
GodelNumbering
2mo ago
I think that's a fair offering tbh
32.
▲
by
GodelNumbering
2mo ago
If there were, do you believe it would be in their interest to answer this publicly?
33.
▲
by
GodelNumbering
2mo ago
Correct, but it was preview release. I was referring to GA in my comment.
34.
▲
by
GodelNumbering
2mo ago
This is one of the interesting aspects the 'AI job loss' community doesn't account for. As the technology unlocks things, more startups are created. And even at a lower nominal engineer-to-work ratio, overall demand for talen
35.
▲
by
GodelNumbering
2mo ago
So, in last several months, all the prominent names Google lost: Demis Hassabis (technically still with google but these things are usually presented with a spin), Jeff Dean, Sanjay Ghemawat, Oriol Vinyals, Quoc Le, Noam Shazeer, John Jumpe
36.
▲
by
GodelNumbering
2mo ago
I disagree to the maximum extent possible. I have lived in 6 countries. Americans to me are by far most friendly and supportive of new ideas. When you are visiting a place, you see it with rosy eyes. You are not forced to deal with the day
37.
▲
by
GodelNumbering
2mo ago
I have been writing a 'paper' [1] on an adjacent topic for months now. At some point, I decided to make it an empirical paper vs position paper. I am still chasing the experiments (when I get some free time waiting for agentic loo
38.
▲
by
GodelNumbering
2mo ago
Thanks, I will reach out. I have also posted for contributors on localllama https://www.reddit.com/r/LocalLLaMA/comments/1vg40w8/anyone_... > We can do a mix of general use (as in user stories) plus a
39.
▲
by
GodelNumbering
2mo ago
I can manage the infra, have a lot of experience in that area. A benchmark with problems coming from multiple sources and backgrounds would be ideal
40.
▲
Ask HN: Anyone interested in building a harness-only benchmark?
5 points
by
GodelNumbering
2mo ago
|
5 comments
41.
▲
by
GodelNumbering
2mo ago
No, server-side compaction is inherently tied to the model's current context. This uses a full side prompt to get the compaction result, then swap the context with new compaction result. I did a bunch of tests and the results from Luna
42.
▲
Token Arbitrage: Use Sol for coding, Luna for compaction, save 84%
(dirac.run)
3 points
by
GodelNumbering
2mo ago
|
2 comments
43.
▲
Claude gen-5 models show significant regression in BullshitBench
(github.com)
3 points
by
GodelNumbering
2mo ago
|
0 comments
44.
▲
by
GodelNumbering
2mo ago
"Half the money I spend on advertising is wasted; the trouble is I don't know which half." -John Wanamaker This applies even more strongly to model choosing. I know for a fact that majority of my work doesn't require a
45.
▲
by
GodelNumbering
2mo ago
In the first paragraph, > Anyone who has read my past writing should know that I don’t regard such bans as a useful measure, Later (on banning chip sales to china) > we should crack down on the rampant smuggling and workarounds used t
46.
▲
by
GodelNumbering
2mo ago
> And hire 2 or 3 dev ops to keep it running Not a devops but I'd say one full time is already too many.
47.
▲
by
GodelNumbering
2mo ago
Back of the envelope calculation (could be off, correct me if I am) If you are a large enough company that spends million+ on inference a month, it makes sense to buy a GB300 rack ($6M on top range from what I could find) which has 20.7 TB.
48.
▲
by
GodelNumbering
3mo ago
I posted earlier on reddit that at this point, the closed source lobby wanting to ban open weights is effectively outgunned (musk also publicly supported this). The oddly reminds me of back in the day when SOPA had caused a similar stir ( h
49.
▲
by
GodelNumbering
3mo ago
Proprietary model weights are IP, their outputs are not IP and I can't imagine a court decision that would rule otherwise because that would set an extremely far reaching and dangerous precedent, even for American businesses. I don
50.
▲
by
GodelNumbering
3mo ago
> so you end up having to write and maintain complex rulesets about what specific resources agents have access to. Why wouldn't you simply use the group chat ID (+more) to key the conversations on?
51.
▲
by
GodelNumbering
3mo ago
No, it's a fair question. The answer: I believe(d) in Demis Hassabis's vision of AI Although, I think saying 'wanted them to win' was not accurate. More like, I stuck with them hoping it will get better, and it did get b
52.
▲
by
GodelNumbering
3mo ago
> I was a big proponent of Google and Gemini, but they left us reeling with their abrupt product decisions. Likewise. This seems like a common feel. I have at least spent $4000 and likely a lot more on Gemini API because I really wanted
53.
▲
by
GodelNumbering
3mo ago
> undeniably a little cool that luxuries got cheap, but it's severely un-cool that necessities got expensive Well said. Being told that 'you should be happy without owning' to a party that can't afford to own by anoth
54.
▲
by
GodelNumbering
3mo ago
There was a youtube video sometime ago that argued about how life has vastly improved in every single aspect (with numbers about how 'stuff' has gotten universally more affordable), only housing has gotten more expensive relative
55.
▲
by
GodelNumbering
3mo ago
This was my first thought too. I actually use this in my Agent prompts
56.
▲
by
GodelNumbering
3mo ago
Exactly 4 months ago, the marketshare on openrouter was 60%-40% in favor of closed models. Now it's 63%-37% in favor of open models. On March 19th, the open models processed 888B tokens in aggregate, yesterday, they processed 4.19T tok
57.
▲
by
GodelNumbering
3mo ago
That was the first thing I Ctrl+F'd in the paper, no results haha Broadly, I keep thinking about this over last year or two: while LLMs have nearly eliminated the bar for slop and coding slop, the reviewers are still expected to perfor
58.
▲
by
GodelNumbering
3mo ago
> "Finding 1: Scale Buys Evaluation, Not Control" The attached paper's ( https://arxiv.org/pdf/2604.16009 ) title is "MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition" T
59.
▲
by
GodelNumbering
3mo ago
I am with you in the spirit of openweights but I am trying to hard-avoid bringing countries into this. The narrative of US vs China only benefits those who want regulatory capture in the US since attacking China is politically much easier t
60.
▲
by
GodelNumbering
3mo ago
The link has 6 well-known benchmarks where this beats Fable (out of 14 I counted). If the numbers hold up scrutiny, this is scary good. Forget about their pricing but the companies that do have means to host such models fully on-prem are al
More ›