Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
throw10920
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
61.
▲
by
throw10920
2mo ago
This is such a ridiculous and shallow cop-out. And it also is completely irrelevant to my challenge to show how proprietary benchmarking can be gamed, because it presumes (absolutely insane and divorced from reality) circumstances that have
62.
▲
by
throw10920
2mo ago
> Lots of praise for Pi in this thread, so I'll offer up a diverging opinion. Being reflexively contarian is not what HN is for. https://news.ycombinator.com/item?id=45530593 > For a program that's minimal i
63.
▲
by
throw10920
2mo ago
Please imagine that the benchmark runners are not making mistakes of the fourth grade level - which they won't be. If you assume this level of incompetence, then literally everything is possible.
64.
▲
by
throw10920
2mo ago
Artificial Analysis was an clearly meant to be an example. I was obviously talking about the ideal scenario of proprietary benchmarking, not how it might be being screwed up in practice.
65.
▲
by
throw10920
2mo ago
> Are subagents basic? I found them only useful in very few situations. I've found them to be extraordinarily helpful, because they allow me to much more carefully control context and reduce token spend by using a smart model for
66.
▲
by
throw10920
2mo ago
> e.g. sentiment analysis when the user begins cursing at the agent, or checking whether the user continued another session with the generated code, or started a new session with the same starting point as before, i.e. they git-stashed.
67.
▲
by
throw10920
2mo ago
> Proprietary or open didn't matter, because with enough attempts at something you can eventually sus out its operations and optimize accordingly (or distill, as we've seen with LLMs). Please explain how, if I'm OpenAI and
68.
▲
by
throw10920
2mo ago
> Building a new benchmark won’t solve the problem, either. It will if the benchmark is proprietary. If you can't train on it, then it's extremely difficult to game, and if it's hard enough, then it's economically mor
69.
▲
by
throw10920
2mo ago
It's AI-generated, which is against the guidelines: > Don't post generated text or AI-edited text. HN is for conversation between humans.
70.
▲
by
throw10920
2mo ago
> should not discount that DeepSeek also gets paid in data, which is probably more valuable to them That's agentic feedback loops for training, right? Any more detail on this, such as how they actually tell whether that data is goo
71.
▲
by
throw10920
2mo ago
This doesn't have anything to do with Google or OpenAI. I'm not Ilya Sutskever and it's not 2017, either. I'm not "whining" about anything. You made the claim "The Chinese have actually been very open abou
72.
▲
by
throw10920
2mo ago
> You can replicate the architectural innovations, and try them for yourself with your own dataset. That's not related to my comment. My comment was pointing out that you can't verify something that wasn't published. You h
73.
▲
by
throw10920
2mo ago
That's generally true, but consider this: If the consolidated (stolen) IP of millions of books remains under US LLM company control, then at the very least there's a chance of justice to be served to the creators in the future -
74.
▲
by
throw10920
2mo ago
> The Chinese models are mostly very well documented in terms of architecture and training processes/flows, with what is missing to recreate them being the training data. ...and because that training data is missing, they can't
75.
▲
by
throw10920
2mo ago
The Reddit comments are both older.
76.
▲
by
throw10920
2mo ago
> In-context, your reply reads like you're asserting that anthropic's offerings give you all the same control as a local model It absolutely does not.
77.
▲
by
throw10920
2mo ago
> "Almost like magic" is cliché hyperbole. It also means "not magic" in the same way that "almost like a dog" means "not a dog." It was immediately followed by the explicit clarification that was s
78.
▲
by
throw10920
2mo ago
I'd rather just use my actual human brain to compute the answer at that point. I don't see the value at throughput that is this low.
79.
▲
by
throw10920
2mo ago
> Sessions can be accessed through the web and native macOS/iOS applications ... only native macOS/iOS applications. Not a good look.
80.
▲
by
throw10920
2mo ago
That might be true, but it's irrelevant to the current discussion, which is arguing about whether or not Chinese models can do this now . One of the commentators above is arguing that it's not only possible but significantly ch
81.
▲
by
throw10920
2mo ago
Huh. Reading your comment and the parent above - maybe retreats aren't "therapeutic", but "exacerbatory" - as opposed to making interpersonal relationships uniformly better (bad -> good and good -> better), th
82.
▲
by
throw10920
2mo ago
I've had repeated conversations with Opus about cybersecurity and never gotten a refusal.
83.
▲
by
throw10920
3mo ago
> Capitalism created the computers whose Soviet copies were Tetris' native platform, it made the languages Pajitnov used and the machines he ported it to. It incentivized people to bring Tetris to the rest of the world, and who know
84.
▲
by
throw10920
3mo ago
> if china can train K3 on a fraction of the US compute availability, yet it benchmarks almost equivalent to Fable for a third of the cost, it’s game over for US labs in the long run. I agree. However, as of yet, most/all leading PR
85.
▲
by
throw10920
3mo ago
Some PRC models are backdoored to silently insert extra vulnerabilities when certain conditions are met - https://www.boozallen.com/expertise/cybersecurity/whats-in-a... And that's just the model behavior. Th
86.
▲
by
throw10920
3mo ago
> I can't think of a game studio the size of Riot that is as community-engaged and pro-consumer Valve. This is also just a crazy statement to make. > Examples range from the quality of their patch notes, dev articles and videos,
87.
▲
by
throw10920
3mo ago
> I assume you feel the same about R6, TF2 and CSGO, naturally. It seems like you think you've pulled a "gotcha" on me? I don't play those games. I don't have any feelings about them. And they're not relevan
88.
▲
by
throw10920
3mo ago
> whether your mom thinks you are a moron or not > Yeah, not really, unless your mom is a party member perhaps? > Yeah - saw you playing in the schoolyard, and thought you looked lonely. You're a middle-schooler. I've dis
89.
▲
by
throw10920
3mo ago
Can you please add an email to your HN profile?
90.
▲
by
throw10920
3mo ago
> why GLM has two personalities I never said that. The fact that you have to compulsively lie about my words is...funny. Most people grow out of this in middle school, you know. > go to https://chat.z.ai/ and ask it A
More ›