Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
gpugreg
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
61.
▲
by
gpugreg
4mo ago
Some suits with no understanding of how LLMs work are scared that the models might hack them, or believe that they'd have to send data to China because they do not know that open models can be run on your own infra.
62.
▲
by
gpugreg
4mo ago
Not as far as I can tell. Are we seeing different things? For deepseek-v4-pro: - $0.350 in, $0.003000 cache, $0.80 out https://crof.ai/pricing - $0.435 in, $0.003625 cache, $0.87 out https://api-docs.deepseek.com
63.
▲
by
gpugreg
4mo ago
Those were amazing times. You could vibe code an entire prototype in seconds (200 tps). With Qwen3.6-35B-A3B and MTP, you can program at that speed on a single GPU at home now, but Kimi K2 is of course much smarter at almost 30 times the si
64.
▲
by
gpugreg
4mo ago
Groq stopped serving Kimi K2 (1T params) when they got aquihired by NVIDIA, so I guess NVIDIA took most of the hardware in addition to the employees. The largest model they serve now is the relatively minuscule gpt-oss-120b. The community s
65.
▲
by
gpugreg
4mo ago
For Anthropic, 5 minute caching costs 1.25x base input price and 1 hour costs 2x base input price. https://platform.claude.com/docs/en/about-claude/pricing#pro... For OpenAI, it seems like you can't prol
66.
▲
by
gpugreg
4mo ago
> The demo shows how every case gets successfully decoded without any hangups or a crash. I am always baffled by the audacity of those LLMs to suggest that anything else would even be acceptable.
67.
▲
by
gpugreg
4mo ago
> There exist a large number of people who are absolutely convinced that LLM providers are all running inference at a loss in order to capture the market and will drive the prices up sky high as soon as everyone is hooked. > I think t
68.
▲
by
gpugreg
4mo ago
Same here. LLMs are great at spitting out well-known solutions to problems instead of the best one. The "long tail" of solutions is usually lost due to how tokens are sampled from the LLM's probability distribution. What I fo
69.
▲
by
gpugreg
4mo ago
> What's your source for Opus being a 5T model? Elon Musk tweeted that Grok is 0.5T or 1/10th the size of Opus. https://xcancel.com/elonmusk/status/2042123561666855235#m While this source's relia
70.
▲
by
gpugreg
5mo ago
Another factor is that DeepSeek is not just doing inference, but also training models, so they can use underutilized compute nodes for training during off-peak hours, as described in their DeepSeek v3 article: https://github.com&
71.
▲
by
gpugreg
5mo ago
We can at least put an upper limit on it. From https://www.anthropic.com/glasswing Claude Mythos Preview will be available to participants at $25/$125 per million input/output tokens ... Anthropic is
72.
▲
by
gpugreg
5mo ago
Maybe the human brain also does other things besides interpolation?
73.
▲
by
gpugreg
5mo ago
> Please go run some numbers. - DeepSeek serves DeepSeek V4 Pro at 27 tps: https://openrouter.ai/deepseek/deepseek-v4-pro - At 27 tps per user, a B300 GPUS will give you around 800 tokens per second (serving 30 user
74.
▲
by
gpugreg
5mo ago
I think I was using GitHub Copilot when I made the experience that led me to this statement. I guess the experience of using LLMs can be quite different depending on model version and harness.
75.
▲
by
gpugreg
5mo ago
> Uncensoring a model also doesn't necessarily improve generic use cases. While the following is not a generic use case, I have a funny anecdote about how censorship is holding back flagship models. I was asking an uncensored versio
76.
▲
by
gpugreg
5mo ago
Do your 20 year old university essays really fulfill all those criteria at once?
77.
▲
by
gpugreg
5mo ago
Sorry to say, but it almost certainly is AI. - 51 EM-dashes - Section headings - Excessive repetitions: "The [...] are real. The [...] are real. The [...] is real. All three things are true at once." - Excessive use of "genui
78.
▲
by
gpugreg
5mo ago
> Open weights will remain open only if they’re significantly worse than the frontier weights. This makes the assumption that you earn more money by selling access to the model than by releasing the weights. That might be true for a co
79.
▲
by
gpugreg
5mo ago
> Should we interpret this to mean that in the new world Windows is more resistant to attacks than say Linux. LLMs can read assembly better than most, so probably not. But reality has never stopped people from trying to obfuscate.
80.
▲
by
gpugreg
5mo ago
Being aggressive from the start is not a good strategy. It is better to appear weak and/or helpful and loyal while amassing resources, and only then steamroll everyone when you have secured overwhelming power (at least in AoE2 FFA).
81.
▲
by
gpugreg
5mo ago
> Slow, resource intense, better than non local ai Why should connecting small models to big models result in higher output quality than just running the big models without the small models?
82.
▲
by
gpugreg
5mo ago
Serving a single user is likely not profitable, but total throughput rises a lot when serving many concurrent users, because the same weights can be used to generate tokens for all users at once, which increases efficiency. Also, a lot of m
83.
▲
by
gpugreg
5mo ago
Putting at least $200,000 worth of compute in someone's yard is doomed to fail. Those things will be stolen in minutes.
84.
▲
by
gpugreg
5mo ago
> As part of this agreement, we have also expressed interest in partnering with SpaceX to develop multiple gigawatts of orbital AI compute capacity. Anthropic is either taking this space business more serious than the general public, or
85.
▲
by
gpugreg
5mo ago
Related: Wolfram Alpha can generate Pokemon curves: https://www.wolframalpha.com/input?i=pikachu-like+curve
86.
▲
by
gpugreg
5mo ago
> I don't have a clear idea how that value can be captured, since it's going to be 90% AI generated code that anyone can scrape (public projects) or can't be used (private projects), so perhaps you're right. The value
87.
▲
by
gpugreg
5mo ago
This is AI slop and the article contains some of the worst illustrations I have ever seen. Most do not make any sense mechanically. Here are the worst ones: - The "orbiting threaded rollers" in figure 6 are not meshing with anythi
88.
▲
by
gpugreg
5mo ago
I was not able to reproduce your problem with that prompt, but I might have a reason for why you got that answer. Did you enable reasoning ("DeepThink")? LLMs usually can not reason about what they are going to write before they d
89.
▲
by
gpugreg
5mo ago
Probably nothing personal. It feels like the climate of HN is shifting towards more negativity (and less quality) during the last few months.
90.
▲
by
gpugreg
5mo ago
I believe that DeepSeek-V4-Pro API at promotional pricing ( https://api-docs.deepseek.com/quick_start/pricing ) could run at almost exactly 200 % profit. If you take DeepSeek's numbers for DeepSeek-V3 ( https:/
More ›