Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
gpugreg
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
gpugreg
1mo ago
I tried it and the experience was not great. I first enabled chrome://flags/#force-enable-webgpu-interop which did nothing. Next, I enabled chrome://flags/#enable-unsafe-webgpu which made some WebGPU demos work
32.
▲
by
gpugreg
1mo ago
Not all their research, but certainly a lot: https://github.com/orgs/deepseek-ai/repositories?q=sort%3Ast...
33.
▲
by
gpugreg
1mo ago
The kettlebell in this [1] image looks a bit like the one [2] that was recalled due to radioactivity. It's probably a different one, but I thought I should mention it anyway. Better safe than sorry. [1] https://greenlightnin
34.
▲
by
gpugreg
1mo ago
I could have used more precise terminology. rfind is average case O(n + m), worst case O(n * m). Imho the worst case performance is more important than the average case performance, since it tells us whether there is any risk for attack
35.
▲
by
gpugreg
1mo ago
Because the use case is very niche and nobody optimized it yet. https://github.com/python/cpython/issues/135824#issuecomment... `x in range(n)` is already optimized, but that was easier since the `__contains_
36.
▲
by
gpugreg
1mo ago
Notable pitfalls: - s[i:j] is O(j - i) because it creates a copy instead of a view - max(range(n)) is O(n) - substring search is O(n), which is good, but rfind is O(n m) - iterative string concatenation (for c in ...: s += c) can be O(n^2)
37.
▲
by
gpugreg
2mo ago
Sure! But where?
38.
▲
by
gpugreg
2mo ago
To learn about sentiment analysis, I'd look for related datasets and then look at recent code, e.g. here: https://www.kaggle.com/datasets?search=sentiment+analysis For more LLM-specific stuff, you can pick some agent t
39.
▲
by
gpugreg
2mo ago
Agents usually start with ingesting the existing code base, and DeepSeek can use those code bases for pretraining. And they will have filters on top of that to throw out garbage. I am not sure how they are using the data for post-training,
40.
▲
by
gpugreg
2mo ago
Oh, I messed up. Half-way through, I thought it would be a good idea to double the numbers so I don't have to deal with half millions, but forgot to also double the 98.5. Unfortunately, I can not edit it anymore. I think the margins of
41.
▲
by
gpugreg
2mo ago
Agentic workloads are somewhere around 1%/0.5%/98.5% input/output/cached tokens. Cached tokens are pretty much free for inference providers (if they implement sparse and compressed attention properly) and throughput for
42.
▲
by
gpugreg
2mo ago
Personally, I prefer QDirStat. I just tried to use FileLight to compare, but the package seems to be broken on Lubuntu.
43.
▲
by
gpugreg
2mo ago
> you're having issues handling files properly? I guess they were using ollama, which does not tell you where it puts the models it downloads.
44.
▲
by
gpugreg
2mo ago
I get the following error: Traceback (most recent call last): File "/app.py", line 1, in <module> import spider ModuleNotFoundError: No module named 'spider' Steps to reproduce: 1. V
45.
▲
by
gpugreg
2mo ago
I did not say that it is impossible. I just think that we need architectural improvements, or maybe even a fundamentally different approach to get something like Kimi K3 for cheap. The point I was trying to make was that we shouldn't j
46.
▲
by
gpugreg
2mo ago
Looks like global energy consumption has risen by an order of magnitude from 1900 to 2000: https://www.encyclopedie-energie.org/en/world-energy-consump... Unfortunately, electricity prices did not fall by the same fact
47.
▲
by
gpugreg
2mo ago
> Some of us might be rich I sure wish I had a few 100M of disposable income to train a frontier model. > or in the future it could be useful when training is cheaper. I do not think that physics will allow hardware getting that much
48.
▲
by
gpugreg
2mo ago
Where do you see less than $10/h for 8 * MI354X? I can only find $2.50 for 1 * MI355X (lowest I can find for rent on other websites is $2.65, but maybe they got a better deal).
49.
▲
by
gpugreg
2mo ago
Notably, MXFP4 was introduced at the (much less costly) supervised fine-tuning stage after pretraining, so the number of B200/B300 GPUs could be relatively small in comparison to the number of H200 GPUs used during pretraining (or ma
50.
▲
by
gpugreg
2mo ago
> they also mentioned only having a 20K GPU cluster (unclear if NVIDIA, or Huawei). A few quotes from the transcript: > Our current computing capacity is approximately 20,000 H-equivalent units, most of which have just arrived within
51.
▲
by
gpugreg
2mo ago
dax (coauthor) recently tweeted https://xcancel.com/thdxr/status/2083178051052155182 > because we added the new deepseek which we do not yet have a ZDR with we cannot blanket say we offer ZDR I wonder how the w
52.
▲
by
gpugreg
2mo ago
The memory bandwidth of the 2x RTX Pro 6000 Blackwell setup will be 10x higher, which should have an equivalent effect on the generated tokens per second.
53.
▲
by
gpugreg
2mo ago
Unfortunately, all mentions of ZDR have silently been removed from the OpenCode Go page today.
54.
▲
by
gpugreg
2mo ago
To add to this, the $60 only applies to DeepSeek-V4-Flash and a few other models. For DeepSeek-V4-Pro, the amount is $15. https://opencode.ai/docs/go/#usage-limits Previously, OpenCode Go had higher API prices for
55.
▲
by
gpugreg
2mo ago
According to the leaked call transcript, DeepSeek is working on vision for V4. Not sure when it will land though.
56.
▲
by
gpugreg
2mo ago
I think this is less about winning the AGI race and more about not dropping out of AI entirely.
57.
▲
by
gpugreg
2mo ago
The model is already natively MXFP4-quantized during training, so there is no quality loss.
58.
▲
by
gpugreg
2mo ago
Brussels is set to contribute roughly €5 billion, matched by another €5 billion from European governments, alongside around €20 billion in private investment. You can train a model like DeepSeek-V4-Pro with around $2 billion. The
59.
▲
by
gpugreg
2mo ago
From the Kimi K3 technical report: Kimi K3 supports a context window of up to 1 million tokens. We achieve this through extending the context window progressively as training proceeds, following a four-stage curriculum. The wi
60.
▲
by
gpugreg
4mo ago
Anthropic probably trained Mythos on their own code and found that it is too got at reproducing it.
More ›