Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
npn
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
npn
2mo ago
I don't think so. there is a lot of tools with similar usage, some harness even bring their own internal tools for accurately manipulation. also, even if some models claim that they have full 1M context window, some only work effective
32.
▲
by
npn
2mo ago
wait for Deepseek Harness (yes it is the official name) release then try again. for your kind of task, harness tools matter.
33.
▲
by
npn
2mo ago
I still believe this is not the full potential of pro models. I expect they will release another checkpoint later this year.
34.
▲
by
npn
2mo ago
well, I think it is easier for the developers, we don't have to change the model name everytime. deepseek also does not have resource to serve 2 different v4 pro versions.
35.
▲
by
npn
2mo ago
I think both are after qwen 3.8 max release.
36.
▲
by
npn
2mo ago
thank you very much for the input. I once said to my friend that "no actually you do not get $20k worth of API usage for $200 subscription, it is mostly $600". turn out it was a pretty close guess. even if anthropic does not apply
37.
▲
by
npn
2mo ago
I suggest you make two other calculations: 1. Only calculate the output cost, skip input cost 2. Only calculate the cached input cost You will be surprised
38.
▲
by
npn
2mo ago
Agreed. I think people tend to forget that: you do not have to pick a single language. it is trivial to use another language just for that single library. there are many way to communicate, it is not like method calling is always the hard r
39.
▲
by
npn
2mo ago
this is just survivorship bias. I bet you 1000 dollars that there exists hundred languages that are more useful and well thought than those languages you listed, but still nobody cared and it died. I believe the reason lies entirely outside
40.
▲
by
npn
2mo ago
weak argument. deepseek v4 flash is open weight, you can easily find other providers with competitive price with Deepseek (except for input caching), some even half as cheap.
41.
▲
by
npn
2mo ago
From the private conversation of deepseek CEO and the investors, I'm under impression that they take pride for the low API price though. So it is not only training data. If they really want the data there are many other ways to do so.
42.
▲
by
npn
2mo ago
btw, GA usually means it is suitable for production, and available globally. the last time a gemini model was available globally in multiple datacenters was gemini 2.0 era.
43.
▲
by
npn
2mo ago
Similar to aistudio vs vertex. Or rather, Alibaba cloud is equivalent to google cloud or aws.
44.
▲
by
npn
2mo ago
More like it is a generic idea that has been implemented 1000 times before so LLM already have the perfect solution. Well, not like I'm saying the product is doomed though, because like always implementation details matter.
45.
▲
by
npn
2mo ago
weird, for so long I'm damn sure that researchers are obligatory to release research papers to fulfill their quotas. otherwise they will lost their titles. so unless those companies (openai, anthropic) gave the researchers some fortune
46.
▲
by
npn
2mo ago
well usually you can just generate the data using LLM. use 2 or 3 different frontier models from different providers, then compare the results and pick the consensus. yes it is not 100% correct like when you make it manually, but then again
47.
▲
by
npn
2mo ago
There is a list of a few websites that I subconsciously think only have garbage articles and never really pay serious attention to them: quora, medium and substack. Surely dev.to also filled with low level articles, but at least people ther
48.
▲
by
npn
2mo ago
fine tuning a small LLMs or even real small language models (like bert) is what the recommended way since the introduction of LLM. the benefit is pretty much obvious: faster to run, fully controlling the stack, better fit for the custom dom
49.
▲
by
npn
3mo ago
even deepseek, with their current (dirt cheap) price, can earn enough profit to cover the cost (hardware investment?) in 10 months.
50.
▲
by
npn
3mo ago
Huawei chips need advanced 3d packaging in order to keep up. Since the process is too complex the yield is still bad. Not to mention there is a lot of demand from various factors, not deepseek only. Huawei itself is a major consumer.
51.
▲
by
npn
3mo ago
They spend $40mil for lobbying, I'm sure they can also spare some millions to this place (and other places like reddit). They all do.
52.
▲
by
npn
3mo ago
AI means it has npu, Max+ is marking the memory channel, PRO is a normal label for chips that have extra security baked in, it has been this way since forever. I don't see anything wrong with the name.
53.
▲
by
npn
3mo ago
update: the knowledge cut off date is "unknown" now. funny because some people downvoted me believed that there is no relation between knowledge cut off date and real world events. that's not how it works!
54.
▲
by
npn
3mo ago
tested the models on aistudio. despite that the knowledge cut off is march 2026 it still knows nothing about 2025! you can check by asking "list notable world events in 2025, only list unplanned" on aistudio. or you can ask for Ch
55.
▲
by
npn
3mo ago
what a horrible article. full of misinformation and dishonesty. 1. training new base models are expensive for sure, but fine-tuning them are relatively inexpensive enough the labs can continue to do so forever. the main reason why frontier
56.
▲
by
npn
3mo ago
Did they cry about it? No right? Don't apply your own standard then judge them about it, petty people. Stackoverflow aimed to be a knowledge base. And knowledge base has a ceiling limit. They simply reached the point that almost all qu
57.
▲
by
npn
3mo ago
Not worth it. I have just tried a single prompt in the web interface and it is still not finish reasoning. It thinks too much and often repeats the same stuff over and over. Combine with the price it will surely more costly than gpt 5.6.
58.
▲
by
npn
3mo ago
it is funny because nobody ever bother points out that they overcharge you for text input token price. sure it was pretty resource intensity a few years before, but with turbo quant, sparse attention and various techniques, plus the advanci
59.
▲
by
npn
4mo ago
That's for the long term. Anthropic only needs short term solutions for the sake of IPO. They will do whatever they can to sabotage other companies (specially the Chinese ones) to reach the same parity with best claude models.
60.
▲
by
npn
4mo ago
I doubt you can do that. MTP magic happens because for texts, we have a lot of low value fixed tokens that almost always get generated in the sequence (like punctuation, function words, language keywords etc). for most important ones (the e
More ›