Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
fspeech
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
91.
▲
by
fspeech
2y ago
That runs against basic financial reality. The treasuries are the basis of all U.S. liquidity. Owning commercial papers or stocks doesn't help as companies own treasuries. Putting deposits in banks won't help either because banks
92.
▲
by
fspeech
2y ago
Why is it a bad thing if they now like each other better (not sure if there is real evidence for it but assume it is true for argument's sake)? If American policy is what keeps them apart, which I strongly doubt, then it is a good thin
93.
▲
by
fspeech
2y ago
How do you keep it stationary? How do you keep it out of the Earth's shadow a.k.a. night? How do you keep the transmission loss low and receiving station footprint small? How do you avoid harming things that could come into your beams?
94.
▲
by
fspeech
2y ago
It shows up so naturally that one may think it is not whataboutism at work, even though the reasoning is often not well argued. We could choose to see it not as a moral judgement, but as questioning the justification of efforts pushing for
95.
▲
by
fspeech
2y ago
They already can freely travel to and work on the mainland.
96.
▲
by
fspeech
2y ago
Unlikely as Canada ultimately depends on the US as a security provider, but yes, like Vietnam, it makes sense to trade with both sides. It's also why the US doesn't like it. China's tax on rapeseed oil is like Canadian tax on
97.
▲
by
fspeech
2y ago
It may start sooner than you think. Now we will see how Missouri attempts to collect their $24 billon default judgement against China from a federal court for treble damage of "$8.04 billion in direct lost tax revenue" due to Chin
98.
▲
by
fspeech
2y ago
100% tax on EV imports and 25% metal tax imposed on China by Canada about half an year ago. Canada is in a tough spot. It needs to sync up its tade policy with the US in order to maintain its free trade with the US.
99.
▲
Jim Haslam: Covid Origins and Coronavirus Genetic Engineering [video]
(youtube.com)
3 points
by
fspeech
2y ago
|
1 comments
100.
▲
by
fspeech
2y ago
Matrix absorption is unnecessary. What is needed is the order of multiplication associates towards the direction of the absorption. This and the modified Rope are needed to make the caching work.
101.
▲
by
fspeech
2y ago
The drive so far has been to scale up to push performance, not to reduce cost. Although everyone offers cheaper versions of models they don't base their marketing on these. Now that the momentum of scaling up on training is slowing dow
102.
▲
by
fspeech
2y ago
https://github.com/deepseek-ai/open-infra-index/blob/main/20... Statistics of DeepSeek's Online Service All DeepSeek-V3/R1 inference services are served on H800 GPUs with precision consistent w
103.
▲
by
fspeech
2y ago
In addition Deepseek just showed that they are serving all that demand (admittedly inadequately) from a cluster of only 2000 H800 GPUs.
104.
▲
by
fspeech
2y ago
Gross margin in theory. They are using API pricing to project revenue but not all traffic is API. They don't charge for chat traffic so this is theoretical. It is in response to a raging debate on whether their gross margin is negative
105.
▲
by
fspeech
2y ago
There are a lot of shared prefixes from user prompts. You can save by first looking into the cache to find the longest prefix for a prompt. Their MLA makes the KV cache particularly efficient.
106.
▲
by
fspeech
2y ago
It's probably more about the size of the firm. Once a company is large enough with middle layers and silos it seems some form of KPI is inevitable. This could also explain why Deepseek doesn't mind open sourcing as they have not r
107.
▲
by
fspeech
2y ago
It appears that they have really good toolings so they can focus on hiring really smart new grads and giving them a lot of leeway on choosing what to work on without worrying about how productive they will be. One particular memorable respo
108.
▲
by
fspeech
2y ago
Efficiency could lead to cheaper hardware for everyone, themselves included.
109.
▲
by
fspeech
2y ago
You need that to optimize load balancing. Unfortunately that gain is not available to small or individual deployment.
110.
▲
by
fspeech
2y ago
A better approach is to split the model with MOEs running on CPUs and MLAs running on GPU. See the ktransformers project: https://github.com/kvcache-ai/ktransformers/blob/main/doc/en... This takes a
111.
▲
by
fspeech
2y ago
China proves that you don't need 100 times return on capital to get potential innovators to take risk. 100% may be enough. People say China couldn't innovate due to its weak IPR protection, but it is turning out that you don'
112.
▲
by
fspeech
2y ago
You need a Chinese number to sign up and facial verification to make payments, as far as I can tell.
113.
▲
by
fspeech
2y ago
The smaller cluster is not necessarily a good thing. It may not be able to take advantage of unbalanced load on experts.
114.
▲
by
fspeech
2y ago
Deepseek doesn't have the infrastructure to support Apple. I suspect that they are also not interested in tailoring to Apple's needs given their mission and small size.
115.
▲
by
fspeech
2y ago
In a polytheistic context the invocation is much less severe, roughly equivalent to Apollo's Eyes or Zeus's Eyes.
116.
▲
by
fspeech
2y ago
That's due to the translation. The original term 天神 comes from the polytheistic Chinese folk religion so it doesn't have the same connotation.
117.
▲
by
fspeech
2y ago
Deepseek's open source inference code, while correct, may not be fully efficient. For example the MLA needs the right associative matrix multiplication order to be efficient.
118.
▲
by
fspeech
2y ago
Do you have any benchmark run yet? I am interested in knowing how many tokens/sec you can get to. Though in the end it should be more efficient to run the model on distributed server clusters.
119.
▲
by
fspeech
2y ago
Thanks for the thought provoking take.
120.
▲
by
fspeech
2y ago
Right. And the number is based on rental rates of GPUs so how many GPUs they own is irrelevant to the claim.
More ›