Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
rfoo
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
18 ms
·
211.
▲
by
rfoo
2y ago
For example, you have if check_something(): wall_of_text = textwrap.dedent(""" this is a multiline string """.strip("\n")) I hope you agree it's ugly. And
212.
▲
by
rfoo
2y ago
> though not sure how this works across the wire You send all header files to 1,000 boxes. Same as nocc. Oh, and if you really want, you can make Bazel ship your /usr/bin/gcc and co, too. It is just slow so nobody would li
213.
▲
by
rfoo
2y ago
> DeepSeek's next-gen models are already outperforming us, especially in reasoning and long-context capabilities This is interesting - this is the first time I heard someone claim that DeepSeek may excel at long-context capabilities
214.
▲
by
rfoo
2y ago
I know that, I'm in this game. I was comparing API throughput/ttft/ttbt of DeekSeek's own R1 API before it went viral in the West, and o3-mini. I remain unconvinced that DeepSeek themselves didn't optimize their o
215.
▲
by
rfoo
2y ago
But... they are? o3-mini is faster than DeepSeek-R1 and has comparable capability. And while I hate "AGI achieved internally" meme, o3 is significantly better than o1. Though I doubt how long until DeepSeek-R3 happens. They could
216.
▲
by
rfoo
2y ago
Definitely. My question was, porn (or other censored thing) could be in text, too. I don't understand why those who want censored content are primarily interested in graphics.
217.
▲
by
rfoo
2y ago
It's GPT-3.5 GPT-4 GPT-4-turbo GPT-4o-2024-08-06 GPT-4o GPT-4o-mini o1-preview o1 (low) o1 (medium) o1 (high) o1-mini o3-mini (low) o3-mini (medium) o3-mini (high) At this pace we are going to get to USB 3.2 Gen 1 real fast.
218.
▲
by
rfoo
2y ago
> if it allows for something they can’t do (often in gen graphics) What makes gen graphics stand out?
219.
▲
by
rfoo
2y ago
At this point, adopting Western labor laws actually helps China. China is facing increasingly severe challenges caused by the distribution of wealth. Changing labor laws does not fix them all, but it likely helps.
220.
▲
by
rfoo
2y ago
Correction: I meant 5-10 billion USD, too drunk yesterday
221.
▲
by
rfoo
2y ago
No. If you can't run it and most people can never run the model on their laptop, it's fine, let people know the fact, instead of giving them illusion.
222.
▲
by
rfoo
2y ago
Says someone who personally has 50-100 billion USD. And no, it's not net worth through corp shares. The guy is essentially his own LP.
223.
▲
by
rfoo
2y ago
Well, DeepSeek engineers are (desperately) fire-fighting as they don't have nearly as much capacity as needed. Competitors either already rushed release or decided to do an hush release of whatever they had in the pipeline. Sounds like
224.
▲
by
rfoo
2y ago
> none of the people listed on the DeepSeek papers got educated at US universities "You have been educated at foreign universities / worked at foreign companies" is indeed an excuse they have used at least once to refuse a
225.
▲
by
rfoo
2y ago
Well, I know. I still have connections back there. But yeah, I'm just a random guy on the Internet so what I said could be just myth too.
226.
▲
by
rfoo
2y ago
This is completely fake though. It was more like their founder decided to start a branch to do AI research. It was well planned, they bought significantly more GPUs than they can use for quant research even before they start to do anything
227.
▲
by
rfoo
2y ago
Total training FLOPs can be deduced from model architecture (which they can't hide since they released weights) and how many tokens they trained on. With total training FLOPs and GPU hours you can calculate MFU. And the MFU of their de
228.
▲
by
rfoo
2y ago
> Then a US compute provider should be able to launch a similarly-priced competitor Right, you just need a few months to implement efficient inference for MLA + their strangely looking MoE scheme + ..., easy! Oh wait, the inference schem
229.
▲
by
rfoo
2y ago
DeepSeek-V2/V3/R1's model architecture is very different from what Fireworks/Together/... were used to. That's their "business" model (okay, they don't care about business that much for now, but
230.
▲
by
rfoo
2y ago
> Because racks of H100s are not sustainable. Huh? Racks of H100s are the most sustainable thing we can have for LLMs for now.
231.
▲
by
rfoo
2y ago
We are going to see it happen without something like next generation Groq chips. IIUC Groq can't run actually large LMs, the largest they offer is 70B LLaMA. DeepSeek-R1 is 671B.
232.
▲
by
rfoo
2y ago
It's not their model being bad, it's claude.ai having pretty low quota for even paid users. It looks like Anthropic doesn't have enough GPUs. It's not only claude.ai, they recently pushed back increasing API demand from
233.
▲
by
rfoo
2y ago
> which ironically is called "OpenPlatform" for some reason This is pretty weird, the original text is 开放平台, but it basically is another name for "API" in China. Not sure who started this, but it's really popular
234.
▲
by
rfoo
2y ago
I have laptop equivalents in the same memory range and is at least $2,500 cheaper. Unfortunately, it does not have "unified memory", a somewhat "powerful GPU", and of course no local LLM hype behind it. Instead, I'v
235.
▲
by
rfoo
2y ago
> Second, I think the statistic is that 81% of businesses have had an outage due to certificate expiry. So you need to understand that making certs expire more is inherently damaging. Uh, no. Most of the outage due to certificate expiry
236.
▲
by
rfoo
2y ago
I believe the law mentioned here isn't focused on which organization it is. The law itself basically said you can't export recommendation algorithm. Yes, in the very similar wording as in "you can't export certain GPU ch
237.
▲
by
rfoo
2y ago
> But these buyers don't want the actual value of TikTok to drop to zero Quick reminder: TikTok is available for most of the planet (except China), so a US ban does not make the actual value of TikTok to drop to zero. It makes a sel
238.
▲
by
rfoo
2y ago
Maybe. It's pretty weird and I'm still thinking about it. You can't throw junior engineers working on an issue under the bus when they clearly can't do that. Or at least it takes some effort. In return you may coach them
239.
▲
by
rfoo
2y ago
Devin does ask for help when it can't do something. I think I have it asked me how to use a testing suite it had trouble running. The problem is it really really hate asking for help if it had a skill issue, it would prefer running in
240.
▲
by
rfoo
2y ago
... which means automation was not setup correctly and 90 days is still too long that you just tolerated it. If it was 6 days after a few turns you would have decided "fuck it I'm going to spend time fixing it once and for all&quo
More ›