Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
calebkaiser
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
calebkaiser
2mo ago
I'm typically very skeptical of most content marketing/corporate PR. In the case of this incident, I struggle to see the clear upshot for OpenAI. It seems pretty unlikely they'd ever okay this intentionally as some sort of ma
32.
▲
by
calebkaiser
2mo ago
Implementing models directly from papers is typically pretty doable (and is of course more straightforward when the full implementation is open sourced). Often there is some amount of specific knowledge, like particular hyperparameters, tha
33.
▲
by
calebkaiser
2mo ago
OpenAI's head of strategic futures publicly stated that you can't explain the quality of the newest Kimi via distillation. Further, you can just read the papers released alongside most open models. Plenty of hugely influential res
34.
▲
by
calebkaiser
2mo ago
Here's his original post https://x.com/i/status/2078133895766114412
35.
▲
by
calebkaiser
2mo ago
The guy who said those things just joined OpenAI in the last month. He was previously a senior policy advisor, specifically on AI, to the current administration in the White House. I don't know that the US government has a unified posi
36.
▲
by
calebkaiser
2mo ago
The conversation going on in the industry is a bit broader than that. OpenAI's head of strategic futures just said last week that open models are inherently decelerationist, ungovernable, and will slow development on the frontier. You&
37.
▲
by
calebkaiser
2mo ago
Money is a big part of it. The notable thing about Chinese open models is their size. Most open source models out of US labs have not tried to compete with the enormous flagship models you see from frontier labs now. Cohere, Arcee, Poolside
38.
▲
by
calebkaiser
3mo ago
I'm sympathetic to this POV. I think specifically in the case of Slack, it's not their responsibility in a cosmic sense, but it is in line with the direction and general promise they've been making to customers since forever.
39.
▲
by
calebkaiser
3mo ago
I am strongly in favor of open models, open source ML more broadly, and am pretty critical of the cynical positions adopted by major US labs vis a vis open models. But this is an insane characterization. Literally every single researcher an
40.
▲
by
calebkaiser
3mo ago
OpenAI's Head of Strategic Futures just this week posted this about the latest Kimi release: "It's a very good model! I don't think its performance can be explained away by distillation or anything like that." It wa
41.
▲
by
calebkaiser
3mo ago
This was my initial reaction to reading this post as well. Additionally, as I get older, I find the sentiment of "we're not trying to build a perfect system here" is less about "let's just go fast vroooom" and
42.
▲
by
calebkaiser
3mo ago
This is not my experience working in the field the last 8 years. There is not a dearth of talented researchers, engineers, etc. who are willing to contribute to open models. Just look at the ecosystem generally, from academia to industry. S
43.
▲
by
calebkaiser
3mo ago
Gwern's absurdly catalogued personal site is one of those online artifacts that I hope never changes.
44.
▲
by
calebkaiser
3mo ago
Based on a cursory read of the situation, it seems similar (at least on its face) to the Waymo vs Uber situation. In that case, Uber payed a Waymo an equity stake and signed an agreement about which technology they would/wouldn't
45.
▲
by
calebkaiser
3mo ago
I mean, OpenAI delayed the public release of GPT-2 back in 2019 because it seemed capable of authoring interesting blog posts (that also happened to be untrue). It was a pretty big deal the first time Transformer models were capable of gene
46.
▲
by
calebkaiser
3mo ago
I think the author largely agrees with you re: type systems and LLMs. He's pretty explicit that Haskell should be very well positioned to be a power language for LLM-assisted programming, but that the Haskell ecosystem presents the bot
47.
▲
by
calebkaiser
3mo ago
I've been a power user of LLMs for software development for a while now, and I've found two things to be true: - The benefits of more "extreme" type systems are more accessible and valuable than ever. I have a fairly inv
48.
▲
by
calebkaiser
3mo ago
Lots of researchers have done just this! There's a really rich history of research + lots of contemporary work on different encoding/representation strategies. This might be interesting to you: https://sbert.net/
49.
▲
by
calebkaiser
3mo ago
Nah, optical compression is a thing. You see it in a lot of different areas in ML. In this case, the "trick" has been known for a while, and belongs to a whole world of compression research. But I think where you're maybe get
50.
▲
by
calebkaiser
3mo ago
This has been a (noble) goal of lots of different projects in the community for a long time. Federated learning projects like Flower have been chipping away at it for a long time. There are many many hurdles to be cleared before anything in
51.
▲
by
calebkaiser
4mo ago
This is a good starting point: https://huggingface.co/docs/peft/developer_guides/model_merg... But yes, in general, merging refers to techniques that directly blend the weights of different models mathematica
52.
▲
by
calebkaiser
4mo ago
I don't understand this line of criticism exactly. By putting new information in the context window, you are materially changing the activations at your point of sampling, which is literally "customizing with mere markdown files.&
53.
▲
by
calebkaiser
4mo ago
Eh, Watson was a classic open domain QA system originally, no deep learning or much of what we think of in an "AI platform" today. It was one of a bunch of such systems that were built in that early 2000s period. They all failed b
54.
▲
by
calebkaiser
5mo ago
My experience has been that this is not unique to tech, and is common in all large enough industries. I think it's just the natural emergence of reward hacking i.e. if you're an executive at Pepsi and your job is largely to increa
55.
▲
by
calebkaiser
5mo ago
If I'm remembering right, it was weirder than that, as Llama's originally release strategy was sort of bizarre. You did have to apply for access, but if you met their criteria (basically if you were the right profile of researcher
56.
▲
by
calebkaiser
5mo ago
2 years? 2 years ago, gpt-4o was OpenAI's flagship model. The gap is real, but much smaller than 2 years.
57.
▲
by
calebkaiser
6mo ago
I've worked on two open source infrastructure projects that raised money now, and am friends with people involved in many more. I'd put a couple of asterisks next to the claims in this article: - VCs definitely cared about our Sta
58.
▲
Opik – The missing observability layer for OpenClaw
(github.com)
1 points
by
calebkaiser
7mo ago
|
0 comments
59.
▲
by
calebkaiser
7mo ago
It's funny how "the real split" is always between the intellectually and morally superior (me) and the inferiors (them).
60.
▲
Opik – An Observability Layer for OpenClaw
(github.com)
2 points
by
calebkaiser
7mo ago
|
0 comments
More ›