Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
bugglebeetle
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
91.
▲
by
bugglebeetle
2y ago
I would say “for now” and that Anthropic’s inability to execute (3.7 was only released in response to competition from DeepSeek) raises a lot of questions about their future. While I agree that DeepSeek is still behind, Anthropic also does
92.
▲
by
bugglebeetle
2y ago
Where do you see open source benchmark results that confirm Mistral’s performance?
93.
▲
by
bugglebeetle
2y ago
Congrats to Mistral for yet again releasing another closed source thing that costs more than running an open source equivalent: https://github.com/DS4SD/docling
94.
▲
by
bugglebeetle
2y ago
Unitree just open-sourced their robot designs: https://sc.mp/sr30f China’s strategy is to prevent any one bloc from achieving dominance and cutting off the others, while being the sole locus for the killer combination of in
95.
▲
by
bugglebeetle
2y ago
Anyone doing this work should be building on top of the OpenAlex API and its Unpaywall integration. This allows you to traverse the world’s largest research graph and access full text, where it’s available in any publicly available form.
96.
▲
by
bugglebeetle
2y ago
I see you’ve never heard of the market for fine art.
97.
▲
by
bugglebeetle
2y ago
> LLMs. I really don't get how so many people can get LLMs to code most of their tasks without those issues permanently popping up. You write tests in the same way as you would when checking your own work or delegating to anyone els
98.
▲
by
bugglebeetle
2y ago
The problem isn’t the prestige it’s that prestigious institutions in America don’t produce high-quality talent. They’re instead mostly corrupt credentialing mills for the rich and well-connected. From what I understand, DeepSeek also only h
99.
▲
by
bugglebeetle
2y ago
It’s funny because I also have found myself doing this exact with R1+Sonnet 3.5 recently. Windsurf allows you to do a chat mode exchange with one model and then switch to another to implement. The reasoning models all seem pretty poorly imp
100.
▲
by
bugglebeetle
2y ago
If Buffet is fearful, the people being greedy are so beyond your means and chasing an irrational self-destructiveness that you have zero chance of seeing an upside by being aligned with them.
101.
▲
by
bugglebeetle
2y ago
My favorite thing about LibreOffice is that it can open CSVs without breaking them. Pretty much everything else about it is a slow, unintuitive mess, but this one killer feature makes me always install it and something that somehow Excel d
102.
▲
by
bugglebeetle
2y ago
o3 mini’s date cut-off is 2023, so it’s unfortunately not gonna be useful for anything that requires knowledge of recent framework updates, which includes probably all big frontend stuff.
103.
▲
by
bugglebeetle
2y ago
This is not creating art, but illustration. Art is where the meaning is inherent to the work, whereas illustration is when it is derived from something else that it accompanies (e.g. a text). There’s obviously a spectrum here, but GenAI blo
104.
▲
by
bugglebeetle
2y ago
Interested to see what folks do with putting DeepSeek-style RL methods on top of this. The smaller Mistral models have always punched above their weight and been the best for fine-tuning.
105.
▲
by
bugglebeetle
2y ago
Like all those papers with their long lists of citations OpenAI has been releasing?
106.
▲
by
bugglebeetle
2y ago
Yeah, #1 is way worse and #2 falls under “turnabout is fair play.”
107.
▲
by
bugglebeetle
2y ago
> “people should not be surprised when they are occasionally surprised, and they should anticipate being surprised some more in the future.” It’s mildly hilarious that you’re deriding Taleb for thinking (and investing) at a level that al
108.
▲
by
bugglebeetle
2y ago
So context size actually helps with this, relative to how LLMs are actually deployed as applications. For example, if you look at how the “continue” option in the DeepSeek web app works for code gen, what they’re likely doing is reinserting
109.
▲
by
bugglebeetle
2y ago
It doesn’t call it into question- they’re not. OpenAI has been bleeding researchers since the Anthropic split (and arguably their best ones, given Claude vs GPT-4o). While Google should have all the data in the world to build the best mode
110.
▲
by
bugglebeetle
2y ago
I mean what’s also incredible about all this cope is that it’s exactly the same David-v-Goliath story that’s been lionized in the tech scene for decades now about how the truly hungry and brilliant can form startups to take out incumbents a
111.
▲
by
bugglebeetle
2y ago
Uh, they invented multilatent attention and since the method for creating o1 was never published, they’re the only documented example of producing a model of comparable quality. They also demonstrated massive gains to the performance of sma
112.
▲
by
bugglebeetle
2y ago
> AI is DOA. LLMs have no successor, and the transformer architecture hit it's bathtub curve years ago Tell me you didn’t read the DeepSeek R1 paper without telling me you also don’t know about reinforcement learning.
113.
▲
by
bugglebeetle
2y ago
Termjuicer
114.
▲
by
bugglebeetle
2y ago
4o is more expensive than DeepSeek-R1, so…? Even if we took your premise as true and we say they are as good as DeepSeek, this would just mean that OpenAI is wildly overcharging its users.
115.
▲
by
bugglebeetle
2y ago
> It can solve that Arc-AGI benchmark thing at huge compute cost Considering DeepSeek v3 trained for $5-6M and their R1 API pricing is 30x less than o1, I wouldn’t expect this to hold true for long. Also seems like OpenAI isn’t great at
116.
▲
by
bugglebeetle
2y ago
YC’s own incredible Unsloth team already has you covered: https://huggingface.co/unsloth/DeepSeek-R1-Distill-Llama-8B
117.
▲
by
bugglebeetle
2y ago
I used to think this, but using o1 quite a bit lately has convinced me otherwise. It’s been 1-shotting the fairly non-trivial coding problems I throw at it and is good about outputting large, complete code blocks. By contrast, Claude immedi
118.
▲
by
bugglebeetle
2y ago
The same way you check performance for any problem like this: by creating one or more manually-labeled test datasets, randomly sampled from the target data and looking at the resulting precision, recall, f-scores etc. LLMs change pretty muc
119.
▲
by
bugglebeetle
2y ago
> Indians and Chinese are no longer considered minorities in the US, so it makes sense to curb the influx. Why would this ever be rational a basis for immigration?
120.
▲
by
bugglebeetle
2y ago
>the logical next step is for payers to vertically integrate and bring more care in house where they can better control cost and quality. …except you skipped over the part where UHC billed at higher rates for the clinics they own so they
More ›