Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
oofbaroomf
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
oofbaroomf
1y ago
How big do you think your index is compared to Google?
32.
▲
by
oofbaroomf
1y ago
Things like Hunyuan 3D are nice for game assets and the like, but they aren't able to really do CAD well. That would be like using Stable Diffusion to code.
33.
▲
by
oofbaroomf
1y ago
Yes, there has been. Unfortunately, there are a few core issues blocking this from becoming a big thing: 1. The majority of 3D modeling is not done parametrically, meaning there is not a lot of data. The little data there is is generally in
34.
▲
by
oofbaroomf
1y ago
Many models that use test time compute are MoEs, but test-time compute is generally meant to refer to reasoning about the prompt/problem the model is given, not about reasoning about which model to pick, and I don't think anyone h
35.
▲
by
oofbaroomf
1y ago
First off, they are basically completely different technologies, so it would be disingenuous to act like it's an apples-to-apples comparison. But a simple way to see it is that when you pick between multiple large models that have diff
36.
▲
by
oofbaroomf
1y ago
No. Each expert is not separately trained, and while they may store different concepts, they are not meant to be different experts in specific domains. However, there are certain technologies to route requests to different domain expert LLM
37.
▲
by
oofbaroomf
1y ago
Yeah, it was for the Llama team because they love playing in ball pits instead of releasing good models.
38.
▲
by
oofbaroomf
1y ago
Probably one of the best parts of this is MCP support baked in. Open source models have generally struggled with being agentic, and it looks like Qwen might break this pattern. The Aider bench score is also pretty good, although not nearly
39.
▲
by
oofbaroomf
1y ago
If you would use LLM-assisted CAD for real industrial design, you would have to end up by specifying exactly where everything has to go and what size it has to be. But if you are doing that then you may as well make an automated program to
40.
▲
by
oofbaroomf
1y ago
This is even more incomprehensible to users who don't understand what this naming scheme is supposed to mean. Right now, most power users are keeping track of all the models and know what they are like, so this naming wouldn't hel
41.
▲
by
oofbaroomf
1y ago
Still a knowledge cutoff of August 2023. That is a significant bottleneck to devs using it for AI stuff.
42.
▲
by
oofbaroomf
1y ago
When are they going to release o3-high? I don't think it's in the API, and I certainly don't see it in the web app (Pro).
43.
▲
by
oofbaroomf
1y ago
It's there now in the web app for me.
44.
▲
by
oofbaroomf
1y ago
Same...
45.
▲
by
oofbaroomf
1y ago
Claude got 63.2% according to the swebench.com leaderboard (listed as "Tools + Claude 3.7 Sonnet (2025-02-24)).[0] OpenAI said they got 69.1% in their blog post. [0] swebench.com/#verified
46.
▲
by
oofbaroomf
1y ago
They didn't provide a comparison either in the GPT-4.1 release and quite a few past releases, which is telling of their attitude as an org.
47.
▲
by
oofbaroomf
1y ago
Finally, a new SOTA model on SWE-bench. Love to see this progress, and nice to see OpenAI finally catching up in the coding domain.
48.
▲
by
oofbaroomf
1y ago
Leader is debatable, especially given the actual comparisons...
49.
▲
by
oofbaroomf
1y ago
Oh, ok. But it's still quite telling of their attitude as an organization.
50.
▲
by
oofbaroomf
1y ago
Reasoning models have the o first, non-reasoners have the digit first.
51.
▲
by
oofbaroomf
1y ago
I'm not really bullish on OpenAI. Why would they only compare with their own models? The only explanation could be that they aren't as competitive with other labs as they were before.
52.
▲
by
oofbaroomf
2y ago
I (not GP) would like to be able to choose between the options. Inference provider isn't super necessary though (can do that through huggingface).
53.
▲
by
oofbaroomf
2y ago
Yeah, but Openrouter has a 5% surcharge anyway.
54.
▲
by
oofbaroomf
2y ago
The Seattle/Bellevue area.
55.
▲
by
oofbaroomf
2y ago
Questions.
56.
▲
by
oofbaroomf
2y ago
Do you think Claude Code is "better", in terms of capabilities and token efficiency, than other tools such as Cline, Cursor, or Aider?
57.
▲
by
oofbaroomf
2y ago
Unified memory is great because it's fast, but you can also get a lot of system memory on a "conventional" machine like OP's, and offload MOE layers like what Ktransformers did, so you can run huge models with acceptable
58.
▲
by
oofbaroomf
2y ago
The M4 Max has up to 546 GB/s. The M4 Pro, what GP was talking about, has only 273 GB/s. An M4 Max with that much RAM would most likely exceed OP's budget.
59.
▲
by
oofbaroomf
2y ago
The bottleneck for single batch inference is memory bandwidth. The M4 Pro has less memory bandwidth than the P40, so it would be slower. Also, the setup presented in the OP has system RAM, allowing you to run models than what fits in 48GB o
60.
▲
by
oofbaroomf
2y ago
That was the R&D cost. According to OP, the cost of building one of these is around $1500.
More ›