Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
hodgehog11
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
hodgehog11
2mo ago
Agreed, AI is not capable at the moment of coming up with radical ideas to solve the tough problems. Sadly, I would argue many problems in math are likely to be found to be not actually tough in this sense, and those working in "comfor
32.
▲
by
hodgehog11
2mo ago
Neither Claude nor GPT are acceptable for writing English text. Personally I have found Gemini to be far better, and that is really all I use it for.
33.
▲
by
hodgehog11
2mo ago
That has always been the major strength of GPT, that's the model you use for checking. It often nearly isn't as good for creation though.
34.
▲
by
hodgehog11
2mo ago
Agreed. The benchmark closest to my experience is FrontierMath Tier 4. Fable and Sol (90%) are very far ahead of Kimi K3 (not even 40%). Kimi is trained heavily to basic agentic tasks, like all the other open models right now.
35.
▲
by
hodgehog11
2mo ago
They are likely assessing based on "raw intelligence" benchmarks, rather than agentic ones. Fable crushes in those, but that doesn't necessarily translate to microscopic rigor, which is what most people use these models for.
36.
▲
by
hodgehog11
2mo ago
It really does depend on your application. In my domain (math research), it is substantially better. Fable can solve really hard tasks with surprising consistency. It makes mistakes, and occasionally refuses, but honestly, at the top level,
37.
▲
by
hodgehog11
2mo ago
Tell that to my colleagues. Despite Sol getting the attention, Fable is really starting to have an impact on mathematicians right now. It has unbelievable insights in a lot of cases that can rapidly speed up progress.
38.
▲
by
hodgehog11
2mo ago
I think it is pretty safe to say at this point that having large open LLM models available is better for humanity than them remaining proprietary. Echoing Linus Torvalds' recent comments, AI is genuinely useful right now, and is here t
39.
▲
by
hodgehog11
2mo ago
The "second only to Fable 5" comment is pretty telling here. I remember early on when a lot of naysayers were saying that Fable was barely an improvement on Opus. Like it or not, Anthropic have a genuine moat right now with that m
40.
▲
by
hodgehog11
2mo ago
That would go against everything that Dario believes in (note that I refer to the CEO and not the company; the staff at Anthropic are not so ridiculous). He believes in Anthropic being the sole arbiter of the forefront of this technology, b
41.
▲
by
hodgehog11
2mo ago
Value models are always going to be there; you can always distill from a larger model. Having a really intelligent model, regardless of the size, is much better for building confidence in your brand. That is a big reason why the US companie
42.
▲
by
hodgehog11
2mo ago
I love convex optimization and there are a few SciML projects I am on where I really need results from there. But in AI research with deep neural networks, it's become a liability, because people will just not let go. I'm getting
43.
▲
by
hodgehog11
2mo ago
It's not a matter of whether the theory "works"; it's a matter of whether one is asking the right questions. Convex optimization studies how quickly an optimizer can reach the optimum. In the non-convex case, there are m
44.
▲
by
hodgehog11
2mo ago
No, I have to push back as well, sorry. It takes a very long time to get to the "near-minimizer" stage when training a neural network, and in practice, you never get there (see neural scaling law regimes). What you are saying is t
45.
▲
by
hodgehog11
2mo ago
Very confused by this comment. The older (poorer) parts of the ML literature focus on models with convex and (gradient-)Lipschitz objectives, but that's not representative of reality, not even close. Modern objectives for AI models are
46.
▲
by
hodgehog11
3mo ago
I really envy you. There is a clear divide amongst my colleagues now in terms of who is using Fable and who isn't (this is math work, so it is well and truly better than all the other models). Everyone is quickly becoming reliant on th
47.
▲
by
hodgehog11
3mo ago
For the particular task I'm working on (a mathematical task in validated numerics), even Sol has generally just repeatedly given up. I asked it for the main problems it could not solve, gave them to Fable, and Fable solves them, every
48.
▲
by
hodgehog11
3mo ago
Overengineering is the name of the game with Fable. Sometimes you don't want that, sometimes you really do, especially as a researcher. It's a very nice tool to have around for those special tasks.
49.
▲
by
hodgehog11
3mo ago
For some tasks, there is no amount of "steering" that will produce sensible code. The model needs to be sufficiently capable as a baseline; this is the "intent" that people are referring to with Fable.
50.
▲
by
hodgehog11
3mo ago
I'm sorry to hear you are unable to use Fable; my partner is in the same boat and it frustrates her immensely to see what I've been able to do with it. As someone who is working with developing new linear algebra routines, Fable i
51.
▲
by
hodgehog11
3mo ago
I would be absolutely stunned if this were really the case in general given how irresponsibly large Fable is, and 5.6 Sol most definitely is not. It depends on what your problems are though, I suppose, since there are those that swear Fable
52.
▲
by
hodgehog11
3mo ago
I don't agree that my argument was a "No true Scotsman", since the argument made above (as far as I read it) was that "all PhDs are a waste of time". My counterargument was that the PhD, if done right according to i
53.
▲
by
hodgehog11
3mo ago
Yes, I did get it within the last ten years (with an excellent supervisor fortunately), and I take part in student supervision. I get to talk with the students and determine what is most valuable for them in the long-term. Considering where
54.
▲
by
hodgehog11
3mo ago
I won't try and argue the merits of Bachelor's and Masters. But if you honestly believe that you can pick up the same experience from PhD on the job, then is seems like you learned little from that PhD that you were supposed to an
55.
▲
by
hodgehog11
3mo ago
> Basically we are saying that if a hobbyist really wants to operate and maintain something, they should be allowed to after some amount of time if the studio stops. I don't believe that is the problem as defined by SKG. I think the
56.
▲
by
hodgehog11
3mo ago
Thanks for continuing the discussion! It's not often an outsider would get to speak with someone with your experience. I definitely agree that it is not tenable in any studio to have a full time staff member dedicated to packaging soft
57.
▲
by
hodgehog11
3mo ago
Absolutely! Happy to learn from an industry veteran; hard to argue with that pedigree (part of why I like HN). I figured you might have been, but wanted to push back a little because I think it is important. Here is my impression. 2001 feel
58.
▲
by
hodgehog11
3mo ago
Small like 30 people? I don't think so. Genuinely small studios are able to preserve their games for the long-term, so there is no excuse. AAA gaming isn't something that needs to be protected. Anything of sufficient magnitude and
59.
▲
by
hodgehog11
3mo ago
Judging from the decisions and outputs of the last decade or so, the leadership at Meta, including Mark Zuckerberg, have got to be among the most incompetent I have ever seen. They go all in on the worst decisions; not just the worst in hin
60.
▲
by
hodgehog11
3mo ago
It just requires game engines to develop a tool to make this much easier to do. The rule won't apply retroactively, so it just means different design choices from the start. > For a small game studio it may be incredibly prohibitive
More ›