4 ms·
I actually still don't see the source for them trying several times, but we can take that for granted. Regardless, as I said: 1. It's labeled as "moderately in
by staticassertion 6mo ago
I actually still don't see the source for them trying several times, but we can take that for granted. Regardless, as I said:
1. It's labeled as "moderately interesting"
2. They said that they expect an expert could solve it in 1-3 months
3. They had already come up with the solution that the AI had but weren't convinced it would have worked
So how big was the gap here, do you think?
- famouswaffles 6mo agoYes, a "moderately interesting" Open problem. I can't think of any chores that would take an expert months to complete. I can't think of any chores that I've completed but was then 'unconvinced could work'. Please sit down and think about what you are saying here. Are we still talking about chores ? One of the more strange phenomena with machines getting better and the incessant need (seemingly driven by human exceptionalism) to downplay each result, is that you just end up belittling humans in the process. This is significant. Your analogy is wrong. It's fine to admit it.
- staticassertion 6mo agoWriting a complex parser or certainly a compiler is a 1 - 3 month project, for example. Again, I'm not trying to downplay this, but to frame this accurately. I think an AI being able to build a parser/ compiler is cool too. > One of the more strange phenomena with machines getting better and the incessant need (seemingly driven by human exceptionalism) to downplay each result, is that you just end up belittling humans in the process. I don't believe in human exceptionalism at all, don't attribute positions to me.
- famouswaffles 6mo ago>Writing a complex parser or certainly a compiler is a 1 - 3 month project, for example. 1. Estimating time completion of something that has been done multiple times before and an open problem that has not yet been solved is a different matter entirely. 1 to 3 months is an educated guess and more likely than not, an underestimate. 2. I do not think months long complex compilers and parsers are being routinely completed by LLMs as your original comment implied. Regardless, they are different classes of problems.
- staticassertion 6mo agoI don't get what either of your points is intended to demonstrate. Let's revisit the first post I replied to: > It's deeply surprising to me that LLMs have had more success proving higher math theorems than making successful consumer software As far as I can tell, they absolutely have not had more success in this area relative to making successful consumer software.
- famouswaffles 6mo agoWell we are kind of arguing past each other aren't we ? "More success" is a bit vague in this instance but building a compiler that would take a programmer 1 to 3 months is not comparable to this result regardless of whatever similarity exists in time completion estimates. That's the point. You can publish a paper (and in fact the researchers plan to) off this result. A basic compiler is cool but otherwise unremarkable. It's been done many times before. You are leaning too hard on how long the researchers (who again did not manage to solve the problem in their attempts) estimated this would take and the "moderately interesting" tag of again, what was still an open research problem. This, alongside a few math and physics results that have cropped up in the last few months is easily more impressive than the vast majority of work being done with LLMs for software.
- staticassertion 6mo ago> "More success" is a bit vague in this instance but building a compiler that would take a single programmer 1 to 3 months is not comparable to this result regardless of whatever similarity exists in time completion estimates. That's the point. I guess we just disagree on this. It's not clear to me that these are totally different in terms of what they represent. > You can publish a paper (and in fact the researchers plan to) off this result. A basic compiler is cool but otherwise unremarkable. Publishing papers means very, very little to me. I can publish a paper on a programming language, you know that, right? > You are leaning too hard on how long the researchers (who again did not manage to solve the problem in their attempts) estimated this would take and the "moderately interesting" tag of again, what was an open research problem. I obviously estimate my "leanings" as being appropriate. I'm just using the researchers direct quotes. Factually, they had already come up with the approach that ultimately panned out. Factually, they estimated that a human could do this in some timeframe. What am I overly leaning on here? > This, alongside a few math results that have cropped up in the last few months is easily more impressive than the vast majority of work being done with LLMs for software. I think both are impressive, I don't know that I would draw some sort of big conclusions about it at this point. I definitely wouldn't draw the conclusion that AI is better at formal mathematics than producing software.
- famouswaffles 6mo agoAlso, the full write up does not say the researchers solved it.
- staticassertion 6mo ago> I had previously wondered if the AI’s approach might be possible, but it seemed hard to work out. They didn't solve it, that's fair. They did consider the approach already.