10 ms·
I was recently in Palo Alto, and bumped into a newly founded startup (I don't remember the name unfortunately) who set themselves the grand the vision of exactl
by gregdoesit 3y ago
I was recently in Palo Alto, and bumped into a newly founded startup (I don't remember the name unfortunately) who set themselves the grand the vision of exactly this: winning a gold medal on the international Olympiad using AI. Their plan was to build mostly on LLMs as a start, and iterate as they go. In their barebones office space, they had a poster with a countdown of the number of weeks till the event: it was 36 at the time.
It sounded interesting to wonder how far they could go with this kind of approach. I thought they were aiming for the moon: but also respected the boldness and determination. They had the funding to operate for at least a year, and were very focused to get there.
Seems like this prize will supr hundreds (or thousands) of teams competing in exactly this space. Perhaps it will have a similar effect like the $1M Netflix Prize in 2009 for recommendations algorithms!
- mlengineerio 3y agoCan you share the name of the startup? I'm also in Palo Alto and tinkering LLM for Match competition too (math.llmlab.io).
- anonylizard 3y agoWell math solving is exactly what the rumored Q* is aiming towards too. I don't think it'll take more than 2 years before some LLM + RL system can take the gold medal. I think companies like OpenAI are aiming for something far more ambitious, like solving a millennium prize problem (even with human assistance). That's the kind of news release that'll add another $100 billion to your market cap.
- dimask 3y agoTheir current ambition is to be able to solve school math, which is quite far away from solving unsolved conjectures or math olympiads. I really doubt that any of this is within LLM/transformer scope, except maybe in some auxiliary sense to other, much different architectures.
- anonylizard 3y agoArt isn't an easier problem than math. An artbot would have sounded more sci-fi than a mathbot only 2 years ago. Yet it only took the AI world 1.5 years to go from drawing child scribbles to replicating top artists with like 90% similarity (I can barely tell the difference between AI and human drawn art anymore with the new NovelAI model). It won't be long before AI starts to go superhuman in art skills. It won't take long from a school-math model to math olympiad model (I'd say 1 year is enough), and going to unsolved conjectures won't be that long either (2-3 years?). We know from AlphaGo that its possible to make AI systems far superhuman at solving some abstract math problem.
- RandomLensman 3y agoAlphaGo solved Go?
- anonylizard 3y ago"Superhuman at solving" != "solved". AlphaGo didn't solve go (Ie, can the first mover guarantee a win?). However, it understood go at a far, far superior level to any human. A mathbot doens't have to solve math in general. It merely has to be better at solving math than any human mathematician to be considered ASI. And it only has to be better than the 'average' human mathematician to be extremely useful in accelerating math research.
- RandomLensman 3y agoWhat mathematics was AlphaGo solving? What do you mean by solving there?
- dimask 3y agoProving that the first player has a winning strategy, or the optimal strategy for both players leads to draw.
- antoinexp 3y agoMight depend on the terms you have in mind, but current consensus seems to be more like 4-5 years as we speak on https://www.metaculus.com/questions/6728/ai-wins-imo-gold-medal/ https://www.metaculus.com/questions/6728/ai-wins-imo-gold-me...
- auntienomen 3y agoThat's 4-5 years for solving Olympiad problems. Those are just very tricky high school math problems. They have solutions and can generally be solved by applying some combination of standard tricks. It's very much the sort of thing an LLM should be good at. Solving Millennium problems is a whole different ballgame. It's not known if these problems are solvable within ZFC axioms. (In one case, the Yang-Mills prize, stating the problem mathematically is part of the challenge.) All of the obvious applications of known tricks have been tried and failed. To solve such problems, one probably has to invent new and surprising mathematical definitions, building a framework in which the problem becomes solvable. This is something that LLMs will be crap at; the process of invention is not represented in any training data we have access to.
- EVa5I7bHFq9mnYK 3y agoWe can look at it this way: there are ~1000 chess Grandmasters and one World Champion. It took very short time for AI to go from beating an average GM to beating World Champion. There are ~1000 MO winners and 1 (one) Millenial problem solver ...
- auntienomen 3y agoWe can. But doing so frames math research as the same sort of activity as math problem solving.. it's not. Many imo champions struggle to do any successful math research. And many successful math researchers (e.g., all of the most recent batch of Fields medallists) never did the Oympiad at all.
- EVa5I7bHFq9mnYK 3y ago
- eddtries 3y agoI have a couple friends who did the Math tripos at Cambridge (so a pretty high level!) who work in tech and have unanimously said they have 0% expectations of an LLM doing a millennium problem anytime soon
- Enginerrrd 3y agoYeah, millennium problems almost certainly require truly novel nontrivial ideas to solve. That's a tough thing for AI to do. On the other hand, Terrence Tao had an interesting article on his blog a while back where he was trying to solve a problem and asked chatGPT about it in a high-level strategy sense. ChatGPT suggested several reasonable approaches, one of which turned out to work. That's nowhere near solving a millennium problem, but it is very interesting and suggests fairly sophisticated conceptual understanding of mathematics nevertheless. Current architecture and training methods I don't think are enough to get there. However, with enough compute, I can plausibly envision some sort of meta training of LLMs using an analogy to GANs where one network tries to synthesize new correct ideas and the other shoots them down as not novel, not correct, or not sufficiently interesting. Such an approach I think could perhaps work, but the compute needed would probably be pretty high.
- jacobr1 3y agoAnd also if you could combine that with some kind of representation software like Lean that can validate proofs, perhaps you can generate some kind of targeted search of the problem space and with a combination of a general high level strategy and maybe brute force of some sub-problems, find novel solutions. My understanding is that is the common human approach: gain some kind of intuition of problem, try a few things and then iteratively refine. Sometimes that works, sometimes you need to find a new starting point. It seems plausible we could automate that workflow, with likely mixed but still useful results.
- eddtries 3y agoI’m not trying to understate the stuff you can do with LLMs or formal proof tooling that could be attached to randomly try things till it reaches a solution to a novel problem, but some solutions you see to lesser problems are sometimes so pie-in-the-sky I think stumbling on one is barely better than a random walk. And I think mathematicians like Tao are far better at narrowing that down. As an aide I can see it’s use, I just don’t believe we’re going to have LLMs & tooling solve these grand problems until compute power is orders of magnitudes better, and even then I’m still sceptical. BUT I’m not a mathematician and this is based on my intuition from pub talks :)
- auggierose 3y agoI don't think anything came out of the Netflix prize, did it?
- anonylizard 3y agoI don't think the actual winning algorithm itself was used, because real world systems have more constraints/requirements than what the recommender was trained on. But that was in 2009, pre deep-learning/AI summer, and $1 mil clearly helped stimulate interest in that area. Today we see multiple billion dollar recommender systems, like Tiktok. Netflix ironically benefits the least from recommenders due to the nature of its dataset (Very expensive, low sample size).
- jacobr1 3y agoMy understanding was that their research on what drove engagement shifted quite a bit. Things like social proof, and product patterns like auto-loading the next episode to binge drove engagement metrics. Recently there were some articles about their team custom-identifying which cuts of a video to show as a trailer maximized engagement on a personal level. In some sense that is a recommendation, but it is a broader problem space.
- rcpt 3y agoSocial proof on Netflix?
- jacobr1 3y agoYeah, the top10 in your region carousels get high engagement
- gorkish 3y agoI have a more cynical take; the recommendations declined when Netflix started producing their own content. Prior to this, what constituted a "good recommendation" was aligned between Netflix and the customer, but afterwards not so much. Today Netflix is in the "how do we get our customers to use our service as little as possible but still pay us every month" phase of their mediacom hypocracy. From a business standpoint, that is their best optimization. They are AOL/TW from 20 years ago.
- dist-epoch 3y agoWhat's his plan on being allowed to compete? Or will he be doing it after the fact, when the questions are published.
- lupire 3y agoAfter the fact.
- jasfi 3y agoI'm not sure if such a system would qualify, the prize specifically mentions a model that could solve specific types of problems. I'm building something myself that I hope will be able to work similarly: https://aiconstrux.com https://aiconstrux.com
- alecco 3y agoI wish there was a serious study of Dunning-Kruger effect in Silicon Valley.
- larodi 3y agoyour sense of humor is much appreciated. couldn't have said it better. hah :) basically i would suggest the study to expand to everyone doing anything titled AI at this moment.
- eddtries 3y agoAnd HN!
- goosinmouse 3y agoHN about anything finance related is funny. Basically same level as reddit/gamestonks
- deleted 3y ago[deleted]
- queuebert 3y agoWhile I get your point and agree, there apparently is no Dunning-Kruger effect. It was determined to be yet another case of faulty data analysis in behavioral psych.[1] 1. https://economicsfromthetopdown.com/2022/04/08/the-dunning-kruger-effect-is-autocorrelation/ https://economicsfromthetopdown.com/2022/04/08/the-dunning-k...
- lupire 3y agoThat paper was thoroughly rebutted when posted to HN. But Gwern has written about some earlier debunkings of the D-K effect, ad furthermore D-K was never about the popular misconception of D-K ("incompetent people think they are more competent than competent people").
- alecco 3y agoI don't think that's a mainstream take. But I can settle with: https://en.wikipedia.org/wiki/Overconfidence_effect https://en.wikipedia.org/wiki/Overconfidence_effect Or even https://en.wikipedia.org/wiki/Grandiose_delusions https://en.wikipedia.org/wiki/Grandiose_delusions
- c7b 3y agoI'm sure lots of people have been thinking/working along that direction. I think the idea of combining LLMs with formal verification/proof assistant tools (Lean, Coq, Isabelle,...) is particularly interesting. Anyone is aware of any major groups working on open source solutions for this?
- tomatoadventure 3y agoThat sounds really cool, but I am having trouble finding any information about them to get in touch :(