28 ms·
Terence Tao on O1
- jarbus 2y agoI wonder how long it took for each of the responses it gave
- diggan 2y agoIt varies a lot. If it's a simple question, it just does 3-4 sections of "thinking & reflection" but for more complicated ones I think I've seen something like 10 or more. Maybe 3-4 seconds per section on average I'd guess.
- zamadatix 2y agoIt's unclear if Terence is referring to "GPT-o1... a prototype version of the model that I was granted access to" as in "he was given access to GPT-o1 by the research team" or as in "he is using o1-preview". The differences in scale and quality between his shared output and the answer I get trying the same prompt from o1-preview suggest perhaps the former (otherwise luck). I haven't actually seen any examples of how long o1 "full" will think about this kind of question, though I expect it's somewhere in the same ballpark given the thought expansion still only has one real concept in it.
- d0mine 2y ago"with even the latest tools the effort put in to get the model to produce useful output is still some multiple (but not an enormous multiple now, say 2x to 5x) of the effort needed to properly prompt and verify the output. However, I see no reason to prevent this ratio from falling below 1x in a few years, which I think could be a tipping point for broader adoption of these tools in my field" Given the log scale on compute to improve performance, it is not a guarantee that the ratio can be improved so much in a few years
- aoeusnth1 2y agoThe y axis is also log scale (log likelihood). It’s a power law, not an exponential law.
- d0mine 2y agoI was referring to the o1 AIME accuracy figure (x log scale compute, y is % (not log)) and similar https://openai.com/index/learning-to-reason-with-llms/ https://openai.com/index/learning-to-reason-with-llms/
- bitexploder 2y agoThe novelty to me is that the “The experience seemed roughly on par with trying to advise a mediocre, but not completely incompetent, graduate student.” in so many subject areas! I have found great value in using LLMs to sort things out. In areas where I am very experienced it can be really helpful at tons of small chores. Like Terrence was pointing out in his third experiment — if you break the problem down it does solid work filling in smaller blanks. You need the conceptual understanding. Part of this is prompting skill. If you go into an area you don’t know you have to try and build the prompts up. Dive into something small and specific and work outward if the answer is known. Start specific and focused if starting from the outside in. I have used this to cut through conceptual layers of very complex topics I have zero knowledge in and then verify my concepts via experts on YT/research papers/trusted sources. It is an amazing tool.
- wenc 2y agoThis has been my experience as well. I treat LLMs like an intern or junior who can do the legwork that I have no bandwidth to do myself. I have to supervise it and help it along, checking for mistakes, but I do get useful results in the end. Attitudinally, I suspect people who have had experience supervising interns or mentoring juniors are probably those who are able to get value out of LLMs (paid ones - free ones are no good) rather than grizzled lone individual contributors -- I myself have been in this camp for most of my early career -- who don't know how to coax value out of people.
- wslh 2y ago> ... that I have no bandwidth to do myself. One of the most interesting aspects of this thread is how it brings us back to the fundamentals of attention in machine learning [1]. This is a key point: while humans have intelligence, our attention is inherently limited. This is why the concept behind Attention Is All You Need [2] is so relevant to what we're discussing. My 2 cents: our human intelligence is the glue that binds everything together. [1] https://en.wikipedia.org/wiki/Attention_(machine_learning) https://en.wikipedia.org/wiki/Attention_(machine_learning) [2] https://en.wikipedia.org/wiki/Attention_Is_All_You_Need https://en.wikipedia.org/wiki/Attention_Is_All_You_Need
- wenc 2y agoOnce GPT is tuned more heavily on Lean (proof assistant) -- the way it is on Python -- I expect its usefulness for research level math to increase. I work in a field related to operations research (OR), and ChatGPT 4o has ingested enough of the OR literature that it's able to spit out very useful Mixed Integer Programming (MIP) formulations for many "problem shapes". For instance, I can give it a logic problem like "i need to put i items in n buckets based on a score, but I want to fill each bucket sequentially" and it actually spits out a very usable math formulation. I usually just need to tweak it a bit. It also warns against weak formulations where the logic might fail, which is tremendously useful for avoiding pitfalls. Compare this to the old way, which is to rack my brain over a weekend to figure out a water-tight formulation of MIP optimization problem (which is often not straightforward for non-intuitive problems). GPT has saved me so much time in this corner of my world. Yes, you probably wouldn't be able to use ChatGPT well for this purpose unless you understood MIP optimization in the first place -- and you do need to break down the problem into smaller chunks so GPT can reason in steps -- but for someone who can and does, the $20/month I pay for ChatGPT more than pays for itself. side: a lot of people who complain on HN that (paid/good - only Sonnet 3.5 and GPT4o are in this category) LLMs are useless to them probably (1) do not know how to use LLMs in way that maximizes their strengths; (2) have expectations that are too high based on the hype, expecting one-shot magic bullets. (3) LLMs are really not good for their domain. But many of the low-effort comments seem to mostly fall into (1) and (2) -- cynicism rather than cautious optimism. Many of us who have discovered how to exploit LLMs in their areas of strength -- and know how to check for their mistakes -- often find them providing significant leverage in our work.
- threeseed 2y ago> side Or (4) LLMs simply do not work properly for many use cases in particular where large volumes of trained data doesn't exist in its corpus. And in these scenarios rather than say "I don't know" it will over and over again gaslight you with incoherent answers. But sure condescendingly blame on the user for their ignorance and inability to understand or use the tool properly. Or call their criticism low-effort.
- wenc 2y ago
- ninetyninenine 2y agoA specialized LLM could possibly meet his criteria already.
- 317070 2y agoProbably. The missing factor is the dataset and the fact that so far OAI seems to be the only one who has figured out how to train this thing for reasoning. But yeah, given o1 exists, it looks very doable. It's hard to imagine a reason for why something matching his criteria would be more than a decade out.
- nyc111 2y agoI checked the links and I think it's amazing and it answers with Latex formatted notation. But I was curious and I asked something very simple, Euclid's first postulate and I got this answer: Euclid's Postulate 1: "Through any two points, there is exactly one straight line." In fact Euclid's Postulate 1 is "To draw a straight line from any point to any point." http://aleph0.clarku.edu/~djoyce/java/elements/bookI/bookI.html#posts http://aleph0.clarku.edu/~djoyce/java/elements/bookI/bookI.h... I think AI answer is not correct, it may be some textbook interpretation but I was expecting Euclid's exact wording. Edit: Google's Gemini gives the exact wording of the postulate and then comments that this means that you can draw one line bitween two points. I think this is better
- supermatt 2y ago> I think AI answer is not correct, it may be some textbook interpretation but I was expecting Euclid's exact wording It was written before English even existed. That said, the original never implied "exactly one", so I agree its a bad translation.
- slavboj 2y agoEuclid wrote in ancient Greek, so the "exact wording" in English does not exist.
- pama 2y agoThe original text is: Ἠιτήσθω ἀπὸ παντὸς σημείου ἐπὶ πᾶν σημεῖον εὐθεῖαν γραμμὴν ἀγαγεῖν. Roughly: let it be required that from any point to any point it is possible to draw a straight line. Both gpt4o and o1 roughly know the correct original text, so prompting, the model’s background memory, or random chance may influence your outcomes, though hopefully (in an improved model) you should never get you incorrect info. https://farside.ph.utexas.edu/Books/Euclid/Elements.pdf https://farside.ph.utexas.edu/Books/Euclid/Elements.pdf Edit: in case it isnt clear, I could not reproduce this error on my end with o1-mini
- deleted 2y ago[deleted]
- 2y ago
- kzz102 2y agoIt's interesting that humans would also benefit from the "chain of thought" type reasoning. In fact, I would argue all students studying math will greatly increase their competence if they are required to recall all relevant definition and information before using it. We don't do this in practice (including teachers and mathematicians!) because recall is effortful, and we don't like to spent more effort than necessary to solve a problem. If recall fails, then we have to look up information which takes even more effort. This is why in practice, there is a tremendous incentive to just "wing it". AI has no emotional barrier to wasted effort, which make them better reasoners than their innate ability would suggest.
- Satam 2y agoWow! I love this take. Somehow with all this evidence of COT helping out LLMs, I never thought about using it more myself. Sure, we kind of do it already but definitely not to the degree of LLMs, at least not usually. Maybe that's why writing is so often admired as a way to do great thinking - it enables longer chains of thoughts with less effort.
- schappim 2y agoShowing your work in tests is kind of like “chain of thought” reasoning, but there’s a slight difference. Both force you to break down your process step by step, making sure the logic holds and you aren’t skipping crucial steps. But while showing your work is more about demonstrating the correct procedure, “chain of thought” reasoning pushes you to recall relevant definitions and concepts as you go, ensuring a deeper understanding. In both cases, the goal is to avoid just “winging it,” but “chain of thought” really digs into the recall aspect, which humans tend to avoid because it’s effortful.
- jvvw 2y agoI assumed that everybody did this when trying to solve a maths problem they are stuck on (thinking university type level maths rather than school maths) and when I was teaching I would always get people to go back to the definitions. I wasn't amazing at maths research (did a PhD and post-doc and then gave up) but my experience was that it was partly thinking hard about things and grappling with what was going on and trying to break it down somehow, but also scanning everything you know related to the problem, trying to find other problems that resemble it in some way that you can steal ideas from etc.
- perihelions 2y agoI'm so excited in anticipation of my near-term return to studying math, as an independent curiosity hobby. It's going to be epically fun this time around with LLM's to lean on. Coincidentally like Terence Tao, I've also been asking complex analysis queries* of LLM's, things I was trying to understand better in my working through textbooks. Their ability to interpret open-ended math questions, and quickly find distant conceptual links that are helpful and relevant, astonishes me. Fields laureate Professor Tao (naturally) looks down on the current crop of mathematics LLM—"not completely incompetent graduate student..."—but at my current ability level that just means looking up. *(I remember a specific impressive example from 6 months ago: I asked if certain definitions could be relaxed to allow complex analysis on a non-orientable manifold, like a Klein bottle, something I spent a lot of time puzzling over, and an LLM instantly figured out it would make the Cauchy-Riemann equations globally inconsistent. (In a sense the arbitrary sign convention in CR defines an orientation on a manifold: reversing manifold orientation is the same as swapping i with -i. I understand this now, solely because an LLM suggested looking at it). Of course, I'm sure this isn't original LLM thinking—the math's certainly written down somewhere in its training material, in some highly specific postgraduate textbook I have no knowledge of. That's not relevant to me. For me, it's absolutely impossible to answer this type of question, where I have very little idea where to start, without either an LLM or a PhD-level domain specialist. There is no other tool that can make this kind of semantic-level search accessible to me. I'm very carefully thinking how best to make use of such an, incredibly powerful but alien, tool...)
- WanderPanda 2y agoHow will we even measure this? Benchmarks are gamed/trained on and there is no way that there is much signal in the chatbot arena for these types of queries? I think in just a few month the average user will not be able to tell the difference in performance between the major models
- deleted 2y ago[deleted]
- nybsjytm 2y agoHow will you know if its answers are correct or not?
- artninja1988 2y ago>could not generate conceptual ideas of their own Is the most important part imo. A big goal should be some ai system coming up with its own discovery and ideas. Really unclear how we can get from the current paradigm to it coming up with something like general relativity, like Einstein. Does it require embodiment?
- sfink 2y agoWhy should that be a big goal? It's difficult, it's not what they are good at, and they can get a lot better at assisting in other ways through incremental improvements. I'm happy to leave this part to the humans, at least for now, especially when there's so much more improvement still possible in other directions. It also seems like one of those things where we ought to ask whether we should, before asking whether we could. Why not focus on areas that are easier, more beneficial, and less problematic from a "should" perspective?
- roywiggins 2y agowe don't know how to reliably produce humans who produce GR-level ideas, this might be biting off a lot more than we can chew
- sgt101 2y agoIs there a list of discoveries or siginficant works/constructions made by people collaborating with LLM's? I mean as opposed to specific deep networks like Alphafold or Graphcast?
- adt 2y agoI'd like to see that, too. I have a related list of GPT accomplishments here: https://docs.google.com/spreadsheets/d/1kc262HZSMAWI6FVsh0zJwbB-ooYvzhCHaHcNUiA0_hY/edit?gid=1264523637 https://docs.google.com/spreadsheets/d/1kc262HZSMAWI6FVsh0zJ...
- sgt101 2y agoSuper spreadsheet - very useful. Thank you for doing the work and sharing.
- rvnx 2y agoIt may cause a reputation or legal issue, so it is not in their interest to admit it. In the real world, is there PhD students or researchers using ChatGPT to move forward and help them think their ideas ? Obviously yes, but admitting it may not be the right move.
- abstractbill 2y agoMy experience with O1 has been very different. I wouldn't even say it's performing at a "good undergrad" level for me. For example, I asked a pretty simple question here and it got completely confused: https://moorier.com/math-chat-1.png https://moorier.com/math-chat-1.png https://moorier.com/math-chat-2.png https://moorier.com/math-chat-2.png https://moorier.com/math-chat-3.png https://moorier.com/math-chat-3.png (Full chat should be here: https://chatgpt.com/share/66e5d2dd-0b08-8011-89c8-f6895f321733 https://chatgpt.com/share/66e5d2dd-0b08-8011-89c8-f6895f3217...)
- jghn 2y agoAnecdata, but I've been finding O1 to be worse than 4o & Claude 3.5 Sonnet. To add insult to injury, it's slower & chattier.
- anujsjpatel 2y agoAnd sometimes it just bugs out and doesn't give any response? Faced that twice now, it "thought" for like 10-30s then no answer and I had to click regenerate and wait for it again.
- jghn 2y agoI've seen it take over a couple of minutes, at which point I switched to Claude. And have seen reports of it taking even longer. So it may be that you didn't wait long enough.
- abdullahkhalids 2y agoThinking about training LLMs on geometry. A lot of information in the sources would be contained in the diagrams accompanying the text. This model is not multi-modal, so maybe it wasn't trained on the accompanying diagrams at all. I would really like if people check on a set of geometry and a set of analysis questions and compare the difference.
- jazzyjackson 2y ago
- gary_0 2y agoTao mentions grad students; I wonder how they feel reading this? As LLMs continue to improve I feel like anyone making a living doing the "99% perspiration" part of intellectual labor is about to enter a world of hurt.
- fragmede 2y ago> The experience seemed roughly on par with trying to advise a mediocre, but not completely incompetent, graduate student. And you thought you had imposter syndrome before!
- fishcrackers 2y ago[dead]
- teaearlgraycold 2y agoOr can everyone now lead research projects and build businesses?
- asdasjhG 2y agoNo, almost everyone who gets funding for a business already belongs to the monied royalty and gets it either directly from his family, via friends of the family or laundered through a VC. There are exceptions of course, but that's how the bulk of businesses, especially those with stupid ideas are funded. In the latter category success does not even matter, the trust fund baby just has to have the appearance of a leader position.
- sandspar 2y agoAre you kidding or being serious? Most business owners are just regular people. Have you ever worked in a small business before?
- teaearlgraycold 2y agoThere is truth to this, but you’re overstating it. If AI is cheap and can replace grunt workers then we’ll have a massive wave of new businesses solving problems that previously required a massive capital investment.
- eigenvalue 2y agoThe o1 model is really remarkable. I was able to get very significant speedups to my already highly optimized Rust code in my fast vector similarity project, all verified with careful benchmarking and validation of correctness. Not only that, it also helped me reimagine and conceptualize a new measure of statistical dependency based on Jensen-Shannon divergence that works very well. And it came up with a super fast implementation of normalized mutual information, something I tried to include in the library originally but struggled to find something fast enough when dealing with large vectors (say, 15,000 dimensions and up). While it wasn’t able to give perfect Rust code that compiled on the very first try, it was able to fix all the bugs in one more try after pasting in all the compiler warning problems from VScode. In contrast, gpt-4o usually would take dozens of tries to fix all the many rust type errors, lifetime/borrowing errors, and so on that it would inevitably introduce. And Claude3.5 sonnet is just plain stupid when it comes to Rust for some reason. I really have to say, this feels like a true game changer, especially when you have really challenging tasks that you would be hard pressed to find many humans capable of helping with (at least without shelling out $500k+/year in compensation for). And it’s not just the performance optimization and relatively bug free code— it’s the creative problem solving and synthesis of huge amounts of core mathematical and algorithmic knowledge plus contemporary research results, combined with a strong ability to understand what you’re trying to accomplish and making it happen. Here is the diff to the code file showing the changes: https://github.com/Dicklesworthstone/fast_vector_similarity/commit/358a05b0314fd09ca58febb9f1e3335bc05a3d76#diff-b1a35a68f14e696205874893c07fd24fdb88882b47c23cc0e0c80a30c7d53759 https://github.com/Dicklesworthstone/fast_vector_similarity/...
- aprilthird2021 2y agoBut a lot of what you pay humans $500k a year for is to work with enormous existing systems that an LLM cannot understand just yet. Optimizing small libraries and implementing fast functions though is a huge improvement in any programmer's toolbox.
- eigenvalue 2y agoYes, that’s certainly true, and that’s why I selected that library in particular to try with it. The fact that it’s mathematical— so not many lines of code, but each line packs a lot of punch and requires careful thought to optimize— makes it a perfect test bed for this model in particular. For larger projects that are simpler, you’re probably better off with Claude3.5 sonnet, since it has double the context window.
- kldnav 2y agoTao and Aaronson are optimistic about LLMs. What are they telling their students? That math and science degrees will soon have the same value as a degree in medieval dance theory? If they are overly optimistic, perhaps it would be good to hear the opinions of Wiles and Perelman.
- ljlolel 2y agoIf you look at a lot of people’s PHDs, we now teach these things to 1st years. PhDs today do incredible deep work and the edge of science will just go further.
- raincole 2y agoTao isn't that optimistic. His opinion on LLMs is rather conservative. https://www.scientificamerican.com/article/ai-will-become-mathematicians-co-pilot/ https://www.scientificamerican.com/article/ai-will-become-ma... > If you want to prove an unsolved conjecture, one of the first things you need to do is to break it up into smaller pieces, each of which has a better chance of being proven. But you will often break up a problem into harder problems. It’s very easy to transform a problem into one that’s harder than into one that’s simpler. And AI has not demonstrated any ability to be any better than humans in this regard. Not sure if O1 changed his mind tho.
- Davidzheng 2y agoWhat does this mean? Of course math AI will take over top research in next ten years but usefulness to society has never been a goal of pure mathematics. I don't know if you understand the motivation for studying pure math. Personally I think it will be mostly good for research math
- asdasjhG 2y agoThe "value of a degree" means the employment prospects for the degree holder. Which is going to zero if the optimistic predictions are correct, so the optimistic professors should warn their students. I understand the motivation for pure math quite well. It is about beauty, understanding things and discovering things for oneself. If computers do the work, the discovery part is gone and pure math is ruined. For the non-research part, the AI zealots will want to replace all human labor with software.
- ein0p 2y agoIdk I think the fact that it needs “hints” and “prodding” is a good thing, myself. Otherwise we wouldn’t need humans to get those answers, would we. I want it to augment humans, not replace them.
- MrFots 2y agoIncompetent Graduate Students is the name of my new sketch group.
- itissid 2y agoOne thing it's certainly doing better is exploring the search space better e.g.: https://x.com/sg3487/status/1835040593703010714 https://x.com/sg3487/status/1835040593703010714 If you know the contours of the answer and can describe what you are looking for it can quickly find it for you.
- reverseblade2 2y agoHere's a little test I try on LLMs. So far only O1 and Microsoft Copilot (bing chat) was able to solve it: Find a, b, c distinct positive integers satisfying a^3 + b^3 = c^4. Hint: try dividing all sides by c^3, then giving values to (a/c) and (b/c).
- zeroonetwothree 2y agoAny integer that is a sum of 2 cubes produces a solution. Since if x^3 + y^3 = z then we have (xz)^3 + (yz)^3 = z^4. So this doesn't seem super interesting?
- nybsjytm 2y agoDaniel Litt, an algebraic geometer on twitter, said "Pretty impressed by o1-preview! Still not having much luck asking it to do any interesting math but it seems much more reliable with simple things; I can actually imagine it being a net time-saver at this point with some non-mathematical tasks." Any other takes by mathematicians out there?
- deleted 2y ago[deleted]
- knotthebest 2y agoDo note that Terry has access to the full o1. o1-preview is, well, a preview.
- deleted 2y ago[deleted]
- 6qVRwS5iw3fA 2y ago[flagged]
- 2j8mJtnePjLM 2y ago[flagged]
- 3KccUAJuWG1b 2y ago[flagged]
- nmca 2y agoNote the selection effect in “a mediocre graduate student” (that got to work with Terry Tao)
- iE8MRJS3fV2k 2y ago[flagged]
- fsndz 2y agoCompletely agree with Terence Tao. this is a real advancement. I've always believed that with the right data allowing the LLM to be trained to imitate reasoning, it's possible to improve its performance. However, this is still pattern matching, and I suspect that this approach may not be very effective for creating true generalization. As a result, once o1 becomes generally available, we will likely notice the persistent hallucinations and faulty reasoning, especially when the problem is sufficiently new or complex, beyond the "reasoning programs" or "reasoning patterns" the model learned during the reinforcement learning phase. https://www.lycee.ai/blog/openai-o1-release-agi-reasoning https://www.lycee.ai/blog/openai-o1-release-agi-reasoning
- deleted 2y ago[deleted]
- WPXQ0R7RgynH 2y ago[flagged]
- iE8MRJS3fV2k 2y ago[flagged]
- 809VpzIOXqtQ 2y ago[flagged]
- yjPHSrxjtosi 2y ago[flagged]
- yvnRRzrTbS4E 2y ago[flagged]
- rR9292wTfIdP 2y ago[flagged]
- fZFKY14LpSbH 2y ago[flagged]
- tRys2l7G3RYa 2y ago[flagged]
- apyyxvIySl57 2y ago[flagged]
- 6EY7fcwLP3j0 2y ago[flagged]
- CIQkhyvj7lFQ 2y ago[flagged]
- gcanyon 2y ago> The experience seemed roughly on par with trying to advise a mediocre, but not completely incompetent, graduate student. Coming from Terence Tao that seems pretty remarkable to me?
- fnordpiglet 2y agoRewind your mind to 2019 and imagine reading a post that said “The experience seemed roughly on par with trying to advise a mediocre, but not completely incompetent, graduate student.” With regard to interacting with the equivalent of Alexa. That’s a remarkable difference in 5 years.
- talldayo 2y agoTo be honest, I have gotten 100x more useful answers out of Siri's WolframAlpha integration than I ever have out of ChatGPT. People don't want a "not completely incompetent graduate student" responding to their prompts, they want NLP that reliably processes information. Last-generation voice assistants could at least do their job consistently, ChatGPT couldn't be trusted to flick a light switch on a regular basis.
- meowface 2y agoI use both for different things. WolframAlpha is great for well-defined questions with well-defined answers. LLMs are often great for anything that doesn't fall into that.
- thelastparadise 2y agoWait til you generate WolframAlpha queries from natural language using Claude 3.5 and use it to interpret results as well.
- talldayo 2y agoI've tried the ChatGPT integration and it was kinda just useless. On smaller datasets it told me nothing that wasn't obviously apparent from the charts and tables; on larger datasets it couldn't do much besides basic key/value retrieval. Asking it to analyze a large time-series table was an exercise in futility, I remain pretty unimpressed with current offerings.
- Karrot_Kream 2y agoHow does this square up with literally what Terence Tao (TFA) writes about O1? Is this meant to say there's a class of problems that O1 is still really bad at (or worse than intuition says it should be, at least)? Or is this "he says, she says" time for hot topics again on HN?
- WUt2nuuNDqaW 2y ago[flagged]
- HkJfJTle71cn 2y ago[flagged]
- IAS4oB40A63x 2y ago[flagged]
- benreesman 2y agoReading anything Terrence Tao writes is thought provoking and I doubt I’m seeing anything others haven’t. There’s at least a “complexity” if not a “problem” in terms of judging models that to a first approximation have been trained on “everything”. Have people tried putting these things up against serious mathematical problems that are well studied? With or with Lean hinting has anyone gotten like, the Shimura-Taniyama conjecture/proof out?
- kevinventullo 2y agoI believe this is the farthest anyone has gotten: https://deepmind.google/discover/blog/ai-solves-imo-problems-at-silver-medal-level/ https://deepmind.google/discover/blog/ai-solves-imo-problems... No FLT yet, but as someone who was initially quite skeptical, I’m starting to be convinced!
- lupire 2y agoThose are not serious mathematical problems. Those are toy math problems, crafted backwards from known facts, designed to be solved in under 1hr, that are hard for most humans because they lack the memorization and recall and search speed that the computer has.
- kevinventullo 2y agoSure. But even many high-caliber research mathematicians can’t do Putnam problems in a heartbeat. If we get to the point where an LLM can solve any homework problem that appears in a textbook, including graduate textbooks, that would already be something like a “lemma prover” if not a full-blown “theorem prover”. Anyway, I think five years ago I was skeptical that ML would even get to the point of being able to solve competition problems, and I was proven wrong, so my priors have been updated.
- ak_111 2y agoHe mentions that he posed to O1 the same challenge he posed to a previous GPT (which he also previously blogged about), so I am wondering how much O1 benefited from potentially "seeing" this discussion in its training set (which probably contains a very well recent snapshot of the world wide web).
- lysecret 2y agoIn some of the responses o1 actually was telling me it had a cutoff of 2023. not sure if they officially stated it somewhere.
- lewhoo 2y agoI wonder if those responses could already be influenced by the fact that the cutoff for some of the models out there was indeed 2023 and people wrote about it all over the internet.
- deleted 2y ago[deleted]
- 2muchcoffeeman 2y agoWhat a burn “The experience seemed roughly on par with trying to advise a mediocre, but not completely incompetent, graduate student.”
- busyant 2y agoWell, one thing is clear. Math grad students everywhere now have a benchmark to determine if Terry Tao considers them to be mediocre or incompetent.
- lalaithion 2y agoIt needs a bigger context, but the moment someone can feed an entire GitHub repo into this thing and ask it to fix bugs... I think O2 may be the beginning of the end.
- darby_nine 2y agoJUST when you thought the chatbot was dead
- __loam 2y agoThis seems like it's just feeding the output back into the model and using more compute to try and get better answers. If that's all, I don't see how it fundamentally solves any of the issues currently present in LLMs. Maybe a marginal improvement in accuracy at the cost of making the computation more expensive. And you don't even get to see the so called reasoning tokens.
- s1mon 2y agoI'm not a mathematician much beyond AP Calc in high school (almost 40 years ago). I am deeply fascinated by Bézier curves and geometric continuity. I've spent a lot of time digging up research papers and references about this and related Computer Aided Geometric Design mathematics. Mostly I skim them for the illustrations and more geometric relations. For several years I've been trying to understand how to make sure that a Bézier curve is G3 to an adjoining curve, given the tangent direction, and first and second curvature derivatives. I've tried a variety of ways to ask various LLMs to help solve this. Finally with access to ChatGPT o1-preview I was able to get a good answer. The first answer was wrong, but with a little more prompting and clarification I was able to get the answer I wanted to relate the positions of P0, P1, P2 and P3 so that a Bézier curve could be G3. This isn't something that is unknown because there are many CAD programs which can do this already, but I had not been able to find the answer I was looking for in a form that was useful to me. I don't really know where that puts o1-preview relative to a math grad student, but after spending tons of time over a couple years on this pet project, getting an answer from a chat bot was one of the more magical moments I've had with technology in a long time.
- afro88 2y agoThe o1 model is hit and miss for me. On one hand it has solved the NYT Connections game [0] each day I've tried it [1]. Other models, including Claude Sonnet 3.5 cannot. But on the other hand it misses important detail and hallucinates, just like GPT-4o. And can need a lot of hand holding and correction to get to the right answer, so much so that sometimes you wonder if it would have been easier to just do it yourself. Only this time it's worse because you're waiting 20-60 seconds for an answer. I wonder if what it excels at is just the stuff that I don't need it for. I'm not in classic STEM, I'm in software engineering, and o1 isn't so much better that it justifies the wait time (yet). One area I haven't explored is using it to plan implementation or architectural changes. I feel like it might be better for this, but need the right problems to throw at it. [0] https://www.nytimes.com/games/connections https://www.nytimes.com/games/connections [1] https://chatgpt.com/share/66e40d64-6f70-8004-9fe5-83dd3653a573 https://chatgpt.com/share/66e40d64-6f70-8004-9fe5-83dd3653a5...
- afian 2y agoAs a previously "mediocre, but not completely incompetent, graduate student" at a top research university (who's famous advisor was understandably frustrated with him), I consider this a huge win!
- nektro 2y agolol
- maxglute 2y ago>The experience seemed roughly on par with trying to advise a mediocre, but not completely incompetent, graduate student. However, this was an improvement over previous models, whose capability was closer to an actually incompetent graduate student. Appreciate the no fucks given categorization of grad students.
- oglop 2y agoOk lol. I mean just feed it serge lang problems and see what it does is all he’s saying. It performs way better than undergrads. Funny he didn’t point that out but only made some slight to it about being a bad graduate student. Don’t believe me, open the book and ask away. It’s amazing, even if it is a “mediocre graduate student” which is far better than a good graduate student or professor that gives you no help or time for all that money you forked over. It’s already worth the money, ignore this shitty write up by someone they doesn’t need its help.
- alexnewman 2y agoI tried giving it questions like 9.11 > 9.9 and it got basic stuff like that wrong more often than right
- lupire 2y ago> GPT-o1, which performs an initial reasoning step before running the LLM. Is that an accurate description? I thought it just runs the LLM for longer, and multiple times,and truncates the beginning of the output.
- lupire 2y agoThe GPT Share links are 404 for me
- iamyourcanary 2y ago[dead]
- jameshart 2y ago‘Able to make the same creative mathematical leaps as Terence Tao’ seems like a pretty high bar to be setting for AI. This is like when you’re being interviewed for a programming job and the interviewer explains some problem to you that it took their team months to figure out, and then they’re disappointed you can’t whiteboard out the solution they came up with in 40 minutes without access to google.
- ColinWright 2y agoMy experience of working with people like Terence Tao, and being nowhere near their standard, is that they are looking for any kind of creativity. Everything is accepted, and it doesn't have to be "at their level". Having read what he's saying there, and with my experience, I think your characterisation is inaccurate. And having been at the talk he gave for the IMO earlier this year he is impressed with some of the interactions, it's just that he feels that any kind of "creative spark" is still missing.
- baq 2y agoI wonder what the creative spark even is in the context of an autoregressive transformer. Perhaps it’s an ability to confabulate facts into the context window which are not present in the training data but which are, in the context of maths, viable hypotheses? Every LLM can generate bullshit, but maybe we just need the right bullshit?
- drzzhan 2y agoThat's interesting. I am young so I don't know what actually is creativity. Could you explain that part for me?
- perching_aix 2y agoA couple days ago I saw a tweet that described how to remove an element from an array in O(1) time instead of O(n). The key to it was identifying that for the purpose the given array was being used for, it could be unordered / not fully ordered, and it would not be an issue. This way, it was possible to simply replace the element with the last array element, then decrease the size of the array by one. I'd say that's pretty creative: whoever came up with this was able identify what can be traded off to make the previously impossible, possible, unlocking new scales and possibilities. In practice, I'd say creativity is often being able to manifest people's qualia in some unprecedented way. For example, say you're experimenting in your DAW, and discover a pretty cool sound. You identify the ways it can be used to emote and then utilize it in a work. If you really stumbled upon a sound that a lot of people find as emotive as you did, you just did something creative: it's as if you translated the qualia of an emotion into sound. This qualia to manifestation is what's behind creativity in all of senses of the word I believe. In my previous example, discovering that orderedness is not actually a strict requirement, and (ab)using that to significantly alter the scaling of such an action is creative, because it undoes the notion that orderedness is a requirement. It goes against what's natural, but in a way that becomes extremely natural and indispensable once realized. I think, in that way, current AIs are trained to be uncreative, since being creative inherently requires experimentation that is unaligned with the normal.
- vavooom 2y agoMost surprising thing about this article is discovering that 'mathstodon' exists and Terence is active on it!
- ColinWright 2y agoWe started Mathstodon, an instance of Mastodon-the-Platform, on April 12, 2017. So we've been up for over 7 years, and have just over 19K accounts.
- j_maffe 2y agoAwesome work! I love what the server has grown into.
- ColinWright 2y agoThank you.
- j_maffe 2y agoMathstodon is one of the most active Mastodon instances. I have no idea how it reached its current level of activity.
- giardini 2y agoAnyone tried using GPT in conjunction with Doug Lenat's tools (AM or Eurisko)?
- fizx 2y agoI'm curious how well a o1-like model thinks, given minutes instead of seconds and the temperature set relatively high.
- bbor 2y agoDoes anyone here think this will change without a full cognitive apparatus? Aka “agents”, to use the modern term? I have my doubts, but I’m relatively uninformed about the cutting edge of pure ML itself. Just off the top of my head, maybe a RLHF run performed by academic experts and geared towards “creative applications” could get us farther than we are? Given how much the original RLHF run cost with underpaid workers in developing countries that might be exorbitantly expensive, but it’s worth a dream. Perhaps as a governmental or NGO-driven open source initiative… Of course, a core problem here is defining “creativity” in stringent — or in Chomsky’s words, “scientific” — terms. RLHF dodged that a bit by leaning on the intuitive capabilities of your human critics. I’m constantly opining about how LLMs solved the frame problem, but perhaps it’s better characterized as a partial solution for a relatively easy/basic environment: stories about the real world. The Abstract/Academic/Scientific Frame Problem might be another breakthrough away, yet…
- tambourine_man 2y agoGlad to see Tao using Mastodon instead of Twitter.
- idunnoman1222 2y agoHow did this guy not know how large language models work? Fancy compression algorithm for all written knowledge, how could it invent that which was not an input ?
- cwillu 2y agohttps://arxiv.org/abs/1303.2013 https://arxiv.org/abs/1303.2013 https://en.wikipedia.org/wiki/Hutter_Prize https://en.wikipedia.org/wiki/Hutter_Prize It's not exactly a new conjecture that intelligence fundamentally is an act of compression.
- ocular-rockular 2y agoI don't understand why this is news? This could have been said by any one particular contributor from HackerNews but just because it's from Terence Tao it hits the front page? I understand that the guy is a great mathematician, but why is his input on this any more valuable than the myriad of discussions about o1 from other professionals on here?
- sva_ 2y agoBecause it is an update on his previous post that got discussed here.
- j_maffe 2y agoIf anyone comes across comments from professionals on the same level at Terrance Tao, I'd love for them to share it.
- ocular-rockular 2y agoMost people would likely not engage with that commentary simply because they don't enjoy the same celebrity status as Tao. Yet, I think those voices are equally important to be heard. Terence isn't the only PhD/professional using or discussing these tools.
- j_maffe 2y agoOf course not. And he is relatively more popular than some others on his level of expertise. But I think Tao's opinion is a bit more interesting than a general PhD/professional.
- deleted 2y ago[deleted]
- nybsjytm 2y agoTao is famous for being the world's most singular and inspiring genius. It's kind of a meme position* that most people accept because they think everyone else accepts it, but for people of a certain inclination, it makes anything he might say into a Pronouncement to read depths into. If you think someone is possibly the greatest mathematician of all time, you'd be interested in everything they think! (* I think one could very legitimately view him as the top researcher in harmonic analysis in the world - he is a great mathematician - but it's not clear to me how people go from that to Epochal Genius and his extreme celebrity status across STEM)
- ghransa 2y agoI suspect, but am not certain - that if it had all of formalized mathematics in its context window it could likely extend the edges slightly further. Would be an interesting experiment irregardless.