5 ms·
> I don't think researchers in math/TCS will be made obsolete, but I think it will instead no longer make sense to work on any low-hanging, or even medium-hangi
by rakel_rakel 3mo ago
> I don't think researchers in math/TCS will be made obsolete, but I think it will instead no longer make sense to work on any low-hanging, or even medium-hanging (you know what I mean) fruit. We'll be needed for problems where actual novel approaches are needed.
I wonder how this compares to what we see happening with "juniors" in software development?
In math research, do you also get the training for the profession from working on the low hanging fruits for a while, to then move to the medium-hanging, and later go on to work on previously unsolved stuff?
- skybrian 3mo agoThis apparently required a 10-page prompt. It seems like someone needs to know enough to write it?
- jvanderbot 3mo agoYeah, back to the gold-in-gold out use of LLMs.
- bredren 3mo agoI was thinking this past week I have gotten so lazy w my prompting via CLIs. Back in the before I had put such discipline into my prompting and supporting context. Now I’m like, “look here and here and here are some tools, and /skill /skill okay go.” Or “restate this request in your own words and enrich it as appropriate handling any gaps. Okay go”
- Quothling 3mo agoWe're also at the point where you can roll out context to your entire organisation. I created an app for our m365 Cowork and deployed it to everyone who develops software. It does a couple of things, but it main knows our compliance policies and can guide developers through writing the documentation needed for NIS2 compliance. It also guardrails against non-approved packages, and helps developers find alternatives, or if none can be reasonably found, how to get a new package/dependency approved (or rejected). A few months back this would be something every developer kind of did on their own. Maybe they shared skills, we certainly encouraged it and tried to do all the change management things, but nobody really had the same versions of the skills. Which was horrible in the deployment pipelines, something like the compliance documentation often had to go back and forth several times before it could be approved. Now it's just there, for everyone. In a year or two, I expect a lot of these things to have become even more standardized. So that we don't even really have to build our own apps, but can simply use the ones in the catalog with minimal configuration (and that config will likely only be necessary because I'm from a tiny country that nobody will maintain standards for).
- bredren 3mo agoYes. I worked on two large monorepos, about one quarter per project. Maybe 80+ devs on first, maybe 40+ on second. Both were not sure how to describe this but ultra-high-velocity agentic driven development efforts. On the first, there were ~no shared skills. There were some requirements set up but they were not minded properly and became stale / ate context for little gain. The hardest hit was in E2E tests which would flake and create long running, too-often failing CI. People would disable them, because they were not reliable and velocity was so high, no one was happy w them. I maintained my own set of skills and CLIs to back them. I'd share them if they came up but it was like the old days of manage your own stuff. Not much credit for building and sharing devex tooling to the team. But then on the second one we were in better shape--we had vendoring set up to distro skills automatically. Before the project was well underway, I put time into understanding how all of our tests aught to be written. Finding the forbidden things, etc, getting review from our best test folks and ultimately landed on a `/test` that routed across all possible test types. Like night and day. Instead of finding out while trying to get a release out the door that some corner of the project had a handful of flakes, tests were written the right way from the start. Like, it was beautiful. And I don't think devs noted difference while building. Only that there was an absence of BS in CI. Hard to quantify the lack of pain, but it was big!
- danielbln 3mo agoThis made me chuckle because it's so true. So much detailed steering and finagling in the past, now I point the agent to a bunch of information sources, skills, similar repositories that might hold useful input and tell it very roughly what I need and off it goes, I'll grab coffee.
- bredren 3mo agoHave you tried this: Look here and here and here are some tools, and /skill /skill [repo of folder paths etc] and here is what needs to happen: [stuff]. --- Restate this request in your own words and enrich it as appropriate handling any gaps. ? It is an ultra-lite way to plan, I suppose. I like the format because: - I still get to put all my thinking into the request but then easily override the instruction - It is interesting to see my casual typo-riddled blast professionalized and improved upon. - Sometimes it surfaces useful questions that can save some time up front. I think the models are doing this anyway, but I find the words "enrich" and "gap" are well understood by models and they demonstrate it in the response to the above pattern. Anyhow, to get back to the point, there are still prompt-level tricks--but ultimately if repeated, should probably also be built into skills themselves!
- ch4s3 3mo agoCertainly. This feels similar, to me, to how building complex software with LLMs works today in practice. You need to know a lot to set up goals and guardrails and verify outputs. For me, making the bits change was always the fun part, not tangling with text in my editor, though that had its moments.
- dwohnitmok 3mo agoThe author also used GPT-5.6 to write the prompt. This did involve giving GPT-5.6 access to his previous work and a back and forth process (so definitely still used the author's expertise to some degree), but the prompt itself is also largely AI generated.
- lucianbr 3mo agoWhat's the difference between using GPT to write the prompt to GPT, and "thinking"? The LLM uses the first tokens to predict more tokens, and then uses those tokens to predict even more tokens.
- petra 3mo agoThe ability to correct the plan/prompt.
- dwohnitmok 3mo agoRoughly speaking it is the difference between having a contractor go out and do some work and having that same contractor first come up with a plan to do some work, run that by you, and then go out to do that work. Part of it is as a another comment in this chain mentions the chance to review the prompt. Part of it is that it forces the AI system to plan things in a certain order, in much the same way that forcing the contractor to write the plan out first forces the contractor to proceed in a certain predefined order that may (or may not!) be better at getting to a final answer.
- JustFinishedBSG 3mo agoMy experience may not be entirely representative because to be entirely honest I’m not exactly a great researcher and there are brilliant PhD students. That said it indeed was my experience that in the pre-PhD / early PhD period ( or even longer … ) your advisor proposes (gives) you pretty low hanging stuff that he mostly already knows how to solve, at least at a high level, with the expectation that it will teach you to use the mathematical tools you need.
- Quothling 3mo agoAround here AI isn't really more of a threat to juniors than it is to seniors. It's a threat to the people who have been taught "recipies" rather than applied computer science. You can have excellent seniors who can do TDD, DRY, SOLID and so on, who also happen to have no idea what a L1 cache miss is. The current AI models know all of those things, but they struggle applying them correctly without someone piloting them. Even in the energy industry where I work, where you'd think it would be obvious from the context that you should prioritize runtime safety over debug safety, the current AI models struggle to do so. As far as seniority goes, though. If we can find a young developer with little experience who actually knows computer science, we're much more likely to hire them... Since they are cheaper. This isn't something which is unique to software development though. We're currently building enterprise AI apps that we can deploy into the AI agents working for anyone of our employees. The key thing we're currently seeing is that the people in a team who are the ones that everyone turn to for advice, are the only people who aren't in "danger". Even people who are great at their jobs are being outperformed by AI in many cases. I think it'll be a massive challenge for our society in the coming years. Maybe we're even going to get to the point where the AI will also be capable of replacing a lot of the "domain experts". Right now that seems far out, but then, if you had asked me about AI four months ago I would've told you it was all hype.
- marcosdumay 3mo agoSo... The AIs with no model of the world are replacing software developers that have no model of the world?
- rakel_rakel 3mo agoInteresting, thanks. I don't know where "around here" is, but the signals I've seen in a lot of articles is that the demand for junior software people has taken a dive since a year or two back, with student programs etc getting cancelled. One googler said they were getting a junior to their team and that was kind of a big deal because it hadn't happened in that whole department for a long time. In relation to that, I guess my question becomes: if the same thing will happen in math research, who will write the ten page math proof prompts in the future?
- vatsachak 3mo agoMath is way more automatable than programming. In math, a proof is a proof. We don't know if we can get there and so getting there is the hard part. In software, we always know that we can solve the problem. So HOW to solve the problem is the hard part. Because the type of solution involves maintainability, which involves planning, LLMs suck at it. This leads to "LLM slop code" whereby the LLM creates ad-hoc convoluted logic with redundancies and no reuse of existing standard library batteries. Unless you're a Grothendieck who gets mad at Deligne for not solving the Weil's conjecture "THE RIGHT WAY", software is fundamentally different than math in this respect. So I'll say it again, AI will win a fields medal for before managing a McDonald's simply because there are enough big problems within arms reach than their current capacity to plan over time
- sashank_1509 3mo ago> So I'll say it again, AI will win a fields medal for before managing a McDonald's simply because there are enough big problems within arms reach than their current capacity to plan over time AI can manage a McDonald’s already. If manage means directing humans to do something to ensure the store is running. If manage means running robots, then yes maybe that is 5 years away but just directing humans to run a store, that is possible right now.
- vatsachak 3mo agoNo it can't. Show me a business which uses in context learning to manage a McDonald's
- wanderlust123 3mo agoWell that’s a problem of incentives. Why would a manager outsource their own job to an AI?
- vatsachak 3mo agoIt's not a problem of incentives. Every executive wants to inject LLMs everywhere these days. If they haven't somewhere it means that it does not work.
- nicf 3mo agoI was trained as a mathematician and worked as a math researcher for a little while (now working as a private tutor), and based on my experience I'd say this description is basically right, with one extra wrinkle. In order to get a Ph.D., you have to do some sort of original research, so in that sense you're working on "previously unsolved stuff" basically right from the start. But that doesn't entail doing anything all that ground-breaking; most Ph.D. dissertations (very much including mine!) contain work that a more senior researcher in the same subfield could probably have produced without too much difficulty. The software development analogy is a pretty good one: a lot of the point of getting junior researchers to do research is to help train them to one day become senior researchers, and often the work itself is nothing all that special. Given the trajectory of these LLM proofs, this seems like it's going to have to change pretty soon, and to be honest I'm pretty grateful that I'm not in charge of deciding what that's going to look like, because I don't have any good ideas! I'm actually pretty worried about the future of the field.
- darkstarsys 3mo agoIndeed. Perhaps my article here will be of interest to some: https://blog.oberbrunner.com/blog/ai-math-as-humanities/ https://blog.oberbrunner.com/blog/ai-math-as-humanities/
- pfdietz 3mo ago> In order to get a Ph.D., you have to do some sort of original research, China has now introduced "practical PhDs" where you have to build some practical machine instead.
- phillip_kerger 3mo agoI would agree with your take. I (author of the post & paper) learned a ton from working on small parts of problems my PhD advisor was doing a lot of the heavy lifting on, and later also from getting some results that were essentially putting together the right pieces that already existed followed by some deep-in-the-weeds analysis.