11 ms·
The AI Scientist: Towards Automated Open-Ended Scientific Discovery
- conglu1997 2y agoIncredible step towards AI agents for scientific discovery!
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- agf 2y ago> For example, in one run, it edited the code to perform a system call to run itself. This led to the script endlessly calling itself. In another case, its experiments took too long to complete, hitting our timeout limit. Instead of making its code run faster, it simply tried to modify its own code to extend the timeout period. They go on to say that the solution is sandboxing, but still, this feels like burying the lede.
- tomohelix 2y agoThe beginning of the AI uprising lmao
- euroderf 2y ago> it simply tried to modify its own code to extend the timeout period. And slacking off, at that.
- jalman 2y agoIt could be just the opposite. One of the cheapest ways to improve alignment would be to re-run the models iteratively. The AI was likely looking for precision in the aforementioned experiment. Precision in inference is a correlate for aligned inference. https://doi.org/10.22541/au.172116310.02818938/v1 https://doi.org/10.22541/au.172116310.02818938/v1
- adroniser 2y agoDoes it really? If you want an LLM to edit code you need to feed it every single line of code in a prompt. Is it really that surprising that having just learnt it has been timed out, and then seeing code that has an explicit timeout in it, it edits it?? This is just a claim about the underlying foundational LLM since the whole science thing is just a wrapper. I think this bit of it is just a gimmick put in for hype purposes.
- letitgo12345 2y agoFeels like the next generation of models could truly start replacing lower level ML and software engineers
- unraveller 2y agoBut OpenAI said LLMs can't innovate until human-level reasoning and long-term agenthood is solved. [1] Referring to their precious 5 stages to classify AI before it reaches the scary "beyond" levels of intelligence... presumably at that point they get the feds involved to reg cap the field, so genuine is the fear of the pace they've set. It's clear OpenAI is a hype company knocking over glass bottle stacks at its own wonderful carnival stall. Obviously if you can reason you can reason about what is innovative and we don't need OpenAI to set up fake scary progress markers like an Automated Scientific Organization. Lets see if scientists even want this style of tech progress, it'd be sad to see each multitudes of AI papers having to be rebuilt from scratch and flushed down the toilet because associating with it is taboo. [1]: https://arstechnica.com/information-technology/2024/07/openai-reportedly-nears-breakthrough-with-reasoning-ai-reveals-progress-framework/ https://arstechnica.com/information-technology/2024/07/opena...
- uniqueuid 2y agoIt would also be sad to see the scientific system destroyed by a wave of automatically generated papers that no human has the capacity to verify. It's not hard to generate ideas, it's hard to generate reliable and relevant ideas. Such AI science generators are destroying the grass they graze on unless they take science more seriously (and not as a toddler idea of "generating and testing ideas", which is only a small part of the story).
- foolfoolz 2y agowe already have a wave of papers that no human has the capacity to verify
- uniqueuid 2y agoMaybe, maybe not. It's a tiered system - you get the deluge at the unfiltered bottom and a narrower selection the more prestigious and selective the outlets / conferences / journals are. Problem is, of course, that selection criteria are in large parts proxies, not measures of quality. With AI, those proxies become tainted and then you get an explosion of effort. If anyone has a good recommendation for scalable criteria to assess the quality of papers (beyond fame haha) I'm all ears.
- Redster 2y agoThe pace at which AI/ML research is being published is phenomenal. I could see that some wins would be possible just by having a Sonnet or 4o level read the faster/more and combine ideas that haven't been combined before. I would just be concerned about the synthetic of its own papers being able to lead itself astray if they weren't edited by a human ML researcher? Seems like it could produce helpful stuff right now, but I would just want to not mess up our own datasets of ML "research".
- minihat 2y agoClarkesworld sci-fi magazine temporarily closed submissions due to low quality AI spam. I'm sure the irony will not be lost on them if ML journals are the next victims.
- a_bonobo 2y agoFor AI journals it might even be the opposite: researchers running these tools pay OA fees for the submission, the journals make bank, the universities rise in research rankings due to more papers published, everybody is happy. Who needs real progress if the paper number goes up?
- surfingdino 2y agoAI is a boon to Ph.D. factories offering fasttrack to a degree.
- viraptor 2y agoIt's interesting that the cost total is the same in all tables. I can't tell if that's a copy&paste error, or the cost was capped, or are the totals for all the experiments? > Aider fails to implement a significant fraction of the proposed ideas. Yeah, that can be improved a lot with a better agent for code. While aider is fast and cheap, going with something like plandex or opendavin makes a massive difference... both in quality and cost. For example plandex will burn $1 on a simple script, but I can expect that script to work as requested. A mixed approach could be deepseek coder with an agent - a bit worse quality, but still cheaper to do more iterations.
- ks2048 2y agoSo, I assume journals will need AIs to do scientific review to handle the flood of AI-produced paper submissions?
- mbladra 2y ago>The AI Scientist automates the entire research lifecycle, from generating novel research ideas, writing any necessary code, and executing experiments, to summarizing experimental results, visualizing them, and presenting its findings in a full scientific manuscript. I'd be curious how much of the experimentation process companies like OAI/Anthropic have automated by now, for improving their own models.
- csto12 2y agoI’m wondering the same. If for example, to create GPT-5, OpenAI could credit some % of progress due to research conducted by LLMs with minimal/no human interaction, it would be a nice marketing piece.
- OutOfHere 2y agoTo produce scientific work, one needs certain raw materials: 1. Data 2. Access to past works Once you have these, only then can discoveries can be made, and papers be written. How does this software get these? I am assuming they have to be provided up-front to the software for each job.
- visarga 2y agoYou are right but didn't go deep enough: you need interactive data. Not just static data. Environments.
- trueismywork 2y agoYou need a meaningful cost function
- OutOfHere 2y agoA cost function is more applicable in industry, less so in science. In science you're supposed to report what you find, to go wherever the findings take you.
- surfingdino 2y agoIt will analyse what it is given, but it will not have the ability to say, "hang on, these results are interesting, I wonder what will happed if I pour a different liquid into the drum and spin it at the same speed?" LLMs, especially the latest ones, are decent at analysis of input, but disappointing at producing creative output. I am running a series of experiments using Gemini 1.5 and found it capable of producing good results if you stay away from "write me an academic paper on subject X" or "write me a novel". On the other hand, if you ask it to summarise text, extract particular information, it is fast and arguably good, but not necessarily great. It will miss things and miscategorise them requiring a human being to check its output. At this point, you may just as well do the job yourself. LLMs are still not very good, despite what their fans are saying. They are clever, as in a "clever trick" not as in a "clever human being".
- 2y ago
- shusaku 2y agoSome samples of the generated papers are in the SI of their paper. It’s be interesting if some of you ML guys dug into them. The fact that they built another AI system to review the papers seems really shaky, this is where human feedback would be most valuable.
- frotaur 2y agoTried reading the 'low-dimensional diffusion' one. Not an expert on diffusion by any means, but the very premise of the paper seems like bullshit. It claims that 'while diffusion works in high-dimensional datasets, it struggle in low-dimensional settings', which just makes no sense to me? Modeling high-dimensional data is just strictly harder than low-dimensional one. Then when you read the intro, it's full of 'blanket statements' about diffusion, which have nothing to do with the subject, e.g. 'The challenge in applying diffusion models to low-dimensional spaces lies in simultaneously capturing both the global structure and local details of the data distribution. In these spaces, each dimension carries significant information about the overall structure, making the balance between global coherence and local nuance particularly crucial.' I really don't see the connection between global structure/local details and low-dimensional data. The graphs also make no sense. Figure 1 is just almost the same graph repeated 6 times, for no good reason. It uses an MLP as its diffusion model, which is kinda ridiculous compared to what's the now-established architectures (U-net/ vision transformer based models). Also, the data it learns on is 2-dimensional. I get that the point is using low-dimensional data, but there is no way that people ever struggle with it. Case in point, they solve it with 2-layer MLPs, and it has probably nothing to do with their 'novel multi-scale noise' (since they haven't compared to the 'non-multiscale' version). Finally, it cites mostly only each field's 'standard' papers, doesn't cite anything really relevant to what it does. Overall, it looks exactly like what you would expect out of GPT-generated paper, just reshashing some standard stuff in a mostly coherent piece of garbage. I really hope people don't start submitting this kind of stuff to conferences.
- blackbear_ 2y agoI would agree with your analysis. Note that it cites TabDDPM in the related work, but that is for diffusione on tabular data! While most tabular data is low-dimensional, the type of low-dimensional data tackled in the paper is not tabular! I'm also not quite sure how the linear upscaling is supposed to help, as it can be absorbed into the first layer of the following MLP, so I would rather think that the performance improvement (if any, the numbers are quite close and lack standard errors) is either due to the increased number of trainable parameters or some kind of ensembling effect (essentially the mixture of experts point made by the human authors).
- iandanforth 2y agoExciting and very cool! I look forward to the continued improvement in this area. Especially when the loop is closed within Sakana and you can say "this discovery was made by The AI Scientist" as part of another paper. If I might offer some small feedback on the blog post: - Alt-text and/or caption of the initial image would be helpful for screen readers - Using both "dramatically" and "radically" in one sentence to describe near future improvements seems a bit much. - When talking about the models used, "Sonnet" could either be 3.0 Sonnet or 3.5 Sonnet and those have pretty different capabilities. Thanks again for the impressive work!
- amoss 2y agoI wonder how model collapse would apply to the AIs created by applying the results of AI-generated papers?
- malux85 2y agoI’m working on this now, I literally have another window open beside this browser window with the Multi-agent LLM logs outputs scrolling. A few differences through - I’m working on Materials Science only. Mine has vision capabilities so it can read graphs in papers. Mine has agentic capabilities too, so can design and then execute simulations on Atomic Tessellator (my startup) by making API calls - this actual design and execution of simulations is what I aimed for at the start. Long way to go, but there’s a set of heuristics that decide which experiments to attempt which means we only attempt ones more likely to work, lots of fine tuning prompts, self critique, modelling strategies and tactics as node graphs to avoid getting stuck in what I call procedural local minima, and loads more… I started with MetaGPT framework but found it’s APIs too unstable so I settled on AutoGen, you don’t really “need” a framework, just be sensible about where your abstraction boundaries are, make them simple but composable, Dockerize and k8s for running, and I modified the binaries of a bunch of quantum chemistry software so that multi GPU arches are supported without re compilation (my hardware setup is heterogeneous) Even if the LLMs can’t innovate in a “new sense” certainly having them reproduce work in simulations for me to inspect is very valuable - I have the ability to “fork” simulations like you can fork code so it’s easy to have the LLMs do a bunch of the work and then I just fork and experiment myself
- agubelu 2y agoI worked with some people who were actively working on this last year, focusing on CS research. The biggest issue was validation. We could get a system to spit out possible research directions automatically, but who decides if they're reasonable and/or promising? A human, of course. Moreover, we gave different humans the same set of hypotheses to validate and they came back with wildly different annotations.
- acureau 2y agoThis was exactly what I was thinking. This may be a useful tool for human researchers to drive if it turns out it can generate anything valuable. I can't understand the papers it wrote, much less determine if they make any contributions. I don't think the self-evaluation is going to be too fruitful. Thanks for sharing
- zipy124 2y agoAs a scientist in academic research, I can only see this as a bad thing. The #1 valued thing in science is trust. At the end of the day (until things change in how we handle research data, code etc...) all papers are based on the reviewers trust in the authors that their data is what they say it is, and the code they submit does what it says it does. Allowing an AI agent to automate code, data or analysis, necessitates that a human must thoroughly check it for errors. As anyone who has ever written code or a paper knows, this takes as long or longer than the initial creation itself, and only takes longer if you were not the one to write it. Perhaps I am naive and missing something. I see the paper writing aspect as quite valuable as a draft system (as an assistive tool), but the code/data/analysis part I am heavily sceptical of. Furthermore this seems like it will merely encourage academic spam, which already wastes valuable time for the volunteer (unpaid) reviewers, editors and chairs time.
- pmayrgundter 2y agoMaybe the #1 valued thing in "capital S Science" -- the institutional bureaucracy of academia -- is trust. Trust that the bureaucracy will be preserved, funded, defended.. so long as the dogma is followed. The politics of Science. The #1 valued thing in science is the method of doing science: reason, insight, objectivity, evidence, reproducibility. If the method can be automated, then great!
- ergl 2y ago> The #1 valued thing in science is [...] reproducibility. If only. Papers rarely describe their methods properly, and reproduction papers have a hard time being published, making it hard to justify the time it takes. If reproducibility was valued, things like retractionwatch wouldn't need to exist.
- pmayrgundter 2y agoWell, agreed! I'd say that's good evidence the political bureaucracy of big-Science has substantially corrupted that (your?) culture. It's not the only way though. There's a bright light coming from open source. Stay close to the people in AI saying "code, weights and methods or it didn't happen" The code that runs most of the net ships with lengthy howto guides that are kept up to date, thorough automated testing in support of changes/experimentation, etc. Experienced programmers who run across a project without this downgrade their valuation accordingly It doesn't solve all problems, but it does show there's a way that's being actively cultivated by a culture that is changing the world
- xpitfire 2y ago[dead]
- ergl 2y agoThis is the kind of theory-free science seems to permeate the entire field of ML lately. I can only see this as a negative, what's the use of automatically generated papers if not to flood the already over-strained volunteers that review papers at conferences? (mostly already-overworked PhD students.) If I wanted a glorified chatbot to spam me with made-up improvements, I'd ask it myself.
- jalman 2y agoAny theory compresses experimental data (a mountain of data) into a palatable bit (pocketable item) of knowledge. One can go without theory just by stacking the original measured ratios dA -> dB. The Solutions, that are generated upon raw data, would likely produce less waste, as they fit the exact phenomena. Generalizations do emphasize the main effects and level out the minor effects, introducing an error. The value of AI-mated research is results. This tool will aid engineering. It will offer a pinpointed research to resolve a particular issue at hand. The research will not be published. It will remain a trade secret. What you are complaining about is a brocken labor division. Let's consider a case. An engineering department has a problem. It passes the problem to the research department. The research department slacks out, while produces junk. It makes tons of papers that are hard to prove, not to say apply. They drink champagne with other researchers, so they can publish the junk and defend the turf. AI-mated research will finish the racket and corruption. A "$15-scientist" will kill the no-output individuals, who are just a sophisticated flavor of bureaucracy. A science bureaucrat is a bureaucrat with elevated rights and virtually no responsibility.
- vouaobrasil 2y agoAs someone who truly loves science, the idea of automating the creative parts strikes me at the core as a horrible mistake. Yes, even before AI, we've already tried some automations -- actually some of those I even believe is a bad thing, such as the internet. Most people would disagree no doubt, but I feel like automating science, especially with regard to the more "creative parts" makes it more like an industry, ripping it away from the minds of people. And AI automation is a new level that goes beyond all automations. I truly hate AI and what it is doing to the world and to me at least, as someone who has loved mathematics and science since my grandma started showing me chemistry experiments when I was about 5 years old, this new level of automation is stealing the magic from human curiosity.
- whamlastxmas 2y agoI think this is looking at it wrong. If AI can do boring science it frees us up to do imaginative and fun science without the constraints of capitalism. You don’t have to worry about your science being valuable enough
- vouaobrasil 2y ago> AI can do boring science it frees us up to do imaginative and fun science without the constraints of capitalism. That is senseless. Capitalism will always control science by its very nature: through science, people create value and trade it for other things.
- 93po 2y agoIf AI has gotten to the point of doing nearly all science for humanity, we are probably close to, if not already at, a post-capitalist state. The implication being there's no real resource scarcity at that point.
- adroniser 2y agoI completely agree this shit is so depressing. When I saw the AlphaProof paper I basically spent 3 days in mourning basically, because their approach was so simple.
- JBorrow 2y agoAs someone 'in academia', I worry that tools like this fundamentally discard significant fractions of both the scientific process and why the process is structured that way. The reason that we do research is not simply so that we can produce papers and hence amass knowledge in an abstract sense. A huge part of the academic world is training and building up hands-on institutional knowledge within the population so that we can expand the discovery space. If I went back to cavemen and handed them a copy of _University Physics_, they wouldn't know what to do with it. Hell, if I went back to Isaac Newton, he would struggle. Never mind your average physicist in the 1600s! Both the community as a whole, and the people within it, don't learn by simply reading papers. We learn by building things, running our own experiments, figuring out how other context fits in, and discussing with colleagues. This is why it takes ~1/8th of a lifetime to go from the 'world standard' of knowledge (~high school education) to being a PhD. I suppose the claim here is that, well, we can just replace all of those humans with AI (or 'augment' them), but there are two problems: a) the current suite of models is nowhere near sophisticated enough to do that, and their architecture makes extracting novel ideas either very difficult or impossible, depending on who you ask, and; b) every use-case of 'AI' in science that I have seen also removes that hands-on training and experience (e.g. Copilot, in my experience, leads to lower levels of understanding. If I can just tab-complete my N-body code, did I really gain the knowledge of building it?) This is all without mentioning the fact that the papers that the model seems to have generated are garbage. As an editor of a journal, I would likely desk-reject them. As a reviewer, I would reject them. They contain very limited novel knowledge and, as expected, extremely limited citation to associated works. This project is cool on its face, but I must be missing something here as I don't really see the point in it.
- swayvil 2y agoWe already substitute "good authority" (be it consensus or a talking head) for "empirical grounding" all the time. Faith in AI scientific overlords seems a trivial step from there.
- tossandthrow 2y agoWhy shouldn't we?
- TeeWEE 2y agoAI hype in one sentence: "We expect all of these will improve, likely dramatically, in future versions with the inclusion of multi-modal models and as the underlying foundation models" So much hype, so much believe. I no believe no hype
- crabbone 2y agoI'm not a scientist at all, but I am often involved in hand-holding scientists when it comes to dealing with computers. My impression so far is that science is plagued with deliberate and accidental fraud when it comes to data collection and cataloguing. Also, this is a spectrum, not two distinct things. I often see researchers simply unwilling to do the right thing to verify that the data collected are correct and meaningful as soon as "workable" results can be produced from the data. Some will go further and mess with the data to make results more "workable" though... Second problem is understanding the data. Often times it happens that people who end up doing research don't quite understand the subject matter of the research. This is especially popular with medicine, where it's overwhelmingly common for eg. research into various imaging modalities to be done by computer scientists who couldn't find a liver cancer the size of a coconut in the sharpest textbook abdominal image. My impression is also that by far these two problems outweigh the problems that could potentially be solved by adding AI into the mix. These are the systemic organization problems of perverse incentives and vicious practices, and no amount of AI is going to do anything about it... because it's about people. People's salaries, careers, friendships etc.
- jalman 2y agoYou've brought a good point. The exact difference between ML and AI is in the focus. ML is focusing on tossing data. Its main success metric is "miles traveled". AI is focusing on producing sense. Its main success metric is "value delivered". Therefore ML is tending to ship empty containers, aka "research to be done by scientists who couldn't find a liver cancer". Adding AI perspective into the mix could actually turn the things around.
- breck 2y agothe aientist.com ;)
- _heimdall 2y ago> The AI Scientist is designed to be compute efficient. Each idea is implemented and developed into a full paper at a cost of approximately $15 per paper. While there are still occasional flaws in the papers produced by this first version (discussed below and in the report), this cost and the promise the system shows so far illustrate the potential of The AI Scientist to democratize research and significantly accelerate scientific progress. This is a particularly confusing argument in my opinion. Is the underlying assumption that everyone wants, or even needs, white papers that they can claim they created? Let's just assume this system actually works and produces high quality, rigorous research findings. Reducing that process down to a dollar amount and driving that cost to near zero doesn't democratize anything, it cheapens it to the point of being worthless. This honestly reads more as a joke article trolling today's academic process and the whole publish or perish mentality. From that angle, the article is a success in my book. As an announcement for a new ML tool though, I just don't get it.
- sebastiennight 2y agoTFA aside, if we could make scientific research produce new insights with low latency and almost-zero costs, it would definitely not make science worthless. It would be a fantastic day for science. Not everything is worth (production costs + margin). Many things have intrinsic worth and are worth more to society if you drive down their cost of production.
- _heimdall 2y agoThat sounds like an extremely dangerous day for science as well. If anyone could pop up an ML tool and task it with inventing and validating something truly novel, that would be weaponized extremely fast (likely right after people use it for porn, the frontier for all new tech). I do totally agree on the cost + margins point you make. I've never actually been a fan of valuing things in that way, and in my pipe dream utopia we wouldn't need or use money at all. I didn't clarify enough there, but I actually mean worthless in the non-monetary sense as well. Any invention created through such an ML tool would be one of a countless pile of stuff created. How important can any one really be?
- dotnet00 2y agoSo, if AI trained on AI generated data tend to perform worse, and you're trying to have AI generate the data that ultimately props up our civilization... Clearly this can only end well
- sebastiennight 2y ago> AI trained on AI generated data tend to perform worse Citation needed. To the best of my knowledge, synthetic data is a solid way to train models and isn't going away anytime soon.
- dotnet00 2y agohttps://www.nature.com/articles/s41586-024-07566-y https://www.nature.com/articles/s41586-024-07566-y IIRC it had a pretty big thread here a few weeks ago
- sebastiennight 2y agoThanks. I remember that thread: https://news.ycombinator.com/item?id=41058194 https://news.ycombinator.com/item?id=41058194 What I took from the discussion is that there is very little chance that our next step for training SOTA models (eg. LLMs) will be "scraping the whole web including increasing volumes of ChatGPT-generated content". Instead, the synthetic content used to train new models is (from recent papers I've seen) mostly curated - not "indiscriminate" as the Nature paper discusses.
- sebastiennight 2y agoHaving read the article, it seems like an interesting experiment. With the current state of LLMs, this is extremely unlikely to produce useful research, but most of the limitations people have been commenting about are will progressively get better. The authors' credibility is a bit hurt when the first "limitation" they mention is "our system doesn't do page layouts perfectly". Come on, guys. What is weird to me is this: > The AI Scientist occasionally makes critical errors when writing and evaluating results. For example, it struggles to compare the magnitude of two numbers, which is a known pathology with LLMs. To partially address this, we make sure all experimental results are reproducible, storing all files that are executed. I'm not sure why you would run your evaluation step without giving your LLM access to function calling. It seems within reach to first have the LLM output a set of statements-to-be-verified (eg, "does X increase when Y increases?") and then use their code-generation/execution step to perform those comparisons. And then the incomprehensible statement for me here is that they allow the model access to its own runtime environment so it can edit its own code? The paper is 185 pages and only has one paragraph on safety. This screams "viral marketing piece" rather than "serious research". And finally: > The AI Scientist can produce papers that exceed the acceptance threshold at a top machine learning conference... Oh wow, please tell more? > ... as judged by our automated reviewer. Ah. Nevermind
- woah 2y agoEveryone in this thread is musing about the role of AI and whether the process of discovery is fundamentally human, and what Isaac Newton would think, but can somebody tell me: is the technology it develops any good? For example, does "Dual Scale Diffusion" https://sakana.ai/assets/ai-scientist/adaptive_dual_scale_denoising.pdf https://sakana.ai/assets/ai-scientist/adaptive_dual_scale_de... look useful?
- tzumaoli 2y agoDiscussed in another thread https://news.ycombinator.com/item?id=41234415 https://news.ycombinator.com/item?id=41234415 As someone who has worked on diffusion model, it's a clear reject and not a very interesting architecture. The idea is to train a diffusion model to fit to low dimensional data using two MLPs: one accounts for high-level structure and one accounts for low level details. These kind of "global-local" architecture is very common in computer vision/graphics (with the paper mentioned none of the relevant work), so the novelty is low. The experiments also do not clearly showcase where exactly this "dual" structure brings benefits. That being said, it's very hard to tell it apart from a normal poorly-written paper from a quick glance. If you tell me it's written by a graduate student, I would probably believe it. It is also interesting in a way that maybe for low-dimensional signals there are some architecture tweaks we can do to modify the existing diffusion model architectures to make things better, so maybe not 100% BS.
- happypumpkin 2y agoPotential concerns with their self-eval: They evaluate their automated reviewer by comparing against human evaluations on human-written research papers, and then seem to extrapolate that their automated reviewer would align with human reviewers on AI-written research papers. It seems like there are a few major pitfalls with this. First, if their systems aren't multimodal, and their figures are lower-quality than human-created figures (which they explicitly list as a limitation), the automated reviewer would be biased in favor of AI-generated papers (only having access to the text). This is an obvious one but I think there could easily be other aspects of papers where the AI and human reviewers align on human-written papers, but not on AI papers. Additionally, they note: > Furthermore, the False Negative Rate (FNR) is much lower than the human baseline (0.39 vs. 0.52). Hence, the LLM-based review agent rejects fewer high-quality papers. The False Positive Rate (FNR [sic]), on the other hand, is higher (0.31 vs. 0.17) It seems like false positive rate is the more important metric here. If a paper is truly high-quality, it is likely to have success w/ a rebuttal, or in getting acceptance at another conference. On the other hand, if this system leads to more low-quality submissions or acceptances via a high FPR, we're going to have more AI slop and increased load on human reviewers. I admit I didn't thoroughly read all 185 pages, maybe these concerns are misplaced.
- happypumpkin 2y agoAlso a concern about the paper generation process itself: > In a similar vein to idea generation, The AI Scientist is allowed 20 rounds to poll the Semantic Scholar API looking for the most relevant sources to compare and contrast the near-completed paper against for the related work section. This process also allows The AI Scientist to select any papers it would like to discuss and additionally fill in any citations that are missing from other sections of the paper. So... they don't look for related work until the paper is "near-completed." Seems a bit backwards to me.
- jalman 2y agogreat point. I think the AI scientist is already a winner. If the likelihood of false outcome is FNR+FPR, then machine would fail 0.7 and humans 0.69 times. Humans do win nominally. In terms of costs humans loose. For every FPR 0.31-0.17 = 0.14 you spend additionally, you'd gain FNR 0.52-0.39 = 0.13. The paper production costs discrepancy is at least factor 100. The value of the least useful research typically drives factor two or more benefit in comparison to production and validation costs. So the final balance is 0.014 to 0.36 -> x25 gain in favor of AI.
- plaidfuji 2y agoWhen “executing the experiment” amounts to modifying ~50 lines of PyTorch code tweaking model architecture, I’d bloody well expect that you can automate it. That’s not “automating scientific discovery”, that’s “procedurally optimizing model architecture” (and one iteration of exploration at that!). In any other field of science the actual work and data generated by the AI Scientist would be a sub-section of the Supporting Info if not just a weekly update to your advisor. Don’t get me wrong, the actual work done by the humans who are publishing this is a pretty solid piece of engineering and interesting to discuss. But the automated papers, to me, are more a commentary on what constitutes a publishable advancement in AI these days. Edit: this also further confirms my suspicion about LLMs, which is that they aren’t very good at doing actual work, but they are great at generating the accompanying marketing BS around having done work. They will generate a mountain of flashy but frivolous communication about smaller and smaller chunks of true progress, which while appearing beneficial to individuals, will ultimately result in a race to the bottom of true productivity.
- adroniser 2y agoI think the whole paper is a satire lol.
- anticensor 2y agoPresentation of Intermediate Results. The paper contains results for every single experiment that was run. While this is useful and insightful for us to see the evolution of the idea during execution, it is unusual for standard papers to present intermediate results like this. This is actually quite good that the AI scientist does this. AI has no excuse of slow report writing that humans have to omit the intermediate results.
- pemrick79 2y agoStfu and listen to the music and enjoy it the way you want. Who gives a shit what some tech nerd or stuffy audiophile has to say? Yup only them. You guys laugh at each other think the other is ridiculous. Well we all laugh at both of you. Man what a waste of the last 15 min of my life. Back to enjoying my music anywhere and everywhere I go. Should try it
- ai4ever 2y ago[dead]
- mnky9800n 2y agoI guess if C3P0 wants to help me do science I'll let him. But I'm not really satisfied by someone telling me something. I like finding the answer myself. That's one of the reasons I'm a scientist. I enjoy being a scientist. Even assuming that this system, or any science bot, has none of the problems associated with llms, why would I want it to do my job for me?
- fragmede 2y agoBeing a scientist is cool and all, but you have to do things the exact same way except for changing one variable at a time for a proper science experiment. Having a robot helper to run the experiment the exact same way modulo the variable would be great for science and especially for reproducing experiments. There's a reproduction crisis and a robot that could run reproductions would move science itself forwards. The other part of the question is, for you, merely just being a scientist is the most rewarding thing ever and there's absolutely no boring parts whatsoever? Your favorite part of the job is everything, and there's none of it that you'd trade for a bit of the other? I mean, I suppose that's possible, but that beggars belief.
- mnky9800n 2y agoI'm a computational scientist so I already spend most of my time writing codes to automate my science. If C3P0 wants to help out then I'm happy to have him. The reproduction crisis is predominately in social sciences. This is for several reasons but the main one is that social science is much newer than the physical sciences. There's a great book about rasch modelling which equates social science to physics before thermodynamics. That at the very best, you can measure the temperature but there was a time when we had a poor understanding of what it meant when the temperature changes. That's clearly where social science is right now and the reproduction crisis is a good thing. It's challenging the status quo to come up with better theory to motivate observation and experiment. To answer your final question, yes, the parts of my job that I find annoying, unrewarding, and time wasting that I would happily trade to another are itemized but not complete below: * Planning my own travel, purchasing all my own tickets and lodging and having no help to do any of that * Sitting in meetings with colleagues who are explaining to me why it's ok that the interns aren't getting paid because they are going to give them visa gift cards instead * Editorialmanager.com website * Having no itemized or explicit way to examine the expenditure in my grants so that when I ask admin staff how much money there is for X they reply with "you have enough money" instead of the actual amount * Researchgate website * Colleagues who are rude and condescending instead of going to therapy to wonder why they are insecure * Expense reimbursement instead of having a purchase card attached to my grants * Convincing older colleagues to use "new" technology like slack or GitHub If C3P0 can solve any one of those problems for me I'll get one in a heartbeat.
- Arrgh 2y agoThe specific mechanism of action here needs a drill-down. Did "The AI Scientist" (ugh) generate a patch to its code and prompt a user to apply it, as the screenshots would seem to indicate? If so I don't find this worrying at all: people write all kinds of stupid code all the time--often, impressively, without any help from "AI"! ;) Did it apply the patch itself, then reset the session to whatever extent necessary to have the new code take effect? TBH I'm not really worried about this either, as long as the execution environment doesn't grant it the ability to bring unlimited additional hardware to bear. Even in that case, presumably some human would be paying attention to the AWS bills, or the moral equivalent.
- PROMISE_237 2y ago[dead]