4 ms·
The current models were trained on a corpus that is essentially all fiction with no basis in reality. If the training corpus has real world data (like experimen
by experimental123 3y ago
The current models were trained on a corpus that is essentially all fiction with no basis in reality. If the training corpus has real world data (like experimental results from agricultural experiments with crops and planting schedules along with their yields) then the neural network should uncover some patterns that wouldn't be obvious simply because finding correlations in large data sets is a hard problem but it is very well suited to analysis by large neural networks.
- ceejayoz 3y agoI mean, you can go try this now; feed some agricultural scientific journals into a model. I suspect it's going to be substantially harder than you expect.
- experimental123 3y agoI agree it is a very easy to do which is why it's surprising someone hasn't already tried it. Most of what I see are toy projects with LoRA for generative models bolted onto existing LLMs for fiction instead of scientific applications. These models already work for software so I see no obvious obstructions why they shouldn't work for agricultural experiments.
- gamblor956 3y agoIt's not very easy to do. LLMs aren't capable of understanding, they can merely regurgitate what they've read based on statistical analysis of what words appear to be linked to each other. That doesn't help you when you need to do something new; at best an LLM can tell you what someone else has already done. There are computer programs that do the kind of thing you're thinking about, for example, for protein structure analysis. They're incredibly complicated and generally require a lot of processing power.
- PaulHoule 3y agoHow about https://www.frontiersin.org/articles/10.3389/fpls.2023.1128388/full https://www.frontiersin.org/articles/10.3389/fpls.2023.11283... ? That's a simple application of machine learning algorithms you might find in scikit-learn. Here is a special issue of another alleged "predatory journal" that is full of papers on the subject https://www.mdpi.com/journal/agronomy/special_issues/E18K759IAF https://www.mdpi.com/journal/agronomy/special_issues/E18K759...
- gamblor956 3y agoLLMs are a type of machine learning. The stupidest type of machine learning. The OP did not suggest machine learning in generally, they suggested LLMs specifically, which time and time again have been shown to be incapable of this task as a matter of fundamental design. Worse, because LLMs can't understand the their training data, the output of an LLM must be verified, which in situations like this would probably take more time than simply conducting novel research in the form of random experimentation. Also, you really need to read your citations. The first one found that machine learning was unsuited for the task of agricultural prediction...
- thfuran 3y ago>These models already work for software Do they? You're talking the agricultural equivalent of something like "devise a new sorting algorithm with sota performance on x, y, z", not "write me some crud boilerplate".
- agronomicon 3y agoI just mean finding correlations in data sets that are hard to find in other ways. The main idea is that there are plenty of data sets on various cultivars and experiments for how to increase yields. There are probably patterns in the data that would be amenable to analysis by neural networks. The article gives an example for how scheduled flooding can increase yields and I bet there are a lot of low hanging fruits like that to pick. This doesn't require discovering anything novel but simply surfacing some patterns in the data that is buried across several papers and hard to uncover by classical meta-analysis and statistical techniques. Neural networks are very good for uncovering non-obvious statistical correlations which can then be verified by experimentation. After reading the article I'm sure there are plenty of low hanging fruits to uncover in yield optimization by trying different schedules for flooding and soil enrichment with different kinds of fertilizers. A neural network doesn't have to understand anything to point out useful statistical correlations just like it doesn't have to understand code semantics for incomplete code fragments to suggest potential completions which are then verified by the programmer/compiler/type system.
- Karrot_Kream 3y agoI would speculate, but don't concretely know, that this is what will happen. I know papers in other fields that were just this; analyzing conditions that successful and failed experiments were performed in and then using ML to derive optimal conditions.
- PaulHoule 3y agoPeople do all kinds of meta-analysis and literature reviews today, I am sure somebody is already applying A.I. to the document handling for this task but doing a quick search it is hard to differentiate it from literature reviews on the subject of A.I. in agronomy such as https://www.frontiersin.org/articles/10.3389/fsufs.2022.1053921/full https://www.frontiersin.org/articles/10.3389/fsufs.2022.1053... It's a big problem that ChatGPT has seduced a large number of people into thinking chatbots = AI and those people have convinced most other people that it is a scam. I find 77,000 or so articles on "rice" in PubAg https://search.nal.usda.gov/discovery/search?query=any,contains,rice&tab=pubag&search_scope=pubag&vid=01NAL_INST:MAIN&offset=0 https://search.nal.usda.gov/discovery/search?query=any,conta... Just like many other areas, agriculture responds to knowledge and is a highly competitive international business. For instance, rice is cultivated by very different methods in Louisiana and Bangladesh and rice from either place could make it to your table. See https://en.wikipedia.org/wiki/System_of_Rice_Intensification https://en.wikipedia.org/wiki/System_of_Rice_Intensification for a method which is heavy on labor input and light on fossil fuel input.
- agronomicon 3y ago> I find 77,000 or so articles on "rice" in PubAg Analyzing this data set with an LLM would be a very good research project.
- PaulHoule 3y agoExactly, and not that hard. My RSS reader has ingested about 250,000 articles from random sources since the beginning of this year and does a cluster analysis of about 50,000 of them every day in under two minutes.
- semi-extrinsic 3y agoMost software dev is repetitive as hell, monkey see monkey do within a computer readable language that has well defined syntax. LLMs can do fairly well in this niche. Research is by definition not repetitive, the text is free form and the data is never formatted in a way that makes comparison between different papers straight forward.
- agronomicon 3y agoThat's exactly the type of data set that can be analyzed by large neural networks. Heterogeneous data with hidden and non-obvious statistical correlations which would be hard to uncover with classical statistical tools and techniques.
- throwbadubadu 3y agoNot at all convinced that this is true, the contrary. Do you have a reference, or something similar that did this successful in another field? (No, that's not ChatGPT and writing some limited software).
- galactician 3y agoFacebook's Galactica.
- gamblor956 3y agoWhich failed miserably at this task... https://www.technologyreview.com/2022/11/18/1063487/meta-large-language-model-ai-only-survived-three-days-gpt-3-science/ https://www.technologyreview.com/2022/11/18/1063487/meta-lar... "A fundamental problem with Galactica is that it is not able to distinguish truth from falsehood, a basic requirement for a language model designed to generate scientific text. People found that it made up fake papers (sometimes attributing them to real authors), and generated wiki articles about the history of bears in space "
- tomrod 3y agoI have my doubts an AI will reliably generate results that are scientifically verifiable. AI interpolates across it's parameter space, but typically performs poorly in extrapolation exercises.
- vkou 3y agoThere's no shortage of ideas for improving real-world processes. Most of those ideas are bunk, and we're constrained by the amount of experiments[1] we are willing to run/fund, and the quality of data[1] that those experiments can collect, and the reproducibility[1] of those experiments. Having an AI shout random ideas is very easy for software people to grok, but isn't going to help. If you want AI to assist with this, you'd need to build an 'AI' that can run the real-world experiments, and that's a few orders of magnitude harder than feeding a text corpus to an LLM. 'Thinking' about this problem isn't the hard part, the hard part is doing it. Even using an LLM for something like a meta-analysis of existing research is unlikely to find many profitable avenues of exploration. [1] Experimental research is incredibly difficult, which is a fact that's highly underappreciated by people working in abstract and theoretical disciplines.
- grokist 3y ago[dead]