5 ms·
From the paper's method section, a bit more about which type of ML algo was used: An RF machine-learning model was developed to predict lithium concentrations
by folli 2y ago
From the paper's method section, a bit more about which type of ML algo was used:
An RF machine-learning model was developed to predict lithium concentrations in Smackover Formation brines throughout southern Arkansas. The model was developed by (i) assigning explanatory variables to brine samples collected at wells, (ii) tuning the RF model to make predictions at wells and assess model performance, (iii) mapping spatially continuous predictions of lithium concentrations across the Reynolds oolite unit of the Smackover Formation in southern Arkansas, and (iv) inspecting the model for explanatory variable importance and influence. Initial model tuning used the tidymodels framework (52) in R (53) to test XGBoost, K-nearest neighbors, and RF algorithms; RF models consistently had higher accuracy and lower bias, so they were used to train the final model and predict lithium.
Explanatory variables used to tune the RF model included geologic, geochemical, and temperature information for Jurassic and Cretaceous units. The geologic framework of the model domain is expected to influence brine chemistry both spatially and with depth. Explanatory variables used to train the RF model must be mapped across the model domain to create spatially continuous predictions of lithium. Thus, spatially continuous subsurface geologic information is key, although these digital resources are often difficult to acquire.
Interesting to me that RF performed better the XGBoost, would have expected at least a similar outcome if tuned correctly.
- jandrese 2y agoDid they actually verify the predictions? In my reading of the article I didn't see any core samples being made to verify the model is correct.
- jofer 2y agoThere wouldn't be any core for this. It would be a holdout of the brine samples used in training. The thing that would be being produced is brine, so lithium concentrations in brine samples are the validation dataset as well. In other words, this is spatial interpolation.
- tomrod 2y agoRF is a heavy hitter when it comes to tabular data. XGBoost is good as well, but more often than not needs and autotuner to really unlock it (e.g pycaret).
- jncfhnb 2y agoXGBoost models are random forest models. They’re also just consistently better for very little effort.
- prog_1 2y agoyou surely mean that both are ensemble models. RFs and GBMs differ in how they fit the data
- jncfhnb 2y agoA GBM like XGBoost is an ensemble of trees. It may be that when you load RandomForest modules they fit based on entropy or whatever the typical DecisionTree does but imo the term “random forest” should really convey nothing more than “ensemble of trees”. I’m saying XGBoost would be a subclass of RF
- tomrod 2y agoNot only RF, they incorporate GBM too as I understand it. Often they are the best "just run it and forget it" but compared to tuning they don't always achieve top -- sometimes surprisingly so. XGBoost and similar are solid first stops in model building.
- lordgrenville 2y agoSo it turns out that there's no theoretical reason that gradient boosting will always outperform RF (which would violate the "no free lunch" theorem). But it does usually seem to be the case in practice, even with small and noisy data. I would hazard a guess that with better tuning, XGBoost would still have won. (The paper notes that the authors chose a suboptimal set of hyperparameters out of fear of overfitting - maybe the same logic justifies choosing a suboptimal model type...)
- levocardia 2y agoThat's been my experience. RF tends to do quite well out of the box, and is very fast to fit. It's less of a pain to cross-validate too, with fewer tuning parameters. XGBoost has a huge number of knobs to tune, and its performance varies from god-awful with bad hyperparameters to somewhat better than RF with good ones. Giant PITA with nested cross-validation, etc. though. I haven't read in detail what their validation strategy is but this seems like the kind of problem where it's not so easy as you'd think -- you need to be very careful about how you stratify your train, dev, and test sets. A random 80/10/10 split would be way too optimistic: your model would just learn to interpolate between geographically proximate locations. You'd probably need to cross-validate across different geographic areas. This also seems like an application that would benefit from "active learning". given that drilling and testing is expensive, you'd want to choose where to collect new data based on where it would best update your model's accuracty. A similar-ish ML story comes from Flint, MI [1] though the ending is not so happy [1] https://www.theatlantic.com/technology/archive/2019/01/how-machine-learning-found-flints-lead-pipes/578692/ https://www.theatlantic.com/technology/archive/2019/01/how-m...
- dwattttt 2y ago> your model would just learn to interpolate between geographically proximate locations At a particular scale, this is entirely correct; if what I'm looking for is 'large', a measurement 1m away from a known hit would also be likely to be a hit. That particular issue sounds like it should be addressed with more negative samples.
- youoy 2y ago
- jofer 2y agoPut another way, this is pretty similar to the interpolation approaches that would normally be used for datasets like this in the world of mineral exploration. Kriging/co-kriging (i.e. gaussian processes) is the more commonly used approach in this particular field due to both the long history and the available hyperparameters for things like spatial aniostropy. However, kriging is really quite difficult to use with non-continuous inputs. RF is a lot more forgiving there. You don't need to develop a covariance model for discrete values (or a covariance model for how the different inputs relate, either).
- aaronblohowiak 2y agofor other folks wonder what the acronym means; RF in this context is Random Forest
- f_devd 2y agoFor a moment I was excited that they had done surveys entirely on RF backscattering and ML.
- Loic 2y agoRF is random forest[0]. We had this discussion a couple of days ago: "Why do Random Forests Work? Understanding Tree Ensembles as Self-Regularizing Adaptive Smoothers". https://arxiv.org/abs/2402.01502 https://arxiv.org/abs/2402.01502 https://news.ycombinator.com/item?id=41873968 https://news.ycombinator.com/item?id=41873968 [0]: https://en.wikipedia.org/wiki/Random_forest https://en.wikipedia.org/wiki/Random_forest