5 ms·
Scaling PostgresML to 1M Requests per Second
- tomrod 4y agoI haven't heard of PostgresML before today. How does it compare to Feast?
- pridkett 4y agoFeast, for the most part, is a feature store that serves up features to machine learning models for inference or training. PostgresML is closer to Big Query ML. They provide the framework to building, deployment, and inference of models from within the database - using extensions for PostgreSQL. This allows analysts and other users to build models off primarily tabular data without needing to figure out the whole deployment mess.
- alberth 4y ago4 load balancers with local caches + 5 replica databases. I don’t want to discredit the achievement but the 1MM/sec seems less exciting when you learn they horizontally scaled the architecture. Especially when SQLite, in a client/server config (BedrockDB), achieved 4MM/sec from a single server. https://blog.expensify.com/2018/01/08/scaling-sqlite-to-4m-qps-on-a-single-server/ https://blog.expensify.com/2018/01/08/scaling-sqlite-to-4m-q...
- sidpatil 4y agoI'm not sure the two are comparable—OP's article is about 1 million ML predictions per second, while your link is about 4 million queries per second.
- alberth 4y agoThanks for pointing out the difference. Genuine question: is a ML prediction slower than to fetch a row from a database and/or more difficult to scale?
- mrfox321 4y agoIt's slower than a row lookup. Specifically, this model is an ensemble of decision trees. This involves 1. Row lookup 2. N (tree depth) inequality checks on fields in the row for M trees 3. Weighted sum over M trees
- tucnak 4y agoIn this case, they are talking about xgboost inference, which can verge on the hungry side of things, but this is no match for deep learning, although word embeddings and transformers are supported[1] by PostgresML regardless. GPUs are supported[2] for Transformers as well as xgboost, so this approach is also kind of future-proof in the sense that you're not going to ever be limited by your compute, as well as the size of your replica set. But now, yes, to answer your original question— yes, for the most part typical regression-type task inference is marginally slower than fetching and involves crossing the process boundary, can't be easily parallelised, and so on. Hopefully, this is something that could be dealt with on PostgresML side, as Postgres exposes some neat parallelism primitives AFAIK. COMBINEFUNC is how you make your aggregates parallel and with SERIALFUNC potentially multinode-friendly, although here I'm not sure as I haven't maintained custom aggregates over a high-availability Postgres cluster... At any rate, this is a clear area of future research; now that there's inference, how to best parallelise it in the model of Postgres. [1] https://postgresml.org/user_guides/transformers/setup/ https://postgresml.org/user_guides/transformers/setup/ [2] https://postgresml.org/user_guides/setup/gpu_support/ https://postgresml.org/user_guides/setup/gpu_support/
- ComputerGuru 4y agoCan you train xgboost on a GPU and use that model to make predications on the CPU?
- tucnak 4y agoSure!
- simonw 4y agoFrom the article: > XGBoost cannot serve predictions concurrently because of internal data structure locks. This is common to many other machine learning algorithms as well, because making predictions can temporarily modify internal components of the model.
- gmhafiz 4y agoFor perspective, that single server is - 1TB of DDR4 RAM - 3TB of NVME SSD storage - 192 physical 2.7GHz cores (384 with hyperthreading
- fulafel 4y agoAnd AWS vCPUs = SMT threads, or Hyperthreads(tm) in Intel speak for anyone comparing to the machines described in the PostgresML post. (Which were answering ML ops, not simple DB queries)
- Foobar8568 4y agoSql Server + R 6 years ago, scoring transactions at 1M/sec https://blog.revolutionanalytics.com/2016/09/fraud-detection.html https://blog.revolutionanalytics.com/2016/09/fraud-detection...
- deleted 4y ago[deleted]
- bagels 4y agoRight, if we're allowed horizontal scaling, 1M instances, 1request/s would do the trick.
- tucnak 4y agoWe have used PostgresML to a great success albeit in a very limited setting. The whole philosophy of pushing Postgres to the absolute limit of what is possible is very attractive; think what Timescale[1] does for time series, and what PostgresML does for machine learning by eliminating the need to perform any kind of ETL jobs to cut down on the latencies as to how late the data is consumed after it's initially produced, et cetera. The very fact that we can now train and deploy regressions/classifiers using nothing but a couple views and timely SQL calls is already a fairly attractive proposition, however I'm really looking towards further adoption, library support, case studies and stuff like that. For example, I've been considering how Timescale could be used together with PostgresML to provide a single way to treat data from the point of consumption. In fact, I've reached out to Timescale Cloud on this very matter, and was sad to learn that they can't support PostgresML as part of their Cloud offering due to some security considerations related to it being a Python extension. Self-hosted it is. At any rate, they have some time back introduced what they call continuous aggregates[2]; a materialised view on top of a hypertable that doesn't require explicit refreshing, and is in fact realtime-accurate due to some waterline logic; it performs a normal view-like query for the most recent bits while retaining the materialised component in the compressed partial form. (Normal time-based compression and retention policies apply like they would to any other hypertable which is how the continuous aggregate views are implemented under the hood.) The idea here is that you can potentially reduce hundreds of million data points to a set of particular continuous aggregates within respective time frames; downsampling of past data is something you get for free. This is where I think the value lies for PostgresML; using a continuous aggregates for continuous training and as a store of historical predictions that you would normally want to keep separate from the training set itself. This could prove a reliable way of feeding new data for re-training along with downsampled selection of past data, and comparing various predictions on similar data points over time to keep track of how it goes. The data is compressed away while readily available and as long as it doesn’t have to change (historical predictions apply) the data points in the materialised component are never going to be materialised more than once. There are limitations[3] to this approach, of course, however they can be circumvented. If you're interested in this, you should check out some of the “hyperfunctions” they have to offer such as histogram[4] which is what I’ve had the pleasure of using previously to distil some of the time series data into a fixed-size feature vector, and naturally this would be a good place to start if you’re ever going to seriously consider this as something that you yourself would want to use for similar purpose. To me this is a natural step forwards in the data lifecycle/ supply chain approach. First, you get rid of ETLs as such by adopting PostgresML, and next you specify the “contract" of how your data is going to be produced, reduced, distilled, modelled, sampled, and ultimately evaluated— over prolonged periods of time and multiple iterations of the implementation. [1] https://docs.timescale.com/timescaledb/latest/overview/core-concepts/ https://docs.timescale.com/timescaledb/latest/overview/core-... [2] https://docs.timescale.com/timescaledb/latest/how-to-guides/continuous-aggregates/about-continuous-aggregates/ https://docs.timescale.com/timescaledb/latest/how-to-guides/... [3] https://docs.timescale.com/timescaledb/latest/overview/limitations/ https://docs.timescale.com/timescaledb/latest/overview/limit... [4] https://docs.timescale.com/api/latest/hyperfunctions/histogram/ https://docs.timescale.com/api/latest/hyperfunctions/histogr...
- ehayes 4y agoWhat is a good algorithm-to-purpose map for ML beginners? Looking for something like "Algo X is good for making predictions when your data looks like Y," etc.
- remram 4y agoThis maybe? https://scikit-learn.org/stable/tutorial/machine_learning_map/index.html https://scikit-learn.org/stable/tutorial/machine_learning_ma...
- bagels 4y agoThis is what I would have replied with too.
- sandGorgon 4y agoxgboost. always xgboost. it will scale all the way from college kaggle problems (where it is the top performer almost always) to cloud scale. xgboost is one of the few frameworks supported by Sagemaker, etc
- xdfgh1112 4y agoVowpal Wabbit is not the best anymore, but it is incredibly simple. You train it by piping text files in, then pipe your input into it for predictions.
- FreakLegion 4y agoTsk to whoever downvoted this. Simple linear models are indeed the right starting point for most new projects while you come to grips with your data. In some cases you can stop there or apply a quick nonlinearization like Fastfood to get good, snappy, and generally debuggable results for very little RAM. In other cases you move on to decision tree ensembles or neural networks, depending on whether you already have features or need those to be learned, too. Either way this ratchets up the complexity and resource requirements. Decision trees in particular tend to have bloated implementations. I still use XGBoost or Scikit for training, but wrote my own library to translate the models into a more efficient format (~95% smaller than Scikit) and have thread-safe inference.
- deleted 4y ago[deleted]