5 ms·
> essentially same pandas just for R You are aware the pandas was designed to replicate the behavior of base R's dataframes? I've been a heavy user of both an
by IKantRead 3y ago
> essentially same pandas just for R
You are aware the pandas was designed to replicate the behavior of base R's dataframes?
I've been a heavy user of both and R's data frames are still superior to pandas even without the tidyverse.
Pandas is really nice for the use case it was designed for: working with financial data. This is a big part of why Pandas's indices feel so weird for everything else, but if your index is a time in a financial time series then all of a sudden Pandas makes sense and works great
When not working with financial data I try to limit the amount of time my code touches pandas, and increasingly find numpy + regular python works better and is easier to build out larger software with. It also makes it much easier to port your code into another language for use in production (i.e. it's quick and easy to map standard python to language X, but not so much a large amount of non-trivial pandas).
- slt2021 3y agowith pandas2.0 and using arrow backend instead of numpy - pandas became "cloud datalake native" - you can essentially read from arrow files in S3 very efficiently and at any large scale - and store/process arbitrarily large amounts of files in a cheap serverless infra. Arrow format is also supported by other languages. with s3+sqs+lambda+pandas - and you can build cheap serverless data processing pipelines and iterate extremely quickly
- Karrot_Kream 3y agoDo you have any benchmarks about how much data a given lambda can search/process after loading Arrow data? Not trying to argue, I'm curious because I never thought of this architecture myself, because I would think that the time it takes to ingest the Arrow data and then search through it would be too long for a lambda but I may be totally off base here. I've not played around in detail with lambdas so I don't have particularly robust mental model on their limitations.
- slt2021 3y agoreading/writing Arrow is zero serde overhead operation to/from memory to disk. I think of lambda as a thread, and you can put a trigger on S3 bucket on each incoming file - to get processed. This allows you to get around GIL, and lets you invoke your lambda for each mini-batch. assuming you have high volume and frequency of data - you will need to "cool down" your high frequency data, and switch from row-basis (like millions of rows per second) to mini-batch basis (like one batch file per 100Mb). This can be achieved by having kafka with high partition number on the ingestion side, and sink to s3. from S3 for each new file your lambda will be invoked and minibatch will be processed by your python code, and you can right size your lambda's RAM, but usually I reserve 2-3x size of a batch file for lambda. the killer feature is zero ops. Just by tuning your minibatch size you can regulate how many times your lambda will be invoked
- Karrot_Kream 3y agoVery cool. Do you then further aggregate and load into a DB or vector store or something?
- palae 3y agoR also has data.table, which extends data.frame and is pretty powerful and very fast
- hermitcrab 3y agoR + data.table is a lot faster than Base R. See a benchmark of Base R vs R + data.table (plus various other data wrangling solutions, including our own Easy Data Transform) at: https://www.easydatatransform.com/data_wrangling_etl_tools.html https://www.easydatatransform.com/data_wrangling_etl_tools.h...