3 ms·
>being able to initialize and share data between functions within the R runtime That's right. >rather than having to transfer back and forth between Postgres
by blahi 10y ago
>being able to initialize and share data between functions within the R runtime
That's right.
>rather than having to transfer back and forth between Postgres and the runtime
But that's not. There's a difference between being able to use data outside of the database (from the R runtime) in my UDFs (executed in Postgres) on one hand and being able to attach 2TBs of data straight from an SQL table in the R runtime on the other. I don't even care that much about the algorithms. Moving the data is the bottleneck most of the time. And Microsoft is actually late to the party (but better than never). Oracle, Netezza, Vertica and Hana have been able to do it for quite a while now.
You are spot on about being able to use the algorithms outside of SQL Server. You can use them on Teradata or Hadoop or rent your own VMs on Azure to use them or you can buy standalone licenses too.
- Jweb_Guru 10y agoSo by "attach 2 TB of data straight from an SQL table into the R runtime" you mean that Microsoft taught R to interact directly with SQL Server's storage engine? If so, I agree, data movement is almost always the bottleneck for large data sets, and I don't think PL/R can do that (though I am not sure if that's a necessity due to the way Postgres's language plugins work, or something that could be done with enough effort). However, if all you mean is that SQL Server can transfer the data a tuple at a time to R on the same server (in memory), I believe that PL/R and Postgres interact like that already (again, maybe I'm wrong). And I don't know how much extra overhead that provides over talking directly to the storage engine, anyway.
- blahi 10y ago>Microsoft taught R to interact directly with SQL Server's storage engine They have created 2 new services for SQL Server 2016 - BxlServer and SQL Satellite which facilitate the communication and data exchange. They obviously have additional speedups for the proprietary runtime (that was one of the main selling points of the company they acquired - fast data access to several RDBMS), but it's plenty fast for regular R too. https://msdn.microsoft.com/en-us/library/mt709082.aspx https://msdn.microsoft.com/en-us/library/mt709082.aspx