4 ms·
The types of queries we run against it don't require real-time results, and we do a pretty heavy amount of subsetting. By the time it reaches the point where we
by christopheraden 13y ago
The types of queries we run against it don't require real-time results, and we do a pretty heavy amount of subsetting. By the time it reaches the point where we do numerical summaries and statistics, the largest set I've worked with here was around 30GB. Most times it's around 5-10GB.
- xtacy 13y ago5-10GB seems small for analysis. Is this sampled across your historical records, or just the most recent? What's the turnaround time for stats today, and what's your pain point? (Slow IO? Lack of high level programming frameworks? Or something else?)
- christopheraden 13y agoMost of the work I do involves recent data--cycles are six months at most and 3 months on average. 3 months of data, sifting by a pretty strict filter, it's not unsurprising that hundreds of terabytes of claims over years and years gets filtered into a few GB. Start to finish on jobs is a few hours (though it can be a few days if the filter is less strict or the task is more complicated), including pulling the data from the warehouse. Without a doubt, the bottleneck of the process is the data warehouse query. I'm sure having a more distributed database (it's DB2--I've been more pleased with Teradata's speed) could make the queries faster, but things are slow to change. Second to the query, the latency from working with a remote server (I'm in CA--the server's in Minnesota) adds latency if there's something I need to pull to the local machine. The actual SAS code (Is it still "high level" if the syntax models Fortran? I kid, I kid.) takes a negligible amount of time compared to the query.