2 ms·
Maybe I misunderstood your problem, but how about a relational DB model with: - FEATURES table (~400k rows), primary key: Feature_ID - SAMPLES table (tens of
by tzmudzin 9y ago
Maybe I misunderstood your problem, but how about a relational DB model with:
- FEATURES table (~400k rows), primary key: Feature_ID
- SAMPLES table (tens of thousands), primary key: Sample_ID
- OCCURRENCE table (what you observed), with the fields Feature_ID, Sample_ID, Occurrence_Count.
This is pretty much a standard solution well proven in data warehousing...
- gaius 9y agoIf you process this kind of high-throughput data in R then yep, you get it as 3 data.frames as you describe, it's nicely tabular and easy to "join" by just matching index numbers.