3 ms·
I'm also working in biology (at the UCSF Cancer Center) and one of the reasons why the future package exists in the first place is that we needed a way to proce
by HenrikB 10y ago
I'm also working in biology (at the UCSF Cancer Center) and one of the reasons why the future package exists in the first place is that we needed a way to process a large number of microarrays, RNA / DNA sequencing samples and HiC data (and I'd like to use everything from the R prompt). We have a large compute cluster available for this (sorry @apathy but we're using TORQUE but hope to move to Slurm soon).
Now our R script for sequence alignment basically looks like:
## Use nested futures where first layer is resolved
## via the scheduler and the second using multiple
## cores / processes on each of the compute node.
library("future.BatchJobs")
plan(list(batchjobs_torque, multiprocess))
fastq <- dir(pattern = "[.]fq$")
bam <- listenv()
for (ii in seq_along(fastq)) {
fq <- fastq[ii]
bam[[ii]] %<-% {
bam_ii <- listenv()
for (chr in 1:24) {
bam_ii[[chr]] %<-% DNAseq::align(fq, chromosome = chr)
}
as.list(bam_ii)
}
}
The future.BatchJobs package (https://cran.r-project.org/package=future.BatchJobs https://cran.r-project.org/package=future.BatchJobs), which enhanced the future package, uses the framework of the BatchJobs package as its backend.