3 ms·
We're processing the opinions rather than the courts, so we're dealing with millions of documents. Since we're building a network of their citations, it winds u
by rjsen 10y ago
We're processing the opinions rather than the courts, so we're dealing with millions of documents. Since we're building a network of their citations, it winds up being way too much data to hold in memory on a single node, hence the need for Spark.