3 ms·
I would suggest sentence tokenization followed by clustering to pull the most representative sentence from each cluster. Unless you want to go abstractive, that
by pilotneko 4y ago
I would suggest sentence tokenization followed by clustering to pull the most representative sentence from each cluster. Unless you want to go abstractive, that is.
If you run into issues with the token limit, there are some nice architectures for handling longer inputs (e.g., Longformer, Reformer, etc.).
- subtech 4y agointeresting ideas, thanks!