12 ms·
You can work around this by chunking the data using a sliding window. Take the first X amount of sentences that fit into the token length and summarize. Now sli
by ca_tech 3y ago
You can work around this by chunking the data using a sliding window. Take the first X amount of sentences that fit into the token length and summarize. Now slide your window of sentences so that you still overlap just a bit with your previous selection. Summarize that information. Continue until you have a bunch of summaries for your text. You can then pass that back into the model for a more concise summary of the summaries.
- cced 3y agoI’ll be exploring a combination of pre-processing such as stemming in order to also reduce the token length while preserving as much important information. Thanks for the feedback. Edit: Won’t the sliding window solution above introduce a sort of bias? For instance, with a sliding window of 3 units, unit 1 is captured once whereas units 2 and 3 are captured twice and three times respectively.