3 ms·
I’ve taken ideas from blog posts like that one, but mostly I’ve found ideas and haven’t really seen a well described approach. There’s a lot of details to be wo
by deepsquirrelnet 2y ago
I’ve taken ideas from blog posts like that one, but mostly I’ve found ideas and haven’t really seen a well described approach. There’s a lot of details to be worked out which I feel are all important in the process. I probably spent more time figuring out how to distribute the text into small segments to even perform similarity comparisons than anything else.
That was one motivation I had in doing a full write up. It has some aspects of a rolling sentence method (rolling window similarity), and some aspects of regular chunking (trying to retain paragraph structure).
I wanted to demonstrate an approach that I worked on with all the gory details for people to follow if they want to understand what’s happening under the hood or take ideas for their own experiments.
- magicalhippo 2y agoThanks again. Definitely something I want to play with. The efficiency of your approach is very nice, thought it would be interesting to compare with the "rolling embeddings" approach to see if the quality stays sufficient. If so it's a huge gain in speed.