4 ms·
(I worked on this.) This is on our radar! We de-duplicate exact matches now, but we'd like to do the same for near-similar documents.
by 100k 5y ago
(I worked on this.)
This is on our radar! We de-duplicate exact matches now, but we'd like to do the same for near-similar documents.
- elliottcarlson 5y agoDe-duping exact matches is a game changed -- search has been miserable to use because of the dupes for so long. I can live with near-similar documents. Very excited to test this out.
- colin353 5y agoAnother GitHub Code Search developer here - to add more to this, we rank all the search results, and try to bring the most relevant results to the top. Ideally, if you have 10 pages of results, you shouldn't have to leave page 1 to find what you're looking for :D
- sumtechguy 5y agoThat would be a tough problem. As de-dup you probably want to show/point towards the 'original' tree. But which one is the source? Or even worse someone abandons a project but someone else forked it and kept going should it show that one instead? Or should it show the one it was forked from depending on the version number. Which one is the 'true' repo now? Most certainly an interesting problem.
- adamnemecek 5y agoI kind of don't care about correctness. Just hide results that seem to be duplicate.
- sumtechguy 5y agoI get that. Just remove the 'extra'. That is a good first pass. I was thinking the longer term you want to show the 'original' higher in the list? Wouldnt you? What sort of criteria would you use to make it so it shows one copy vs another? Probably in many cases it probably would not mater much. But if you wanted to figure out linage of imports it could be? Some projects could have thousands of forks. Yet only maybe a dozen of those actually have anything going on. Those would be more useful to show?