35 ms·
The sparse grams solution to deal with stupidly common ngrams such as for or tes is very interesting. I’d love to see more discussion on how they are dealing w
by boyter 4y ago
The sparse grams solution to deal with stupidly common ngrams such as for or tes is very interesting.
I’d love to see more discussion on how they are dealing with the false positives though. It looks like a positional index is being used to achieve this, but that usually blows out your index size.
Additional information about deduplication would be especially interesting to me as well. It seems to solve this quite well. I usually try a search of Jquery to test this and it does not return multiple copies of different versions of it which is a good indicator that it’s slightly fuzzy.
What I find really interesting about all the code search engines I know of is that each one implemented its own index. Nobody is using off the shelf software for this. I suspect that might be down to no off the shelf software providing a decent enough solution, and none providing a solution that scales. At least none that scales with decent costs.
I did a small comparison of GitHub code search a while ago https://twitter.com/boyter/status/1480667185475244036?s=61&t=nfND46d9rReCju7-aw457Q https://twitter.com/boyter/status/1480667185475244036?s=61&t... But I should note a lot has improved since then, and it looks like sourcegraph now also does default AND of terms rather than exact match, so my complaints there are resolved.
Impressive work by GitHub. I am sure some of the people behind it will read this comment, let me say well done to you all. I am very impressed. Also please post more information like this. There is so little out there.
- ZephyrBlu 4y agoImplementing your own index gives you more control over it. I think at this scale you probably want to tweak things specifically to your product rather than using a generic solution. I would guess that what you're indexing on (E.g. language, file, repo, etc) and sharding strategy affects the structure of your index as well.
- boyter 4y agoBelieve me I am aware. I am one of those who implemented their own index for a code search engine :) I did it for my own learning, but find it interesting because something like elastic with trigrams can get you very close, albeit at a far greater cost.
- ZephyrBlu 4y agoI'm reading your blog posts about building your own index now. I started writing my own very simple index and search engine, but quickly decided to just use ClickHouse via https://tinybird.co https://tinybird.co as my backend (Serverless SQL with automatic APIs is pretty sweet) because I wanted to build out the product side of things and my data is really small, so I felt like it was going to be a lot of effort for little reward. Maybe one day I will need to write a custom index or search engine that actually scales though :).
- boyter 4y agoI won’t hijack this thread with details but if you have questions you can find my details on my profile.
- 100k 4y agoThanks! I enjoyed reading your blog posts about building your code search engine. One minor point of clarification, we do not use a positional ngram index, which as you note blows up the index size. Instead, we use the covering sparse ngrams to produce candidate documents and then search the content. An early version of Blackbird experimented with trigrams plus a bitmask of the next character, but it didn't work well because it wasn't selective enough. This is mentioned in the blog post: We tried a number of strategies to fix this like adding follow masks, which use bitmasks for the character following the trigram (basically halfway to quad grams), but they saturate too quickly to be useful.
- boyter 4y agoThat's what I get for a cursory glance at 4am when I wrote this. I will have a much better look after I get some coffee into me. Thanks for the clarification. Looking forward to see what else you and your team end up writing about. Which reminds me to publish some other posts I have about searchcode.
- boyter 4y agoCannot edit previous reply, but I would love to know more about how the sparse grams work. There isn't enough detail in the post, just a few tantalizing crumbs of information. Seems a lot of others in this thread are interested as well.
- therealdrag0 4y ago> What I find really interesting about all the code search engines I know of is that each one implemented its own index I mean GH got a long way using ElasticSearch until now.
- masklinn 4y ago> I mean GH got a long way using ElasticSearch until now. I'm not sure that's true. It used ES for a long time, but the search was also terrible, at the edge of complete uselessness.