4 ms·
Hi, I work on GitHub code search. Grep.app is based on Solr and indexes about 500,000 public repositories. As far as I know, the author has not shared their tok
by 100k 4y ago
Hi, I work on GitHub code search. Grep.app is based on Solr and indexes about 500,000 public repositories. As far as I know, the author has not shared their tokenization strategy, but based on the the results I expect they are indexing trigrams to perform regular expression matching.
The current GitHub code search is based on Elasticsearch and indexes more than 100M repositories. Its tokenization is based on whitespace, case changes (like CamelCase) and punctuation (like kebab-case) and strips out non-letter characters like < or { which is why it can't do exact match search.
Our new code search, currently in beta and indexing about 45 million repositories, uses an search engine we built in house that indexes content using a technique we call sparse ngrams. This allows us to execute searches faster than a trigrams index, while also being smaller than a positional trigram index. My teammate discussed some of the technology behind it in our blog post that was published yesterday: https://github.blog/2023-02-06-the-technology-behind-githubs-new-code-search/ https://github.blog/2023-02-06-the-technology-behind-githubs...