9 ms·
We use this to power things like find-references or jump-to-def, "symbol search" and autocomplete, or more complicated code queries and analysis (even across la
by dons 5y ago
We use this to power things like find-references or jump-to-def, "symbol search" and autocomplete, or more complicated code queries and analysis (even across languages). Imagine rich LSPs without a local checkout, web-based code queries, or seeding fuzzers and static analyzers with entry points in code.
Our focus has been on very large scale, multi-language code indexing, and then low latency (e.g. hundreds of micros) query times, to drive highly interactive developer workflows.
- the_duke 5y agoThis is really cool. Seems like there are only indexers for Flow and Hack though. Will there be more indexers built by Facebook, or will it rely on community contributions?
- dons 5y agoA bit of both I think.
- simonmar 5y agoThere will be more indexers: we have Python, C++/Objective C, Rust, Java and Haskell. It's just a case of getting them ready to open source. You can see the schemas for most of these already in the repo: https://github.com/facebookincubator/Glean/tree/main/glean/schema/source https://github.com/facebookincubator/Glean/tree/main/glean/s...
- soonnow 5y agoDoes that mean you are using the shell or how is it used to enable these functionalities?
- dons 5y agoMost clients hit the Glean server via the network (thrift/JSON) and then mostly via language bindings to the Glean query language, Angle. The shell is more for debugging/exploration. Imagine an IDE plugin that queries Glean over the network for symbol information about the current file, then shows that on hover. That sort of thing.
- soonnow 5y agoAlright gotcha. Thanks for the clarification.
- progval 5y agoHow would it perform for, say, 500TB of source code? And what would be the disk and memory requirements for this? Could they be distributed across a handful of servers?
- dmos62 5y agoI'd be surprised if this question could have an off hand answer. Doesn't sound like something that could have scalability predictable enough to do back of the envelope calculations on.
- gricardo99 5y agoWhat on earth has this much source code? Every open source project ever?
- gurleen_s 5y agoI mean, yeah. Imagine being able to do more rich queries against GitHub.
- progval 5y agoYes, good guess! That's the size we have after deduplication across projects at https://www.softwareheritage.org/ https://www.softwareheritage.org/ . We archive all the source code we can find; and would like to support some sort of full-text search on it at some point, so Glean looks interesting
- pdpi 5y agoBeen away from Fb for a few years. How does this relate to tbgs?
- gaogao 5y agoJump to def is nice when biggrepping a piece of code a la what you can do with codesearch, cs.android.com
- gravypod 5y agoI see you support Thrift and Buck. Would you also be interested in adding Proto and Bazel support? Being able to query the code based on the build graph (sort of) would be very cool.
- zerr 5y agoSince this is HN, could you please share more technical/impl details, e.g. what makes it more scalable and faster in general and also compared to other similar engines?
- gwbas1c 5y agoI'm really struggling to understand what Glean does, and why I would use it. Most important: Your landing page should quickly show what Glean does that a typical IDE (Visual Studio, Visual Studio Code, Eclipse, ect, does.) Specifically, things like "Go to definition," and tab completion have been in industry-leading IDEs for at least 20 years. What's novel about Glean? It seems like a lot of hoops to jump through when Visual Studio (and Visual Studio Code) can index a very large codebase in a few seconds. (And don't require a server and database to do it.) Perhaps a 20-second video (no sound) showing what Glean does that other IDEs don't will help get the message across?
- WastingMyTime89 5y ago> It seems like a lot of hoops to jump through when Visual Studio (and Visual Studio Code) can index a very large codebase in a few seconds. I think you are not thinking large enough. An IDE absolutely can not index a very large codebase and allow users to make complex queries on it. Think multiple millions lines of code here. The use case is closer to "find me all the variables of this type or a type derived from it in all the projects at Facebook" than "go to this definition in the project I'm currently editing".
- Syzygies 5y agoThere's large, and there's scope. I use VSCode to dabble in dozens of projects across a dozen languages at a time, often coming back to fix things after years. VSCode is great at telling me what I did in the current project, but I can't remember library calls or even syntax without looking at something I wrote before. My efficiency is perhaps 50% at recalling where to look; a tool that kept my entire corpus at my fingertips would be extremely welcome. But I'm failing to see how this is that.
- fragmede 5y agoIf you've not had to deal with a codebase that takes VSCode longer than a few minutes to index, then you're probably outside their initial target market. If you've not had to setup a hosted code search tool (eg livegrep https://github.com/livegrep/livegrep https://github.com/livegrep/livegrep ) because there's just too much code, you've been lucky. If your projects can be scoped, and not pull in code from dozens of libraries, across dozens of teams, many of which are on different continents, you're doing a better job of organizing code than I've been able to manage.
- mhitza 5y agoBriefly skimmed the docs and it noted that it doesn't store expressions from the parsed AST. That means it's mostly a symbol lookup system? When doing large system refactoring searching by code patterns is the number one thing I'd like to have a tool for. For example being able to query for all for loops in a codebase that have a call to function X within their body.