5 ms·
Next post from danfox - “how to get 3 job offers in 3 hours”. Already has been publicly contacted by: - GitHub CTO - SerpApi CEO - SourceGraph CEO Search i
by thanatos_dem 7y ago
Next post from danfox - “how to get 3 job offers in 3 hours”.
Already has been publicly contacted by:
- GitHub CTO
- SerpApi CEO
- SourceGraph CEO
Search is hot right now!
- Existenceblinks 7y agoI'm surprised as well, think why big tech companies didn't have this awesome search already.
- thanatos_dem 7y agoIf this were to be offered by an actual company (a first party solution), there are some features that'd be expected that make the problem space a lot harder. Here's an "intro to search" article that's a good read, and I'll use it to highlight some of the things that'd be different in a first party solution - https://medium.com/startup-grind/what-every-software-engineer-should-know-about-search-27d1df99f80d https://medium.com/startup-grind/what-every-software-enginee... (See the "Theory: the search problem" section) Size: This is only indexing ~500k public repos. A first party solution would be expected to index all of it, public and private. Indexing speed: This can take up to a few days to index. A first party solution would be expected to have a much lower index latency - seconds to minutes. Query language: This can (and does) have its own simple query language. A first party solution would need to have support embedded into and not break backwards compatibility with the current query language. Context-dependence: A first party solution would be expected to index private repos as well, and now the query context (logged in user) becomes another variable in an already multi-variate problem space. Latency: Gets harder with scale, and a first party solution would likely provide a SLA/SLO around latency. Access control: Same issue as context-dependence, with private repos being included. There's also unknown but likely considerations around compliance and internationalization, which are quite tricky problems. Note - I don't mean for this to be critical of the author at all. This is an awesome and useful tool, with a fantastic UX. I just want to make it clear that search at scale is a lot harder than it seems at first glance, especially as the feature requirements increase.
- sdesol 7y agoFor GitHub, I would have to imagine only being able to search public repos with regexp would be good enough. GitHub has many strategies, but the main one is, they want to maintain, if not, expand their open source mind share. The more reasons you give people to go to GitHub, the better off they will be in the future. So I do agree with you that as a commercial solution, this may not be viable, but for GitHub's public repos, this can turn into a very positive thing.
- marceloabsousa 7y agoThat might well be true but to scale this type of service to all public repos with decent latency and update ratio is a major technical challenge and likely very costly to maintain.
- sdesol 7y agoThis is my personal observation, but GitHub appears to be a much more ambitious company, now that they are part of Microsoft. With a CEO that understands both the open source and the enterprise world and with Microsoft cash at hand, I don't think spending money to make search better would cause any concerns. Doing technical things that GitLab, Bitbucket, etc. can't is quite valuable. It also helps with recruiting, since smart people want to work on difficult problems. It may well be costly to maintain, but I think the operating cost would be well within the realm of an incumbent that wants to maintain and expand their reach. I've been studying the code hosting space for quite sometime and GitHub, from an outsiders perspective, appears to be much more focused and ambitious, which should cause serious concerns for GitLab.
- fjania 7y agoEngineering manager for code search at GitHub here... this is an excellent summary of many of the concerns we have as we work on code search at GitHub scale!
- sdesol 7y agoIt also probably goes without saying he should be careful with what details to share.
- neonate 7y agoAlso by the co-creator of Django: https://news.ycombinator.com/item?id=22397023 https://news.ycombinator.com/item?id=22397023
- swat535 7y agoActually, It would more be like: "How I failed at 3 interviews, despite being directly contacted by execs."
- runawaybottle 7y agoThat would be quite the dystopian interview nightmare.
- nickthemagicman 7y agoSure you built app on multi 20 core machines with functionality to search hundreds of millions of lines of code almost instantaneously, but are you someone I'd drink a beer with?
- tmpz22 7y agoFor some of those companies it would be "drink a La Croix with"
- deleted 7y ago[deleted]
- yakshaving_jgt 7y agoThis snide remark dismisses the fact that working on software does mean working with other humans, not just unemotional robots devoid of any kind of irrational ideas. Being able to “drink a beer with” (and reasonably substituting the drinking of beer for just about any other social interaction) is an important part of being able to work with someone. Unless of course you believe an office environment consisting of a tyrannical manager barking orders at worker drones is a healthy relationship.
- IshKebab 7y ago100%. I don't really care if you're a super genius if you're also a massive dick that everybody hates.
- nonbirithm 7y agoIf only the answer to "how" was as simple as "writing a web service for searching GitHub repos with regexes," even though the problem is probably in itself non-trivial if there's this much interest in search at all. At least the specification is clear enough. I guess what I mean to ask is, how would people know this is a "correct" answer to the "how" question beforehand? Is the answer literally just "search" because that's simply what's trending right now?