5 ms·
On Quora someone asked what the longest search query time was. I was able to craft a query that took multiple seconds to complete. It used wildcards and undocum
by carvalho 9y ago
On Quora someone asked what the longest search query time was. I was able to craft a query that took multiple seconds to complete. It used wildcards and undocumented iteration allowing one to stuff thausands of queries into a single query. Turns out it is someone's job to measure result response times, and he/she came into the thread to kindly ask us to stop messing up their statistics.
- logicallee 9y agoIt's hard to believe, but years ago, back when Google had what was called "stop words" (like 'the', that it ordinarily ignored) I was able to make Google perform a search that took over 30 seconds. The reason stop words take such a long time is that millions of sites have words like "the" on them, so doing a join on all those simply takes a long time. My method to find a long string consisting entirely of stop words, was to just download a project gutenberg of the complete works of shakespeare, and find the longest string consisting of just stop words in there, then search for it as a literal quote. The longest one I found was: "From what it is to a". Let me see how long Google takes to do it now :) 2.04 seconds! Nice :) - http://i.imgur.com/IhPTpr6.png http://i.imgur.com/IhPTpr6.png that took 30+ seconds 'back in the day'.
- sova 9y agoIs it really technically correct to say that Google was performing web-wide joins on data? Isn't it all about clever indexing?
- logicallee 9y agoThere's nothing to index. How could it have found my Shakespeare quote via an index? It consisted entirely of words 'from what it is to a' but produced only the Shakespeare quote. I don't see how it could have indexed anything.... it must have done a join. (Which makes sense given the 30+ seconds I had to sit and wait before it returned its answer, while also reporting the time it took to produce it. What else could it have been doing?) By the way I believe I wanted to know whether it would return the Shakespeare quote at all. If you mean that it might have cached the results of the query, I doubt anyone else queried that exact phrase, other than me.
- Benjammer 9y ago>There's nothing to index Huh? What do you mean? Google indexes HTML web page content from the entire public internet using web crawlers...
- logicallee 9y agoI'm confused. By "clever indexing" I thought they meant, in the database sense of the word. The reason my search took 30 seconds is because it started by getting a list of every site with "from" on it, every site with "what" on it, and so on, intereseecting them all. That's how it ended up finding my quote. how else do you think it did it? ----- edit: to find the string "from what it is to a" which occurs only hidden in the middle of shaespeare's texts -- what do you think they do? In my opinion they combine the list of sites that have every word - starting with the least common ones. It's easier if you search for something that has a few uncommon words. Then you start with a small list, and have to combine it with other small lists. When every word in the phrase has billions of sites (there are billions of pages that have the word "to" on them, same for "from", "what", "it", "is", "a"), you have to combine them all. Then you have to do a string search within the resulting set, since I put it in quotation marks. There is no easy strategy. Hence the long search time. what else could they be doing?
- Benjammer 9y agoYou said "There's nothing to index," as if Google is making web requests to every domain in existence, parsing the document responses, and seeing which sites have these words on them, all at runtime when you type a search query. Google obviously indexes the web in the sense that they store their own cached versions of web pages "locally," on top of which they then build an insanely complicated, web-facing, search architecture.
- logicallee 9y agowe're talking past each other. sova referred to this meaning - https://en.wikipedia.org/wiki/Database_index https://en.wikipedia.org/wiki/Database_index when they said "clever indexing." the sense you mean is a different sense of the word index - meaning, to crawl. Yes, of course it does that too.
- SamBam 9y agoInterestingly, I wonder if it cached your query. My same query as you took 0.3s, but if I stripped out one word ("From what it is to") it took 2.2 seconds.
- logicallee 9y agoof course it cached my query. :) try it again in a few weeks.
- falsedan 9y agoGoogle never had stop words. The original lexicon only included the most popular 14 million words (for fast bucketing), and rare words were processed specially.
- logicallee 9y agoIt did - I had to use plus signs to force them to use them. Normally it ignored those words. I am fairly certain of this detail. I must have found a list of those words - how else would I have found the string "from what it is to a"? I had a list of its stop words. Edit: for proof, here's someone's screenshot of the same - http://farm3.static.flickr.com/2270/2201828252_45a32da7f4.jpg http://farm3.static.flickr.com/2270/2201828252_45a32da7f4.jp... As you can see, it is Google saying it is ignoring a word because it is too common. It has a list of every site that has that, but that list is huge and it doesn't usually use it.
- falsedan 9y agoStop words are words that are ignored when indexing, not when querying. Since you did find a result, those words must have been indexed.
- dispo001 9y agoAlso long long ago... I was watching a video of a guy channeling aliens. Someone in the audience asked what would be the next big thing for humanity. I immediately typed the query into Google. It showed the page head.... 20-30 seconds later it came up with a single published paper about "nuclear magnetic resonance identification" Today it resolves instantly and finds at least 10 publications from before 2000.
- RugnirViking 9y agoI managed to beat that query with "The thing is what it is"
- vochinch 9y agoNice! Do you still have the link to the Quora question or an example of the query?
- rshaban 9y agolinked in https://news.ycombinator.com/item?id=14372977 https://news.ycombinator.com/item?id=14372977 https://www.quora.com/What-is-the-slowest-Google-query https://www.quora.com/What-is-the-slowest-Google-query
- dsacco 9y agoDo you have a link to that? I'd be interested in reading it his response and I can't see it by searching Quora.
- ahiknsr 9y agohttps://www.quora.com/What-is-the-slowest-Google-query/answer/Kevin-Lacker/comment/536273 https://www.quora.com/What-is-the-slowest-Google-query/answe...
- ClassyJacket 9y ago>and he/she came into the thread to kindly ask us to stop messing up their statistics. Shouldn't someone with a job in statistics know how to account for outliers?
- annnnd 9y agoAnd also, wouldn't they be interested in getting those queries so they could either fix their performance or block them?
- developer2 9y agoThe query was already published on Quora. I assume that Google did patch the issue. They were simply requesting that it not be posted in a public forum, resulting in the potential for denial of service attacks. It wouldn't have been a statistician making a fuss about their pretty graphs being ruined. The problem was the very real performance impact the queries were having on the service, as thousands of visitors copy/pasted from Quora to see for themselves. This is why companies like Google have bounties for such things. "Please submit bugs and performance issues privately so we can patch them before you disclose the details publicly and hurt our services - we'll even pay you for your discretion!"
- developer2 9y agoThis particular issue was posted on Quora, where anyone could pick it up and participate in what is essentially a denial of service attack (whether or not performed intentionally). It wasn't submitted as a private bug report to Google so they could fix the issue. It was spread in a public forum. I think it's fair for Google to politely ask "a few of your own tests to validate an issue you will submit as a bug report is fine, but please don't disclose to the public until we patch it." When you operate at the scale of Google, everything is expected to be airtight; outliers should not be possible. It wouldn't surprise me if their monitoring systems are built without the ability to "massage" (ie: manipulate) statistics, as it is a terrible practice. I don't think a statistician who relies on ignoring outliers would last long working for Google. They're not doing their job if the only thing they care about is silencing warnings to make pretty graphs that falsely show everything is running smoothly. Their job is to work with the truth - not manufacture little white lies to appease management.
- hoschicz 9y agoThis is his answer: "I work on search at Google, and I have to say, very clever answers! Now, please stop. :-p" I don't think that he did it because it's his job to stop random people on the Internet from running slow queries. I think she was just surprised how creative people are and found it funny.