7 ms·
> Yes, Google has a huge index, but most queries aren't in the long tail. I'm not quite sure about that. 15% of Google searches per day are unique, as in, Goog
by dkyc 7y ago
> Yes, Google has a huge index, but most queries aren't in the long tail.
I'm not quite sure about that. 15% of Google searches per day are unique, as in, Google has never seen them before. [1]. That's quite an insane number.
[1] https://searchengineland.com/google-reaffirms-15-searches-new-never-searched-273786 https://searchengineland.com/google-reaffirms-15-searches-ne...
- vanderZwan 7y agoHow many of those are confirmed to be of human origin?
- Broken_Hippo 7y agoProbably quite a few. New things happen. Politics, wars, famous folks, movies, music, diseases, scientific studies, products, brands, model numbers for products, fads and slang. I'm guessing there are other things as well. Some of the new things are probably variation as well - as others have mentioned, sentences and voice commands can give lots of new stuff.
- rocgf 7y agoWow, 15% unique searches is indeed quite an interesting figure. With that said, what OP said is definitely not disproved. Just because 15% of searches are unique, that doesn't mean the most relevant result is buried in the tail end. I mean I can think of loads of my own searches that are probably unique or rare, but lead to the same popular results because of typos, improper wording etc. Without some clear numbers on that from a major search engine, I think this might be very difficulty to infer.
- i_cant_speel 7y agoEspecially with voice searches. People are searching entire sentences rather than specific keywords which are much more likely to be unique.
- b_tterc_p 7y agoDo people do this? Or do you mean queries forwarded by home assistants trying to parse inputs?
- quirkot 7y agocan confirm. i search full sentences even from the keyboard
- JustSomeNobody 7y agoI search full sentences (questions) from the keyboard. I figure I'm not the only to have had the question before, so I ask. Also, I find that blog posts, etc. tend to match well for full sentences.
- have_faith 7y ago> Do people do this? The calling card of the developer realising that real users never act like you expect :)
- steve76 7y agoSearch is nothing. They already have the brand, browser, and devices. If someone wants an online ad, they go straight to google. Not any other network which is probably shadow owned by them, like my ability to buy food is their secret gay fraternity hazing scheme. I suggest you take a look at the fraud associated with Adsense bans. They make a ton of money, and then, drive you into the poor house by just not sending your check. Pretty much, Google knows we took your life savings. Fuck you. Die.
- zadkey 7y agoReal users will use your product in ways you never imagined.
- sct202 7y agoI often do full sentences and then start deleting words from it if it doesn't work.
- Angostura 7y agoYes, sorry - that's me. Copy and pasting Sharepoint error messages
- thaumasiotes 7y agoThose searches are unlikely to be unique.
- Angostura 7y agoHmmm - Error: System.InvalidOperationException: The workflow with id=15f08b34-33f5-4063-8dea-d4ca6212c0d6 is no longer available. is not atypical.
- buildzr 7y agoDoes that actually work? I must be old school, I always delete such IDs before searching, but then again I used Google back when it actually did what you told it instead of misinterpreting everything for you.
- Angostura 7y agoIt doesn't seem to have any particular effect on the results that come up. I always used to delete them, and still do sometimes but Google seems to pretty much ignore them in practice.
- jschwartzi 7y agoWhich is a wonderful behavior except for all the times that the error numbers are not actually GUIDs but rather identify general errors.
- jjoonathan 7y agoIf only :(
- whalabi 7y agoNow I feel bad for putting gibberish like jsjsjdkktkwoapaoalf in my address bar and searching Google to test if my internet is working..
- chrshawkes 7y agoI do that all the time, I wonder how common that is?
- jamiek88 7y agoI would think it’s pretty common. For a lot of people google is the internet. Or at least the reference. If google isn't working it’s almost certain it’s your end. I don’t think anyone else has that reputation for availability amongst the general public.
- ElijahLynn 7y agoI just type "test", hopefully they do that too and it is ignored.
- utefan001 7y agoSharing for anyone who didn't know there is a very good dataset you can use now. If you don't have a nvme ssd in your computer, I highly recommend getting one for fast i/o. http://commoncrawl.org/ http://commoncrawl.org/ http://commoncrawl.org/the-data/ http://commoncrawl.org/the-data/ http://index.commoncrawl.org/ http://index.commoncrawl.org/ related.. Mark's blog is amazing and worth more than any data science degree imho. https://tech.marksblogg.com/petabytes-of-website-data-spark-emr.html https://tech.marksblogg.com/petabytes-of-website-data-spark-... https://tech.marksblogg.com https://tech.marksblogg.com
- kingludite 7y agowow, thanks. [edit] in my experience yacy works really well. You have it crawl the sites you frequently visit and their external links and it quickly accumulates to something more accurate than google.
- jaytaylor 7y agohttps://yacy.net/en/index.html https://yacy.net/en/index.html
- nilkn 7y agoCould this be explained by supposing that people are just searching for current events, sometimes national, sometimes international, sometimes very local? If so, you really wouldn't need much indexed to handle those queries. I imagine many queries are also just overly verbose and sentence-length, which artificially inflates the number of unique queries which are actually seeking roughly the same pages.
- Kovah 7y agoGood point and 15% is indeed much, but the question would be what "unique" means. If it means that the exact same character sequence appeared for the first time, it doesn't mean that the users searches for a term that has never been searched for. I mean with the newest advantages like machine learning it's more and more possible to _semantically_ link queries. If that's the case, those 15% could become 5% truly unique searches or even less. "how dumb is trump" and "how dumb is donald trump" are two different searches but they semantically belong together because they mean the same.
- kjeetgill 7y agoI think they mean that the results are still from the top pages of the internet. They mean long tail of visited pages, not long tail of searches. A unique search query could still land you on Wikipedia.
- madez 7y ago> 15% of Google searches per day are unique, as in, Google has never seen them before. That is impossible, and therefore wrong (I'm wrong, please see below). To know if a search is unique, as in Google has never seen them before, Google must be able to decide if a query it receives was seen before or not. Even if we assume Google needed only one bit for each message it has ever seen, and assuming it only saw 15% of new messages each day since its creation more than 20 years ago, it would need to store more than 2^1471 bits. What could be true is that each day 15% of all searches are unique on that day. Edit: I'm wrong. The 15% of completely unique messages per day are in regards to the messages per day, and not in regards to all messages it has ever seen, therefore exponential growth doesn't apply. To see that, assume Google just received one search query each day for 20 years but it was unique random gibberish, then Google could easily save that even though 100% of all messages per day are unique.
- Strilanc 7y agoHow are you computing that number? It's definitely wrong. Assume Google receives 1 trillion queries per year, and has been around for 20 years. Using a bloom filter you can achieve a 1% error rate with ~10 bits per item. So a 200 terabyte bloom filter would be more than sufficient to estimate the number of unique queries.
- deleted 7y ago[deleted]
- saalweachter 7y agoA Bloom filter is just way overkill. If you have a list of 20 trillion query strings, and each query string is on average < 100 bytes, you're looking at a three line MapReduce and < 1 PiB of disk to create a table which has the frequency of every query ever issued. Add a counter to your final reduce to count how often the # times seen is 1.
- sloppycee 7y agouh, is this sarcasm? A bloom filter is the most appropriate data structure for this use-case. How is it overkill when it uses less space and is faster to query?