24 ms·
Ha, yes, I've done that at https://gigablast.com/ https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper
by gbmatt 5y ago
Ha, yes, I've done that at https://gigablast.com/ https://gigablast.com/ .
The biggest problems now are the following:
1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages.
2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive.
I believe my algorithms are decent, but the biggest problem for Gigablast is now the index size. You do a search on Gigablast and say, well, why didn't it get this result that Google got. And that's because the index isn't big enough because I don't have the cash for the hardware. btw, I've been working on this engine for over 20 years and have coded probably 1-2M lines of code on it.
- justinzollars 5y agoThis is great! I found something other engines do not pick up! apparently I signed an agile manifesto in 2010 https://agilemanifesto.org/display/000000190.html https://agilemanifesto.org/display/000000190.html
- collin128 5y agoHave you ever looked at the Amazon file? I'll see if I can track down the link but I remember somebody sharing a dump with me from Amazon that apparently was a recent scrape. Edit: https://registry.opendata.aws/commoncrawl/ https://registry.opendata.aws/commoncrawl/
- web007 5y agoThat's Common Crawl, they do the spidering of some billions of webpages but that's still a tiny percentage of the web versus Google or Bing.
- cschmidt 5y agoDo you have any stats on that? I've always wondered about the coverage of Common Crawl, if you include all the historical crawl files too.
- visarga 5y agoCommon Crawl is being used to train the likes of GPT-3 and mine image-text pairs for CLIP. I wonder how much useful content is missing, we're going to use all the web text, images and video soon and then what do we do? We run out of natural content. No more scaling laws.
- collin128 5y agoOh interesting, I've played with it a little but not a dev and I've always wondered what the coverage was like.
- carlesfe 5y agoGreat job, I didn't know aboug Gigablast and it looks very interesting. Can I give you a small piece of feedback? I just tried searching for myself on Gigablast, and the first results are profile pages which haven't been updated since like 2005. Meanwhile, my own personal page appears on the very bottom of the results. So my suggestion would be to lower the weight of the ranking of the domain, and promote sites which have a more recent update date. Send me an email (contact in profile) if you want to follow up on this feedback!
- _HMCB_ 5y agoThe Internet is such a fabric of society that I think all nations should contribute to a one-truth index. Not owned by a corporate entity. Tell me I’m wrong and we can consider the alternative: startups of all types with a more even playing field.
- JPKab 5y agoRegarding the Gatekeeper companies like Cloudflare, it sounds like anti-competitive behavior that could potentially be targeted with anti-trust legislation, correct?
- shashashasha___ 5y agoi would assume its mostly anti scraping protection which is mostly for privacy. you don't want to allow everyone scrap your website, pull and use your info. for example from fb, ig, LinkedIn, github, .... you can build a really big profiling db on people that way. so websites need to know you are a legit search engine first
- karmanyaahm 5y agopeople can still be targeted if that information is public. anti scraping sounds like security by obscurity
- gbmatt 5y agoit should be. there should be some sort of 'bots rights' to level the playing field. perhaps this is something the FTC can look into. but, as it is right now big tech continues to keep their iron grip on the web and i don't see that changing any time soon. big tech has all the money and controls access to all the data and supply chains to prevent anyone else from being a competitive threat. look at linkedin (owned by microsoft unspiderable by all but google/bing). github (now microsoft using this to fuel its AI coding buddy, but if you try to spider this at capacity your IP is banned) facebook (unspiderable) .. the list goes on and on .. and as you can see, data is required to train advanced AI systems, too. So big tech has the advantage there as well. especially when they can swoop in and corrupt once non-profit companies like openai, and make them [partially] for-profit. and to rant on (yes, this is what i do :)) it very difficult to buy a computer now. have you tried to buy a raspberry pi or even a jetson nano lately? Who is getting preferred access to the chip factories? Does anyone know? Is big tech getting dibs on all the microchips now too?
- technobabbler 5y agoCloudflare functions kinda like a private security company. They don't go around blocking sites willy-nilly, site owners have to specifically choose to use their service (and maybe pay for it), configuring the bot blocking rules themselves. That's not really Cloudflare's fault. Someone has to do it, whether it's them or a competitor or sys admins manually making firewall rules. Cloudflare just happens to be good enough and darned affordable, so many choose to use them. Hosting costs for small site owners would be much more expensive without Cloudflare shielding and caching.
- mrkramer 5y agoI'm sorry to say but your project is 20 years old and it had no impact at all. You are doing something wrong. Innovation and initiative is needed ala Bitcoin and DeFi not hobby projects which are not picking up in popularity and utility.
- ErrrNoMate 5y agoBitcoin and DeFi don't have utility outside of gambling and pump and dumps. Not everything (tbh not really anything) needs crypto.
- jquery 5y agoCrypto’s biggest achievement is being the financial equivalent of the gulf war oil fires. Just massive pollution. Think of all the good things that computing could be used for… we used to have all kinds of interesting collaboration projects. Instead we are setting those CPU cycles on fire for short term profit.
- Sohcahtoa82 5y agoImagine if all that processing power was used for Folding@Home. The problem is that cryptocurrencies do not inherently need tons of processing power to operate. You could theoretically run the entire Bitcoin network on a Raspberry Pi. But the PoW algorithm was designed to always produce a block every 10 minutes, no matter how much hashing power was dedicated to the network. Everyone wanted a piece of the block reward pie, so the arms race was created. Proof-of-stake algorithms would eliminate this problem entirely, but PoS is a shitty "rich get richer" method. Granted, with how expensive mining power is, even PoW results in the rich getting richer, but at least it doesn't result in the wasting of gigawatts of electricity.
- ilammy 5y ago> Everyone wanted a piece of the block reward pie, so the arms race was created. And that's intentional – getting people pursue the goal for their own egoistic reasons, because that's bound to succeed. As a result, they all increase the security and stability of the network whether they want it or not, the only way to not do this is to not participate. If the network were running on a single Raspberry, someone bringing two Raspberries could effectively outcompete the other person on block rewards. I'm not sure how this can be avoided without fundamental changes in society with regards to competition and adversity.
- thoughtstheseus 5y agoPerhaps trolling the entire web is not useful today? I’d love a search engine where I can whitelist sites or take an existing whitelist from trusted curators.
- laurent92 5y agoIf the user requests a website, you could at least crawl on request, which would be an excuse to bypass the rules in robots.txt. It would be a loophole, let’s say.
- erhk 5y agoTrusted curators is a dangerous dependency
- dragonwriter 5y agoPlus, it scales less well than pure algorithmic search. This fight already happened, with a much smaller internet.
- shituonui 5y agoIt works really, really well for libraries. Research libraries (and research librarians) are phenomenally valuable. I've missed them any time I'm not at a university. Both curators and algorithms are valuable. This goes for finding books, for finding facts and figures, for finding clothes, for finding dishwashers, and for pretty much everything else. I love the fact that I have search engines and online shopping, but that shouldn't displace libraries and brick-and-mortar. Curation and the ability to talk to a person are complementary to the algorithmic approach.
- dragonwriter 5y ago> It works really, really well for libraries It scales extremely poorly. It works very well for situations where there are customers/sponsors are willing to spend lots of money for quality, because then the cost scaling doesn't matter as much; research libraries, Lexis/Nexus Westlaw, etc. all do this, but it's not cheap, and the cost scaling with the size of the corpus sucks compared to algorithmic search. It is among the approaches to internet search that lost to more purely algorithmic search, because it scales poorly in cost.
- indymike 5y agoI've used Gigablast off and on for a long time (I think I first discovered Gigablast in 2006 or so). Would be cool to have a registration service for legitimate spiders. I used to run a team that scraped jobs and delivered them (by fax, email, us mail as require by law) to local veteran's employment staffers for compliance. We were contracted by huge companies (at one point about 700 of the fortune 1000) to do so, and often our spiders would be blocked by the employer's IT department even though the HR team was paying us big bucks to do so.
- lloydatkinson 5y agoWith a slightly fresher coat of paint this could be very popular. For example, no grey background.
- mirker 5y agoIf you have customers, does that mean the incremental gain from an improved index costs too much to store? Or are you talking about computational costs?
- gbmatt 5y agoit's both storage and computational. they go hand in hand.
- afrcnc 5y agohow recent are your results? 1-2h? 1 day?
- gbmatt 5y agoit's continually spidering. just not at a high rate. actually, back in the day i had real time updates while google was doing the 'google dance'. that caused quite a stir in the web dev community because people could see their pages in the index being updated in real time whereas google took up to 30 days to do it.
- 1cvmask 5y agoWhy do you have a user account with a login?
- mrlinx 5y agoWhere did you read that google/alphabet owns part of Cloudflare?
- bloudermilk 5y agoAssuming OP is referring to Google Venture's participation in at least one of Cloudflare's rounds. https://www.crunchbase.com/funding_round/cloudflare-series-d--86e2df31 https://www.crunchbase.com/funding_round/cloudflare-series-d...
- deleted 5y ago[deleted]
- garaetjjte 5y ago>Gigablast has teamed up with Imperial Family Companies Associating with that crank (responsible for recent freenode drama) is very off-putting.
- hdjjhhvvhga 5y agoRegarding the gatekeeper problem: it's a wild guess but maybe if there was a way to involve users by organizing distributed scraping just for the sake of building a decent index, I'm sure many of them would help.
- gbmatt 5y agoyes, large proxy networks are potential solutions. but they cost money, and you are playing a cat and mouse game with turing tests, and some sites require a login. furthermore, people have tried to use these to spider linkedin (sometimes creating fake accounts to login) only to be sued by microsoft who swings the CFAA at them. so you start off with an intellectual desire to make a nice search engine and end up getting sidetracked into this pit of muck and having microsoft try to put you in jail. and, no, i'm not the one microsoft was suing.
- smt88 5y agoWhat if you allowed trusted contributors to "donate" their browsing to your index? AltaVista and Yahoo did that with browser plugins in the 90s.
- InfiniteRand 5y agoNot sure if you're looking for feedback, but the News search could use some work, I searched for "Ethiopia" and almost all of the articles were unrelated to Ethiopia except for the existence of some link somewhere on the page. Your general web search seems pretty good, although I've just given it a casual glance. I think your News search could be improved by just filtering the general search results for News-related content, since the "Ethiopia" content I get there is certainly Ethiopia-related. In any case, an interesting product, I'll try to keep an eye on it.
- ramboldio 5y agomaybe just add small webpages into your index, dont bother yo execute JS and dont download any images. The content quality will be higher and it's a lot cheaper.
- woutr_be 5y agoOut of curiosity, why would not executing JavaScript or not downloading images equal higher content quality?
- melony 5y agoWhat we need is a net neutrality doctrine on the server side. Bandwidth is hardly scarce outside of AWS's business model. Ban the crawler user-agent dominance by the big search engine players. "Good behaviour" should be enforced via rate limiting that equally applies to all crawlers, without exemption for certain big players.
- ColinHayhurst 5y agohttps://knuckleheads.club/ https://knuckleheads.club/
- sockaddr 5y agoNice. I'd pay 5-10$/mo for a search engine that didn't just funnel me into the revenue-extracting regions of the web like Google does.
- RhysU 5y agoA subscriber-supported search engine sounds cool to me. Any precedent?
- xtracto 5y agoCopernic ( https://copernic.com/ https://copernic.com/ ) had Copernic Agent Professional, a for-pay desktop application that had really good search features, a while ago . Not sure if they discontinued it.
- gompertz 5y agoWow blast from the past. I think I was using Copernic all the way back in 2003... Forgot all about them. Thanks!
- gianthockey495 5y agoYou'll like https://neeva.com/ https://neeva.com/
- fxtentacle 5y agoHow do they pay for it?
- yellow_postit 5y agoFrom the FAQ: > …Eventually, we plan to charge our members $4.95/month.
- samcrawford 5y agoKagi.com does this. In closed beta at the moment, but you can email and request access.
- DavidCole1 5y agoInteresting. I had some interests in building a search engine myself (for playing around ofcourse). I had read a blog post by Michael Nielson [1] which had sparked my interest. Do you have any written material about your architecture and stuff like that? Would love to read up. [1]: https://michaelnielsen.org/ddi/how-to-crawl-a-quarter-billion-webpages-in-40-hours/ https://michaelnielsen.org/ddi/how-to-crawl-a-quarter-billio...
- gbmatt 5y agothere's some stuff here : https://github.com/gigablast/open-source-search-engine https://github.com/gigablast/open-source-search-engine
- DavidCole1 5y agoThank you.
- entropie 5y agoHoly, thats a huge codebase. Github even shows no code/syntax hl for many cpp files because they are so big. I fiddled around and searched for some not so well known sites in germany and the results were surprisingly good. But it looks really... aged.
- kingcharles 5y agoHoly shit. Click on random .cpp file. Browser hangs. O_O
- Archelaos 5y agoI tried out four search words with your search engine, and I am not convinced that it is mainly the index size and not the algorithm that is to blame for bad search results. There are way too much high ranking false positives. Here is what I tried: a) "Berlin": 1. The movie festival "Berlinale" 2. The Wikipedia entry about Berlin 3. Something about a venue "Little Berlin", but the link resolves to an online gaming site from Singapure 4. "Visit Berlin", the official tourism site of Berlin 5. The hash tag "#Berlin" on Twitter 6. "1011 Now" a local news site for Lincoln, Nebraska 7. "Freie Universität Berlin" 8. Some random "Berlin" videos on Youtube 9. The Berlin Declaration of the Open Access Initiative 10. Some random "Berlin" entries on IMDb 11. A "Berlin" Nightclub from Chicago 12. Some random "Berlin" books on Amazon 13. The town of Berlin, Maryland 14. Some random "Berlin" entries on Facebook 15. The BMW Berlin Marathon b) "philosophy" 1. The Wikipedia entry about philosophy 2. "Skin Care, Fragrances, and Bath & Body Gifts" from philosophy.com 3. "Unconditional Love Shampoo, Bath & Shower Gel" from philosophy.com 4. Definition of Philosophy at Dictionary.com 5. The Stanford Encyclopedia of Philosophy 6. PhilPapers, an index and bibliography of philosophy 7. The University of Science and Philosophy, a rather insignificant institution that happens to use the domain philosophy.org 8. "What Can I Do With This Major?" section about philosophy 9. Pages on "philosophy" from "Psychology Today". I looked at the first and found it to be too short and eclectic to be useful. 10. The Department of philosophy of Tufts University c) "history" 1. Some random pages from history.com 2. "Watch Full Episodes of Your Favorite Shows" from history.com 3. Some random pages from history.org 4. "Battle of Bunker Hill begins" from history.com 5. Some random "History" pages from bbc.co.uk 6. Some random pages from historyplace.com 7. The hash tag "#history" on Twitter 8. The Missouri Historical Society (mohistory.com) 9. Some random pages from History Channel 10. Some random pages from the U.S. Census Bureau (www.census.gov/history/) d) "Caesar" 1. The Wikipedia entry about Caesar 2. Little Caesars Pizza 3. "CAESAR", a source for body measurement data. But the link is dead and resolves to SAE International, a professional association for engineering 4. The Caesar Stiftung, a neuroethology institute 5. Some random "Caesar" books on Amazon 6. Hotels and Casinos of a Caesars group 7. A very short bio of Julius Ceasar on livius.org 8. Texts on and from Caesar provided by a University of Chicago scholar 9. (Extremely short) articles related to Caesar from britannica.com 10. "Syria: Stories Behind Photos of Killed Detainees | Human Rights Watch". The photos were by an organization called the Caesar Files Group So what I can see are some high ranked false positives that are somehow using the search term, but not in its basic meaning (a3, a11, b2, b3, d2, d3, d4, d6) or not even that (a6). Some results are ranking prominently although they are of minor importance for the (general) search term (a9, a13, b7, b8 -- perhaps a15 and d10). Then there are the links to the usual suspects such as Wikipedia, Twitter, Amazon, etc. (a2, a5, a8, a10, a12, a14, b7, c5, d1, d5); I understand that Wikipedia articles are featuring prominently, but for the others I would rather go directly to eg. Amazon when I am interested in finding a book (or use a search term like "Caesar amazon" or "Caesar books"). Well, and then there are the search results that are not completely off, but either contain almost no information, at least compared to the corresponding Wikipedia article and its summary (b4, b9, d7, d9), or that are too specific for the general search term (c1, c2, c3, c4, c6, c9, c10). That leaves me with the following more or less high quality results (outside of the Wikipedia pages): a1, a4, a7, b5, b6, b10, and d8. The a15 and d10 results I could tolerate if there had been more high quality results in front of them; but as a fourth and second, respectively, good result they seem to me to be too prominent. Also in the case of "Berlin" a4 should have been more prominent than a1, and a7 is somewhat arbitrary, because Humbolt University and the Technical University of Berlin are likewise important; what is completely missing is the official Website of the city of Berlin (English version at www.berlin.de/en/). All in all, I would say that your ranking algorithm lacks semantic context. It seems the prominence of an entry is mainly determined by either just being from the big players like Twitter, Youtube, Amazon, Facebook, etc. or by the search term appearing in the domain name or the path of the resource, regardless of the quality of the content.
- easton 5y agoYou can be whitelisted so Cloudflare doesn't slow you down (or block you): https://support.cloudflare.com/hc/en-us/articles/360035387431-Frequently-asked-questions-about-Cloudflare-bot-products#h_5itGQRBabQ51RwT5cNJX8u https://support.cloudflare.com/hc/en-us/articles/36003538743...
- gbmatt 5y agoIt's not quite that easy. Have you ever tried it? See my post below. Basically, yes, I've done it, but i had to go through a lot and was lucky enough to even get them to listen to me. I just happened to know the right person to get me through. So, super lucky there. Furthermore, they have an AI that takes you off the whitelist if it sees your bot 'misbehave', whatever that is. So if you have a certain kind of bug in your spider, or your bot 'misbehaves', whatever that means is anyone's guess, then you're going to get kicked off the list. So then what? You have to try to get on the whitelist again? They have Bing and Google on some special short lists so those guys don't have to sweat all these hurdles. Lastly, their UI and documentation is heavily centered around Google and Bing, so upstart search engines aren't getting the same treatment.
- gbmatt 5y agoCloudflare is not the only gatekeeper, too. Keep that in mind. There's many others and, as an upstart search engine operator, it's quite overwhelming to have to deal with them all. Some of them have contempt for you when you approach them. I've had one gatekeeper actually list my bot as a bad actor in an example in some of their documentation. So, don't get me wrong, this is about gatekeepers in general, not just only Cloudflare and Cloudfront.
- Lhiw 5y agoI dunno if y'all realise this but I'd pay for a search engine that black holes CloudFlare and any other sites that think bots shouldn't read their sites.
- 5y ago
- fsflover 5y ago> 2) Hardware costs are too high. Which is why the next big search engine should be distributed: https://yacy.net https://yacy.net.
- wruza 5y agoNo way to test it right away, demo peer 502-es.
- fsflover 5y agoYou could search for other public-facing instances, e.g., http://sokrates.homeunix.net:6060 http://sokrates.homeunix.net:6060.
- fcantournet 5y ago"distributed" doesn't make things more hardware efficient... It literally always make them less efficient. If e.g : mastodon had the same number of users as Twitter it would use 10x the ressources for the same traffic.
- betwixthewires 5y agoSure, but it does spread the costs among users and makes them more manageable. One guy shouldering the cost of a search index is less viable than letting users shoulder the costs. Some charge customers as a solution to this, and that works, but then they need a minimum revenue to continue, or have to monetize with investors which usually means changing direction and goals. The other option, letting people host portions of the index, spreads the cost out, and the product gets about as good (best case scenario) as it's utility to people.
- ma2rten 5y agoIt's much more expensive now to build a large index (50B+ pages) Do you have a cost estimate? Also could you be more selective in indexing, e.g. by having users requests sites to be crawled.
- ampersandy 5y agoRequiring users to know what sites they want in advance somewhat defeats the purpose of a search engine, no?
- robbomacrae 5y agoNot at all. You only have to fail the first request. It is an approach I took with my own attempt at a search engine way back! In fact I know personally that there is at least one patent out there that suggests initial 1st time request users being asked to provide the appropriate response as an efficient way to teach systems for future users. Obviously failing first requests isn't ideal but for popular requests it quickly becomes insignificant. Wikipedia might (if they don't already) want to make a similar suggestion for users to contribute when finding a low content/missing page.
- lowwave 5y ago> Obviously failing first requests isn't ideal but for popular requests it quickly becomes insignificant. The first request can also be called asynchronously, and display a message to the user that it is 'processing....'.
- convolvatron 5y agosince sites are so desperate to be indexed, doesn't it seem better to put the onus on them to announce themselves? it would be great if dns registries publshed public keys .. maybe they do in newer schemes?
- ma2rten 5y agoThat works once your search engine is more widely used, but not a lot of sites are going to register with a niche search engines. Many users on the other hand really want a search engine like this and would be willing to invest some time.
- bullen 5y agoDo you have some sort of PageRank?
- SamBam 5y ago> You do a search on Gigablast and say, well, why didn't it get this result that Google got. And that's because the index isn't big enough I wionder how much this is true, and how much (despite all our rhetoric to the contrary) it's because we have actually come to expect Google's modern proprietary page ranking, which counts more than just inbound links but all sorts of other signals (freshness, relevance to our previous queries, etc.). We dislike the additional signals when it feels like Google is trying to second-guess our intentions, but we probably don't notice how well they work when they give us the result we expect in the first three links.
- JacobThreeThree 5y agoI think people also have an inflated recollection of how good Google actually was back in 2005. Back then Google was only going up against indexes and link-rings, not 2021 Google/Bing/DDG/etc.
- more_corn 5y agoI hate google now. Every time I use it by accident I’m reminded how infuriating it is. I know DuckDuckGo is just bing in a Halloween mask, but I’ll gladly use something that’s not awesome as long as it’s also not infuriating. I’d take 2005 google any day.
- pbhjpbhj 5y ago2005? There were loads of other search engines (SE), and many meta-SE: hotbot, dogpile, metacrawler, ... (IIRC), plenty more. There was also indexes, which Yahoo, AOL (remember them!) had but there was, what was it called, dmoz?, the open web directory. When Google started, being in the right web directory gave you a boost in SERPs as it was used as a domain trust indicator, and the categories were used for keywords. Of course it got gamed hard. Google was good, but I used it as an alt for maybe 6 months before it won over my main SE at the time. I've tried but can't remember what SE that was, Omni-something?? One of the main things Google had was all the extra operators like link: inurl:, etc., but they had Boolean logic operators too at one point I think.
- yumraj 5y ago> Cloudflare (owned in part by Google) Please elaborate. Is there a special relationship between Cloudflare and Google?
- spullara 5y agoGoogle Capital is an investor: https://www.forbes.com/sites/katevinton/2015/09/22/google-microsoft-qualcomm-and-baidu-announce-joint-investment-cloudflare/ https://www.forbes.com/sites/katevinton/2015/09/22/google-mi...
- yumraj 5y agoThat is not the same as being owned by Google.
- kragen 5y agoActually, being an investor in a company is the same as owning that company in part.
- vitus 5y agoEspecially since Cloudflare went public back in 2019, at which point any investors cashed out. - Sincerely, a Google employee who has nothing to do with the investment branch of the company
- yumraj 5y ago> at which point any investors cashed out. Well, actually that is also not true. At IPO preferred stocks convert to common but the investors can keep their ownership, they can but don't have to cash out or can only partially cash out. Investors can also keep board seats in many (or most?) cases.
- jefftk 5y agoI don't know anything about this particular case, but it's very common for VCs to cash out at IPO or not long after. VCs identify good investments among early stage companies; they don't want to keep their money tied up in investments outside of their specialty.
- Minor49er 5y agoI really love how the results organize multiple matching pages from the same domain. This is really cool.
- subsubzero 5y agocurious how you implemented the index, memory based or disk based? Either way you are right, HW costs are extremely expensive and you would need a lot of high RAM/high core count machines to return such a large index to the endusers in a low latency fashion.
- 1_player 5y agoIf you're serious about this, add a paid tier. Until it's free, I don't trust you will not ever sell my data to make bank.
- Nasrudith 5y agoWhy do people think a paid tier will prevent their data from being sold after pocketing it? Aside from that if they go bankrupt then it isn't theirs to not give away anymore for one.
- jermaustin1 5y agoYou are going to pay for a generalized web search when DDG/Google/Bing/etc are free?
- Closi 5y agoI would - the problem with those services is that they prioritise the results that generate the search engine the most money rather than give me the best results, and then indexes my searches to track and advertise to me throughout the web. A clear pricing transaction sounds much nicer to me. Should generate better results too.
- 1_player 5y agoYes. I use Brave Search and I hope they add a paid tier, which I think they have confirmed they'll add at a later date. If you don't pay, you are the product. Simple as that.
- duckmysick 5y ago> If you don't pay, you are the product. If not enough people pay, there's no product.
- 1_player 5y agoIf nobody pays, there's even less of it. Not sure what's your point.
- 1vuio0pswjnm7 5y agoI really like GigaBlast. I wrote a "meta" search utility for myself that can query multiple search engines from the command line.^1 It mixes the results into a simplified SERP ("metaSERP"), optimised for a text-only browser, with indicators to show which search engine each result came from. The key feature is that it allows for what I might call "continuation searches". Each metaSERPs contains timestamps in its page source to indicate when searches were executed, as well as preformatted HTTP paths. The next search can thus pick up where the previous one left off. Thus I can, if desired, build a maximum-sized metaSERP for each query. The reason I wrote this is because search engines (not GigaBlast) funded by ads are increasingly trying to keep users on page one, where the "top ads" are, and they want to keep the number of results small. That's one change from 2005 and earlier. With AltaVista I used to dig deep into SERPs and there was a feeling of comprehensiveness; leave no stone unturned. Google has gradually ruined the ability to perform this type of searching with their now secretive and obviously biased behind-the-scenes ranking procedures. Why is there no way to re-order results according to objective criteria, e.g., alphabetical; the user must accept the search engines' ordering, giving them the ability to "hide" results on pages the user will never view or simply not return them. That design is more favorable to advertising and less favorable to intellectual curiosity. Each metaSERP, OTOH, is a file and is saved in a search directory for future reference; I will often go back to previous queries. I can later add more results to a metaSERP if desired. I actually like that GigaBlast's results are different than other search engines. The variety of results I get from different sources arguably improves the quality of the metaSERP. And, of course, metSERPs can be sorted according to objective criteria. This is, AFAIK, a different way of searching. The "meta-search engines" of yesteryear did not do "continuations", probably because it was not necessary. Nor was there en expectation that user would want to save meta-searches to local files. Users were not trying to minimise their usage of a website, they were not trying to "un-google". Today's world of web search is different, IMO. There seems to be a belief that the operator of a search engine can guess what a user is searching for, that a user who sends a query is only searching for one specific thing, and that the website has an ad to match with that query. At least, those are the only searches that really matter for advertising purposes. Serendipitous discovery while perusing results is not contemplated in the design. By serendipitous discovery I do not mean sending a random query, e.g., adding an "I'm feeling lucky" button, which to me always seemed like a bad joke. The only downside so far is I ocassionally have to prune "one-off" searches that I do not want to save from the search directory. I am going to add an indicator at search time that a search is to be considered "ephemeral" and not meant to be saved. Periodically these ephemeral searches can then be pruned from the search directory automatically. 1. Of course this is not limited to web search engines. I also include various individual site search engines, e.g., Github.
- skyde 5y agowhat kind of index is Gigablast using? traditional inverted index like Lucene or something more esoteric? I know Google and Bing both use weird data-structure like BitFunnel https://www.microsoft.com/en-us/research/publication/bitfunnel-revisiting-signatures-search/ https://www.microsoft.com/en-us/research/publication/bitfunn...
- gbmatt 5y ago100% custom.
- xwdv 5y agoHow much cash do you need?
- betwixthewires 5y agoDude, I use your engine regularly, it is spectacular. The amount of work you put into this takes some dedication. I was curious if you ever intend to implement OpenSearch API so that we could use it as default in browser or embed it in applications? Also how can people contribute to help you maintain a larger index and/or keep the service going?
- zandorg 5y agoI wanted to add my site to Gigablast, but it said it would cost 25 cents. How is this a good thing?
- agencies 5y agoRe: crawling being too hard Have you contributed your crawl data to common crawl?
- S5yDyAk3XoQH5 5y ago<div id=content style=padding-left:40px;> </div id=box> lmao, hopefully the C code isn't nearly as bad as your html
- dang 5y agoPlease make your substantive points without snark or swipes. We ban accounts that do the latter, because it's poisonous to the culture we're trying to develop here. If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.
- S5yDyAk3XoQH5 5y agoProbably been here longer than you, so really irrelevant. Anyway every single page of his site has html errors, not pointing it out is more poisonous than doing so.
- dang 5y agoHN users need to follow the site guidelines regardless of how long they've been here or how strongly they feel about HTML errors.
- beamatronic 5y agoStoring information about the pages you can't index, is also useful
- Thespian2 5y agoI hadn't used gigablast before, but a quick test had it find some very old, obscure stuff, as the top hit. Well done. However, the link on the front page to explain privacy.sh comes up with "Not Private" in Chrome. The root Cisco Umbrella CA cert isn't trusted. Oops.
- kazinator 5y agoIs there a way to get the results to be formatted for desktop? It looks like the layout is hard-coded for a mobile browser, in portrait mode.
- deleted 5y ago[deleted]
- znpy 5y agoI just looked myself up in your search engine and I can confirm that it finds stuff old enough that google wouldn't find them (eg: and old patch I submitted on gnu savannah). I tried looking up a game I'm interested in and the second results cluster from your search engine is a reddit thread about linux support for that game... I love this. Great job!
- kingcharles 5y agoI tried searching for an answer, but how do you get a site added to your directory? Who maintains it? Directories are a real PITA to maintain with any quality.
- tasogare 5y agoI just tried it and the UI is kinda old and not mobile friendly but the English results I got were satisfying. Not the case for French though. I'll try again in the future, diversity in this landscape is important.
- mandeepj 5y ago> Hardware costs are too high I want to say - you don’t know what are talking about. But, it’ll be rude. Hardware is much cheaper and powerful now compared to 2005.
- gbmatt 5y agothe complexity of the search algorithm has also increased substantially since 2005 And, in 2005, a billion page index was pretty big. Now it's closer to 100 billion.
- systemBuilder 5y agoThere were ~60B pages on Facebook in 2015 I think your numbers are outdated. - Google search SRE
- BbzzbB 5y agoYou've said it and it is rude, what's the point of that first sentence except to spite him? I'm sure he's well aware of the price per capability trend since 2005, you don't code a search engine without knowing that. Could be the costs of servicing his free users and/or maintaining an ever-growing database/index that is costly - in spite of cheaper hardware on a relative basis.
- kf6nux 5y agoWhat are your sources for hostnames to crawl? I looked into it a long time ago and seem to remember there was a way to get access to registration records, but I imagine combining that with HTTP certificate transparency records would significantly increase your hostname list. Anything else?
- cphoover 5y agowhat heuristics or AI is being used for blocking your spider? If your spider appears human or organuc it will not be blocked right? Is this an issue of rate limiting, or request cadence? could you add randomness to the intervals in which you request the page? Is it more complicated? do they use other signals to ascertain if you are a script or not like checking data from the browser (similar signals to the kind of things browser fingerprinting uses... e.g. screen res, user agent, cache availability, etc...) would it be possible for the browser to spoof this information? I imagine rate limiting the IP address is the major issue... but could you not bounce the request through a proxy network? I've tried this with the TOR network before when writing web scrapers and had mixed success... seems like Google knows when a request is being made through Tor. Perhaps you could use the users of your search engine as a proxy network through which to bounce the request for the scrape/indexing... This way the requests would look like they were coming from any of your users instead of one spiders ip address...Im not sure how cloudflare or any other reverse proxy could determine that thise requests were organic or not... id be ok with contributing to a distributed search service so long as my cpu was not making requests to illegal content, and there were constraints put on the resource usage of my machine. Sorry if this came off as all over the place, I do not know too much about the offense vs defense of scraping. These are just some thoughts...
- rnotaro 5y ago> I've tried this with the TOR network before when writing web scrapers and had mixed success... seems like Google knows when a request is being made through Tor. That's because all the TOR entry/exit nodes and relays IPs addresses are public [1]. [1] https://metrics.torproject.org/rs.html#toprelays https://metrics.torproject.org/rs.html#toprelays
- varispeed 5y agoMake sure to file complaints to any competition market authority you have in your country.
- alok-g 5y agoOh my god! This works so much better than every Internet search engine I have tried.
- readonthegoapp 5y agodid you ever try to raise funds? why/not? not accusing, just curious. did you ever think, let me just focus on Italy-relevant results? or job search only? or some slice of search.