Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
gbmatt
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
15 ms
·
1.
▲
by
gbmatt
3y ago
Only Big Tech (Microsoft,Google,Facebook) can crawl the web at scale because they own the major content companies and they severly throttle the competition's crawlers, and sometimes outright block them. I'm not saying it's im
2.
▲
by
gbmatt
4y ago
We are just robots in a human simulator, reliving our creation.
3.
▲
by
gbmatt
4y ago
Q: how might an AI algorithm be modified in order to return citations with its response? A: There are several ways in which an AI algorithm could be modified to return citations with its responses. Here are a few possibilities: One ap
4.
▲
by
gbmatt
5y ago
I just posted this same comment on the ddg story, but I'm going to post it here as well. Google forced my search engine (gigablast) basically out of business. I had ixquick.com as a big client at one time; I was providing them with sea
5.
▲
by
gbmatt
5y ago
Yeah, Google forced my search engine basically out of business. I had ixquick.com as a big client at one time; I was providing them with search results from my custom web search engine. Then their CEO called me one day and told me he was ca
6.
▲
by
gbmatt
5y ago
everyone needs equal access to public data. right now only big tech can download the many web pages (without thottling or being ip banned) on linkedin (microsoft), youtube (google), facebook, github (microsoft) and billions of more pages. t
7.
▲
by
gbmatt
5y ago
the complexity of the search algorithm has also increased substantially since 2005 And, in 2005, a billion page index was pretty big. Now it's closer to 100 billion.
8.
▲
by
gbmatt
5y ago
thanks ben, you are too kind.
9.
▲
by
gbmatt
5y ago
the javascript is run by your browser, so you can fully audit it.
10.
▲
by
gbmatt
5y ago
hey thanks for the recognition, people. :) finally, all my problems are solved. this comment is here for hacker news karma points.
11.
▲
by
gbmatt
5y ago
I'd argue that a level playing field and more competition in the search space is a good thing.
12.
▲
by
gbmatt
5y ago
100% custom.
13.
▲
by
gbmatt
5y ago
that's tripped out. where did you hear about that?
14.
▲
by
gbmatt
5y ago
I'll admit I had not been working on the quality of single term queries as much as I should have lately. However, especially for such simple queries, having a database of link text (inbound hyperlinks and the associated hypertest) is v
15.
▲
by
gbmatt
5y ago
Cloudflare is not the only gatekeeper, too. Keep that in mind. There's many others and, as an upstart search engine operator, it's quite overwhelming to have to deal with them all. Some of them have contempt for you when you appro
16.
▲
by
gbmatt
5y ago
It's not quite that easy. Have you ever tried it? See my post below. Basically, yes, I've done it, but i had to go through a lot and was lucky enough to even get them to listen to me. I just happened to know the right person to ge
17.
▲
by
gbmatt
5y ago
there's some stuff here : https://github.com/gigablast/open-source-search-engine
18.
▲
by
gbmatt
5y ago
brave 'falls back' to bing. which in my experience is most of the time. in fact, out of all the queries i did a while back, they all seemed to come directly from bing. is there a way to disable the reliance on bing and get pure &#
19.
▲
by
gbmatt
5y ago
yes, large proxy networks are potential solutions. but they cost money, and you are playing a cat and mouse game with turing tests, and some sites require a login. furthermore, people have tried to use these to spider linkedin (sometimes cr
20.
▲
by
gbmatt
5y ago
it's both storage and computational. they go hand in hand.
21.
▲
by
gbmatt
5y ago
both ddg and brave are bing (microsoft) in disguise.
22.
▲
by
gbmatt
5y ago
it's continually spidering. just not at a high rate. actually, back in the day i had real time updates while google was doing the 'google dance'. that caused quite a stir in the web dev community because people could see thei
23.
▲
by
gbmatt
5y ago
I've had extensively dealing with Cloudflare. They have a complex whitelisting system that is difficult to get on, and they also have an 'AI' system that determines if you should be kicked off that whitelist for whatever reas
24.
▲
by
gbmatt
5y ago
it should be. there should be some sort of 'bots rights' to level the playing field. perhaps this is something the FTC can look into. but, as it is right now big tech continues to keep their iron grip on the web and i don't s
25.
▲
by
gbmatt
5y ago
and i work with rasengan on private.sh so yes there's some issue there. one of the back end servers is returning a max capacity error of sorts... we are checking into it.
26.
▲
by
gbmatt
5y ago
Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it reall
27.
▲
by
gbmatt
5y ago
A project I'm involved with https://private.sh/ has search privacy that is more verifiable than ddg. It uses client-side cryptography in the javascript, and routes your request through an anonymizing proxy, similar to
28.
▲
How Big Tech Controls Web Search
(gigablast.com)
2 points
by
gbmatt
5y ago
|
0 comments
29.
▲
How to Fix Web Search
(gigablast.com)
3 points
by
gbmatt
5y ago
|
0 comments
30.
▲
Ask HN: How can I make money with my Web Search Engine?
(gigablast.com)
6 points
by
gbmatt
5y ago
|
2 comments
More ›