7 ms·
Guy running a Google rival from his laundry room
- amelius 1y agohttps://archive.ph/HA7y4 https://archive.ph/HA7y4
- authnopuz 1y agohttps://archive.is/HA7y4 https://archive.is/HA7y4
- cheema33 1y agoI tried the search site at https://searcha.page/ https://searcha.page/ by searching for something random and got the following message: "An error has occurred building the search results."
- authnopuz 1y agohug of death? I fear the temperature will get very high in his laundry room
- DannyBee 1y agoI'm sure it depends on how much laundry he is doing - his dryer is probably heated entirely by servers. He can then exhaust the remaining server heat through the dryer vent stack.
- debo_ 1y agoKeep going. I love dry humor.
- deleted 1y ago[deleted]
- egberts1 1y agoIts dryer sheets soften the soul.
- ArekDymalski 1y agoUntill the exhaust starts "Feeling leaky" I guess.
- robofanatic 1y agoMight not even need a dryer :-)
- ape4 1y agoChange it to a sauna?
- doublerabbit 1y agoI thought of this a whole ago when I was a Datacentre monkey. In the winter it was pleasant to walk down the hot aisles. However the exhausted hot air never had the same feel of a sauna. It left the air stale and dry.
- HelloUsername 1y agoYup; same at https://seek.ninja/s?q=beatles https://seek.ninja/s?q=beatles
- eschulz 1y agoBefore this happened to me, my first search returned an impressive SERP.
- chiefsearchaco 1y agoYep, my usage increased 20x week over week. It was actually the context expansion that was my bottleneck, not the search itself. My usage graph looks almost vertical. Not sure if this counts as a good week or a bad week.
- evanjrowley 1y agoSearch websites by Ryan Pearce: - SearchaPage - Web Search Engine https://searcha.page/ https://searcha.page/ - Seek Ninja - Stealthy Search Engine https://seek.ninja/ https://seek.ninja/
- 317070 1y agohttps://searcha.page/s?q=blog https://searcha.page/s?q=blog https://seek.ninja/s?q=blog https://seek.ninja/s?q=blog Both of them are erroring out right now?
- chiefsearchaco 1y agoYep, it was load. Usage increased 20x week over week, especially today. I think I failed my trial by fire. Got a good plan for scaling capacity and better UX for when its under strain.
- kitd 1y agoWere you trying them via Chrome, by any chance? ;)
- jslakro 1y agofirefox here and it's not working
- thm 1y agoI'm running one for news https://mozberg.com https://mozberg.com - not in my basement though.
- cosmicgadget 1y agoWhere is it?
- BLKNSLVR 1y agoGreat innovation plus cloud-skeptic self-hosting. There should be much much more of this!
- the_real_cher 1y agoI always wondered why someone couldn't do this. Google was invented many years ago by two guys in a dorm room and since then there's been so many white papers and advancements in the public sphere and the actual underlying problem has not changed that much, that it seems like it could be done by a small group or independent person.
- OutOfHere 1y agoThe actual underlying problem has changed altogether. Pagerank is easily gamed by SEO. Search candidates and rankings now require assessment by LLM. Moreover, as a default, users want the results intelligently synthesized into a text response with references rather than as raw results. Crawling too requires innovative approaches to bypass server filters. I doubt any independent person can afford to run a vector database or LLMs at immense scale.
- kcbanner 1y ago> users want the results intelligently synthesized into a text response with references rather than as raw results. The reason I pay for Kagi is that I specifically don't want this to occur.
- Workaccount2 1y agoHow much do you technologically relate to the average person on the street though? Every person I have seen (outside the tiny tech bubble) google something has just read the AI overview without skipping a beat.
- throwmeaway222 1y agoAt this point the web is also so centralized you only need 3 bookmarks these days (your news, youtube and Amazon) A search is just learning what you don't know and AI does a better job than search has ever done for me - and I'm in tech.
- 1y ago
- ourguile 1y agoI greatly prefer Kagi https://help.kagi.com/kagi/company/ https://help.kagi.com/kagi/company/ but it's very nice to see more competition in this space in general.
- tmdetect 1y agoKagi is a polished product. This is drying someones laundry.
- the_third_wave 1y ago[flagged]
- hamdingers 1y agoHave you considered it's a good product that causes its users to become advocates?
- foobarian 1y agoCould also be a form of effort justification. [1] [1] https://en.wikipedia.org/wiki/Effort_justification https://en.wikipedia.org/wiki/Effort_justification
- glenstein 1y agoTIL about effort justification! I think signing up for Kagi is not particularly effort-intensive however.
- tolerance 1y ago> The effect is most likely to occur when there are no obvious reasons for performing the task. Because expending effort to perform a useless or unenjoyable task, or experiencing unpleasant consequences in doing so, is cognitively inconsistent (see cognitive dissonance), people are assumed to shift their evaluations of the task in a positive direction to restore consistency. I’m not following you. https://dictionary.apa.org/effort-justification https://dictionary.apa.org/effort-justification
- vlucas 1y ago> “I think it’s definitely lowered the barrier,” Lin says of the LLM’s role in enabling DIY search engines. “To me, it seems like the only barrier to actually competing with Google, creating an alternate search engine, is not so much the technology, it’s mostly the market forces.” Oh sweet summer child
- luizfelberti 1y agoI was trying to do this in 2023! The hardest part about building a search engine is not the actual searching though, it is (like others here have pointed out), building your index and crawling the (extremely adversarial) internet, especially when you're running the thing from a single server in your own home without fancy rotating IPs. I hope this guy succeeds and becomes another reference in the community like the marginalia dude. This makes me want to give my project another go...
- ge96 1y agoThe IP thing is interesting, I was trying to make this CSGO bot one time to scrape steam's prices and there are proxy services out there you rent, tried at least one and it was blocked by steam. So I wonder if people buy real IPs.
- kccqzy 1y agoYeah people buy residential IPs on the black market. They are essentially infected home PCs and botnets.
- Bratmon 1y agoNot just the black market anymore! https://www.proxyrack.com/residential-proxies/ https://www.proxyrack.com/residential-proxies/
- immibis 1y agoyou can get paid about $0.10/GB in cryptocurrency (at a few GB per month) to run one on your PC. Apparently they also just buy actual connections sometimes. It's not even unethical - it's just two groups of equally bad businesspeople trying to spend money to block the other one.
- typpilol 1y agoI've heard a few horror stories... Since the people using residential proxies aren't necessarily always good people
- Oarch 1y agoI'm sure there's a money laundering joke in here somewhere
- ofrzeta 1y ago"The beefy CPU running this setup, a 32-core AMD EPYC 7532, underlines just how fast technology moves. At the time of its release in 2020, the processor alone would have cost more than $3,000. It can now be had on eBay for less than $200" why do I never get deals like that when I am shopping for the homelab on eBay?
- progval 1y agoYou need to spend a lot of time looking through badly labeled offers, and be willing to buy from sellers with no reputation.
- deleted 1y ago[deleted]
- robrtsql 1y agoI searched "AMD EPYC 7532" and there are a ton of listings for $150-$200. Are you just regretful that it wasn't like this when you were shopping parts for your homelab?
- throwawayffffas 1y agoI got a 7551p plus motherboard and ram for about 600 bucks from China this January. I may have overpaid but it works great, and gets the job done.
- _fat_santa 1y agoNot for a CPU but earlier this year I bought a Thinkpad workstation off eBay for $500. It's a machine from 2020 and when it was new cost $5,700. I see this for pretty much all hardware out on eBay, just go back 5 years and watch the price fall 10x.
- saalweachter 1y agoHas eBay fixed their "and then they ship you a box of rocks" problem? I feel like there was a five year span where everyone I talked to said buying or selling electronics on eBay was a nightmare, so I'm a little curious if I need to re-evaluate my priors.
- iam_saurabh 1y agoI love stories like this—tech history is full of scrappy beginnings. Even if this project doesn’t succeed, it reminds us that giant companies aren’t unshakable.
- renegat0x0 1y agoWell, I created my own domain index. I have not crawled every page inside domains, but it is not my goal. I have 1542766 domains. Might not be much, but it is an honest work. It is available as a github repo, so anybody that wants to start crawling has some initial data to kick off. Links https://github.com/rumca-js/Internet-Places-Database https://github.com/rumca-js/Internet-Places-Database
- hobs 1y agoCant you just request the ICANN’s zone files and have the canonical list of the day?
- egberts1 1y agoAvoiding GIGO (Garbage In, Garbage Out). This is why we have computer-variants of Library Science and Archeology, Forensic Science and a bunch of other advanced knowledge (not AI, mind you).
- hobs 1y agoI don't see how this applies as its aggregating a bunch of stuff from random crawlers - if you want to crawl a list of actual domains that's generally considered the list of things that could resolve, so seems like a good starting place.
- egberts1 1y agoSmashing stuff together by pure probablistic word association like AI do today is really, tsk tsk. That's why it is important to clearly define words like love, oh wait.
- renegat0x0 1y agoAny link list, or domain list is not worth much without any rating, or meta. I lead a hobby project, and I am not expert, so I provide ratings based on what kind of data pages provide (title, social, description), and my own manual voting system. It is not ideal, but it is something. Also I provide tags, so it is easily known what the domain provides, or domains can be filtered by tags. I know that you cannot count and visit every domain, so the list will never be finished, but I am happy with the results.
- elite_barnacle 1y agoReminds me of this XKCD: https://xkcd.com/908/ https://xkcd.com/908/
- tolerance 1y agoThe great thing about this is that with the decentralization/recentralization of the Web, it may become easier for certain people to roll their own search engines for their respective communities and crawl/index pages only according to their shared tastes. The bad thing about this is...read above.
- HardCodedBias 1y agoI know that Google engineers have a cushy life but I actually find it unlikely that a guy, who isn't attempting some radical new type of search (like pagerank back in the day) can hope to compete with the orgs in Google who support search. Again, those orgs are likely too comfortable and less productive than people would like, but we're talking about many-many thousands and depending upon how you define "the work" of search upwards of 10k. I didn't see any new secret sauce in the article and Google is has said that since 2015 (?) Google Brain has been involved in search. This is not to say that Google couldn't be dislodged by search via LLM or similar, that is "new" research.
- freeopinion 1y agoIf you wrote that 100 people could outwork one person, I'd nod my head. If you wrote that 10k people could outwork 1k people, I'd shrug. If you tell me that 100 people can combine to tie my shoe faster than I can, I'd question that. Building a state-of-the-art search engine is not shoelaces. But upwards of 10k workers is not impressive in the right direction. One person starting out with anything at all can quickly grow into one person with one or two really innovative ideas. One or two good ideas can catch fire pretty quickly. Don't be too dismissive.
- mooiedingen 1y agoNothing new as it has been done before, the concept is simple enough: step 1: indexer, solr/lucene Step 2: crawler of which there are several foss, build one yourself? or you just run yacy which is a combo of the above, hook combine with an oldschool searx instance and you will be granted the title as seeker by the spirit of Fravia+ who was elder of the searchlores!!! Not only will you filter crap made by machine learning models, but thou shall find what thou seek! I refuse to call a 16 line long for loop triggering in memory loaded tokenized data where data can be anything from a scientific paper hallucinated by a chatbot to a message between two lovers anything intelligent for it is not intelligence but a blob of tokenized fcking data in memory getting triggered for an output by a derp with a 16 line long for loop!!!
- rurban 1y agoxapian is easier and faster. No Java memory eater. I've once built a good company wide search engine with custom crawlers, and result hooks, eg to crazy SAP or other ticket systems. Gmane was also legendary.
- p3rls 1y agoi've been thinking that google could use its own AI to evaluate URLs instead of relying on pagerank and backlinks which are almost completely valueless as a signal in 2025. in my niche there's more slop than ever being produced daily and it's all hitting rank 1. it's tragic what google is doing to the internet.
- ytrt54e 1y agoCrashed? The curse of Hacker News!
- lucb1e 1y agoIt claims I reached the article limit. The last time I saw a fastcompany link must have been a decade ago! I was nostalgically looking forward to read another article of theirs. Alas... https://archive.is/HA7y4 https://archive.is/HA7y4 Some bits and pieces: > his new search engine, the robust Search-a-Page <https://searcha.page https://searcha.page>, which has a privacy-focused variant called Seek Ninja <https://seek.ninja https://seek.ninja> > The secret to making it all happen? Large language models. “What I’m doing is actually very traditional search,” Pearce says. “It’s what Google did probably 20 years ago, except the only tweak is that I do use AI to do keyword expansion and assist with the context understanding > Fellow ambitious hobbyist Wilson Lin, who on his personal blog <https://blog.wilsonl.in/search-engine/ https://blog.wilsonl.in/search-engine/> recently described his efforts to create a search engine of his own, took the opposite approach from Pearce. > And then there’s the concept of doing a small-site search, along the lines of the noncommercial search engine Marginalia <https://marginalia-search.com https://marginalia-search.com>, which favors small sites over Big Tech And the obvious answer to the title: "Why the laundry room? Two reasons: Heat and noise." It runs on a a 32-core AMD EPYC 7532, half a terabyte of RAM, and "all in, cost $5,000, with about $3,000 of that going toward storage"
- udkl 1y agoI absolutely devoured Wilson Lins articles recently .. they are very high quality and informative for any amateur interested in search engines and LLMs! - https://blog.wilsonl.in/search-engine/ https://blog.wilsonl.in/search-engine/
- wvenable 1y agoReader mode in Firefox (plus sometimes a page refresh) gets me past most paywalls -- including this article.
- phendrenad2 1y agoThis is a cool project, and I hope he has fun with it. I've daydreamed about how I'd create my own search engine so, so many times. But I always run into an impassable wall: The internet now isn't at all the same as the internet in 1999. Discovery isn't really that useful. If you find someone's self-hosted blog about dinosaurs, it probably hasn't been updated since 2004, all the links and images are broken, and it's just thoroughly upstaged by Wikipedia and the Smithsonian. Sure, it's fun to find these quirky sites, but they aren't as valuable as they once were. We've basically come full circle to the AOL model, where there are "hubs" of content that cater to specific categories. YouTube has ALL the long-form essays. Tiktok has ALL the humorous videos. Medium has ALL the opinion pieces. Reddit has ALL the flame wars. Mayo Clinic has ALL the drug side-effects. Amazon has ALL the shopping. Ebay has ALL the collectables. None of these big companies want nasty little web crawlers poking and prodding their site. But they accept Google crawlers, because Google brings them users. Are they going to be that friendly to your crawler? Of course, I still dream. Maybe a hub-based internet needs a hub-aware search engine?
- OJFord 1y ago'Google rival' is quite a stretch, surely 'search engine' is not just more accurate, but clearer too with all that Google does today, as if that's new.
- _joel 1y agoThe photo of the power socket right next to the sink looks safe
- throwway120385 1y agoLooks like a GFCI. Should be fine.
- chiefsearchaco 1y agoI'm planning on running a cord through my wall, I just keep putting it off :D
- zrobotics 1y agoAbsolutely don't run an extension cord through a wall, that's only slightly less of a fire hazard than storing a gasoline can on top of the server. Extension cords are normally very derated, expecting occasional use and ample cooking not being inside a wall. Better to keep it as-is, or have a 20A dedicated circuit run.
- lxe 1y agoThis is a cool hobby project, but why is this notable? Why a FastCompany article? I'm trying to figure out anything that sets this apart from thousands of other little hobby search projects. I understand companies like Perplexity or Brave or DuckDuckGo "rivialing Google", but building a hobby index and crawler is nice, and worthy of a "Show HN: "... but an actual media article?
- gowld 1y agoIt's only notable as a clickbait narrative for ignorant readers -- FastCompany's target market
- risico 1y agoOne of my dream projects as well, sadly it feels a lot harder to crawl the internet these days, as others have said around here as well. What are some good practices these days to ensure a good crawl/scrape? Invest in proxies, preferably residential?
- mips_avatar 1y agoIt’s amazing what indie builders are doing with vector search, but I’m not sure how long it will last. Pure vector search works well today largely because no one is seriously trying to game it yet. Once adversaries start targeting it like they do SEO, we could see the same problems. You can already glimpse the risk in Pinterest, where roughly half the results for many queries are AI slop - since their primary search is image vectors
- freedomben 1y ago> Why the laundry room? Two reasons: Heat and noise. Pearce’s server was initially in his bedroom, but the machine was so hot, it actually made it too uncomfortable to sleep. This is a rite of passage and a badge of honor for homelabbers/tinkerers/hackers to discover for themselves IMHO. If you haven't tried it, you should. The heat is bad enough to warrant moving it, but add the noise too, sprinkle in a few nights of bad sleep, and it becomes an effective form of torture :-D Just don't decide to move it to a closet unless you also install some fans in there. I ended up finding a cozy spot under the staircase which worked quite well
- jp191919 1y agoI wonder what his ISP is, and what speed he has to subscribe to...
- yelling_cat 1y agoI love this project, but that server setup makes me hope that Ryan doesn't live in earthquake country.
- chiefsearchaco 1y agoWell I can't respond to everyone - I am the one running the search engine. And yes, it did crash today from load. Usage increased 20x this week vs last and I was totally unprepared. I don't know if that counts as a good launch or a bad one. For some reason in my head I imagined usage would be some slow steady ramp. Thank you for those who tried it, and I'm sorry if you were one of the people it didn't perform for. As far as load goes this was the first day it truly had a "trial by fire".
- rurban 1y agoJust switched to Search Ninja as my default search engine on my Android firefox. No tracking, faster, better than duckduckgo. Now I'm just looking how to get search suggestions enabled.