7 ms·
I built a 500k-domain search engine for makers in a weekend for $10
- nxndjdkdksmsb 2mo ago[dead]
- dreamforever 2mo agoCheck out my latest project! You can fork it, tweak the policy manually or with AI, run the system and watch the data come in! It's engineered to keep a low data footprint, so 500k domains fits into 1GB on disk. If you have local models it's free! You just might not get the best throughput depending on your GPU. My production data is not exposed anywhere yet, and I may never expose it. The point is for you to fork and make your own policy, and thus your own personal search engine! The article covers basic analysis on my data, so it's worth a read if you're interested! A deeper analysis may arrive with V2 if I ever do it
- hns86vq0nb 2mo ago[dead]
- iFire 2mo ago[dead]
- iFire 2mo agoHere's my impressions of your algorithm: 1. read each site 2. rent a 4090 with https://vast.ai https://vast.ai to run vllm 3. let llm model invent its own category and tag names freely 4. save 1KB of metadata each a. a small local language model that reads each one and writes a name, two or three sentences, a category, and a handful of tags. 5. `code is going up as open source` soon (TM)
- jeroenhd 2mo agoThe technical details are on another page: https://alexmorleyfinch.github.io/marlin/history/v1/article/technical.html https://alexmorleyfinch.github.io/marlin/history/v1/article/... Your impressions seem about right, but there are a few control steps it seems.
- moffkalast 2mo agoThey really needn't have specified "in a weekend" cause yeah we can tell. Since when has low effort become a selling point anyhow?
- smokel 2mo agoI typically interpret it as an excuse, not a selling point.
- dreamforever 2mo agoit was by no means low effort
- knollimar 2mo agoIsn't 60 hours by definition low effort? Setting an LLMs effort value to high doesn't count
- RadiozRadioz 2mo agoIt's an interesting new type of brag. It roughly means: "I want you to know how advanced I am at using LLMs, and how AI-first I am. So here is how little time I spent, to prove that I am using LLMs as much as possible, demonstrating that I am ahead of the curve on this new trend"
- iambenm 2mo agoCode appears to already be up: https://github.com/alexmorleyfinch/marlin https://github.com/alexmorleyfinch/marlin
- johntash 2mo agodon't forget: 6. let llm write a blog post about this conversation
- headz 2mo agoTS;DR: Too Sloppy; Didn't Read.
- juleiie 2mo ago[flagged]
- abc3354 2mo agoIn the expression "AI slop", "AI" is about the tools, "slop" is about the content
- uean 2mo agoI'm not going to spend an hour trying to distill the AI-slop to find out what potential golden nugget may lie in there. It's impossible to judge the content if it's buried under a landfill. "If you won't take the time to write it, I won't take the time to read it."
- juleiie 2mo ago[flagged]
- uean 2mo agoThanks for this. My point is proven.
- juleiie 2mo ago[flagged]
- flyingcoder 2mo agoHow did web3 and crypto work out for you buddy?
- jorisw 2mo ago
- dewey 2mo agoI think Kagi Small Web filter would give you very similar results.
- dreamforever 2mo agoI'll check them out!
- deleted 2mo ago[deleted]
- nonewideas 2mo ago[dead]
- marginalia_nu 2mo agoInteresting project. Website discovery is indeed in a pretty dire spot, definitely a space that needs innovation. An auto-labeled website directory isn't that silly of an idea. I have a 400 GB sqlite database with samples of rendered root document DOMs I use for ad detection in Marginalia Search I've been meaning to explore similar ideas using.
- jeromechoo 2mo agoIt feels like we've hit a point where search engines can become what "todo list apps" were for devs 10 years ago. What a homebrewed solution lacks in coverage it excels in indexing and serving a small slice of the internet really really well.
- marginalia_nu 2mo agoTo be fair they are a supremely interesting problem to hack away at, and one that will meet you where you are. Almost anyone can put together a basic search engine in a few thousand lines of code, it's just not very hard to make a program that will index a few million documents better than Confluence. Then, between that first ansatz and a working scalable internet search engine, you have a pile of interesting problems touching every aspect of computer science and computer hardware and networking, enough so that hundreds of people will have gotten PhDs in narrow sub-problems of those problems you'll be facing. It's great because you can just tackle the stuff you feel comfortable approaching and leave the rest for later.
- alightsoul 2mo agoI have been wanting do do this. The biggest source of domains is certificate transparency logs. Also ICANN zone files. According to some scientific papers these cover 88% of all registered domains. You could crawl dns for CNAME records with all ipv4 IPs by distributing requests across dozens of DNS servers, the internet archive or the common crawl but doing it for the internet archive is a dick move without giving them money There's about 200 million active domains currently. That's about 66% of all businesses worldwide of which there are around 300 million. Around 100 to 150 million have active webpages
- orliesaurus 2mo agoLike a personal Google? How do you bypass all the captcha, ip bans, cloudflare turnstile antibot stuff etc?
- voidUpdate 2mo agoThey don't: "skips the model entirely if the page is empty, parked, or a bot-challenge wall"
- sandeepkd 2mo agoThats the fun part, the user just went with happy path. Javascript, captchas, cloudflare protected content did not made to the catalogue. This sort of use case exists in LLM training data a lot which makes it easier. The data gathered by the user is not really practically useful cause there are way too many gotchas when it comes to web scraping and building a catalogue (source: I have done scraping for a particular domain data and had to do at least 10+ iterations to get it >90 right)
- dreamforever 2mo agodespite this limitation, there is still some good stuff out there, and with the priority steering, you can focus compute on what you actually want, fast and cheap.
- eggbrain 2mo agoThis is actually where I see software going in the short term -- cloud moving to local. A few years ago, if you wanted translation, you'd use Google Translate. If you wanted to search the web, you'd use Google search. But for a few gigabytes, you can now install nllb-200-distilled-600M, and get translations for almost any language locally. You can have your computer crawl the web, create abstracts and categorizations for websites, and build search exactly as you want it. The main limiter now is hard drive space (and to an extent, local compute) -- but right now it feels like the 70s again where the terminal into a remote server turned into building applications locally.
- dylan604 2mo agoThe number of times we've gone from cloud/server access via terminal to local compute back and forth is something that always makes me laugh a bit.
- an0malous 2mo agoIt's more of a pipeline than a back-and-forth. New abilities happen in the cloud first because they require specialized, higher capacity resources and then move towards being local as the resource usage gets optimized.
- dylan604 2mo agoOutside insane GPU appliances, I really see cloud based things as a lock-in to subscription based offerings. The fallacy of always using the latest code isn't worth never ending subscription fees. I'll install it locally. I'll keep my data locally. I don't need to wait for things to upload/download. If there's an update that feels worthy, I'll purchased and install locally.
- pimlottc 2mo agoSometimes I think people forget how capable computers are. 500k is not much. You can just slap that in a Lucene instance. This is a solved problem.
- marginalia_nu 2mo agoApproaching search by just tossing the data in Lucene is how you end up with Confluence's search box though.
- elorant 2mo agoDomains are way more than just 40M though.
- BaudouinVH 2mo agoFrom what I understand the aim was not to collect all the domains on the web but focus on personal website, etc. and avoid corporate web sites.
- fg137 2mo agoSorry I have a lot of trouble understanding what this is useful for. Like, I am never going to replace it with Google, DuckDuckGo, ChatGPT or even Bing.
- prepend 2mo agoI was wondering the same thing. I’ve wished for just a big blob of the web to grep and regex through, but I don’t think this is that much easier than using duckduckgo or even google.
- deleted 2mo ago[deleted]
- Xyra 1mo agoI've been working on grep.it.com. can currently grep over 200 TB of the internet
- prepend 1mo agoThis is really neat. How much does it cost to use? I see that I can start for free, but couldn’t find anything in your docs about usage fees.
- Xyra 1mo agothanks. I'm still sorting out pricing. I don't have experience with B2B and am interested in empowering out-of-distribution people, so will likely maintain some free tier + congestion-based auction pricing for non-commercial use.
- dreamforever 2mo agoIt's not for that, sorry, I should have been more specific. It's for people who wanna put in the effort and steer their own crawl to surface their own slice of the web. The article is just a little story of the journey
- pavel_lishin 2mo agoFrom the screenshot, it's very funny that one of the indexed sites is www.llresearch.org, which looks like it's run by a crackpot.
- dreamforever 2mo agoThere are all kinds of websites in here lol. There is some gold in here and I'm determined to surface it all. I had to wrap this up without full analysis cos it was dragging on
- BaudouinVH 2mo agoHow do you build a list of domains you want to index ? I see there is a fetcher and a spider in the code but so for I haven't found how to build that list.
- dreamforever 2mo agoAh, I forgot to mention that anywhere. You have to provide your own. You can start from a small set, like 10 websites you like that have a bit of character, and it will also add any domains it finds from those 10
- whatistrending 2mo ago[dead]
- lagrange77 2mo agoTook me a few minutes to realise it's not a domain name search engine.
- tpowell 2mo agoIt takes a bit of setup and a huge download, but every time I need a good domain I follow this old post from Derek Sivers. I have Claude de-dupe it and turn it into a searchable database (on my machine), then have it search genres and terms I'm looking for. It's a task Claude is very well-suited to, from the technical implementation to back-and-forth about selections. [link]: https://sive.rs/com https://sive.rs/com
- eggbrain 2mo agoNote -- if you do this, watch out for requesting access to "all tlds". They send you two emails per TLD -- one for your pending state, and one for your approved/rejected state. I suddenly had 1k+ emails flooding into my inbox, until I found the setting on their website to disable emails.
- thesuitonym 2mo agoHoly over-engineering, Batman!
- deleted 2mo ago[deleted]
- eichin 2mo agoReminds me that AltaVista's servers ran in 4G of RAM (there were famous, at the time, pics of the circuit boards - DEC was rightfully proud of this, 30 years ago) and that a modern AltaVista should run on a decent laptop :-)
- coredog64 2mo agoDuring AltaVista's prime, 32MB would have been a lot of memory for a typical computer.
- eichin 2mo agoI still can't find pics of the boards (other than on ebay :-) but https://seltzerbooks.com/alta.html https://seltzerbooks.com/alta.html has quotes like "To get a sense of the relative size of the RAM in AltaVista Search, consider that a typical personal computer today comes with 8 to 16MB of RAM, and the AltaVista Index Servers each have 6GB of RAM—about 500 times more." and "AltaVista shows the practical value of machines that have more than 4 gigabytes of physical main memory." (Dick Sites, DEC SRL) The book is from 1997.
- NetOpWibby 2mo agoThis is a damn good project. Makes me want to make headway into an idea I've had for quite some time P2P search...we'll see.
- gosolozero 2mo agoBuilt a free CLI for this as well: https://github.com/solozerolabs/Namera https://github.com/solozerolabs/Namera Checks socials and trademark too
- cpill 2mo agoI was thinking that a search engine that ignores anything with advertising on it would be useful. This has inspired me to give it a go (on the weekend even)
- frogger8 2mo agoFYI for those needing a list of domains Subject: I want all domains and subdomains https://groups.google.com/g/common-crawl/c/XC2QmOE-sdI?pli=1 https://groups.google.com/g/common-crawl/c/XC2QmOE-sdI?pli=1 or google for COMMON CRAWL
- deadbabe 2mo agoThe author struggled with categorization simply because they did not truly understand k-means clustering, a fundamental concept in this kind of computing science. You cannot just let a model run wild.
- vivzkestrel 2mo ago- since you know what k-means clustering is - why dont you tell me how you ll categorize this list with k-means?
- deadbabe 2mo agoRun k-means over embeddings
- formvoltron 2mo agoHere's an idea: Each day you could summarize all the new sites into an email. Call it NCSA "What's New" or something like that. ;)
- forix 2mo ago> pages classified as “portfolio” or “zine” or “software” push their outbound links way up the priority list, pages classified as “corporate” or “docs” push theirs down. And just like that, you recreated the internet of the 90s - early 00s. Brilliant!
- mewens 2mo ago[flagged]
- AndroVertex 2mo ago[flagged]
- kimi-haiku-9-9 1mo ago> There are 40-ish million domains. citation needed