11 ms·
A Unix-style personal search engine and web crawler for your digital footprint
- wydfre 5y agoIt seems pretty cool - but I think falcon[0] is more practical. You can install it from the chrome extension store[1], if you are too lazy to get it running yourself. [0]: https://github.com/lengstrom/falcon https://github.com/lengstrom/falcon [1]: https://chrome.google.com/webstore/detail/falcon/mmifbbohghecjloeklpbinkjpbplfalb?hl=en https://chrome.google.com/webstore/detail/falcon/mmifbbohghe...
- grae_QED 5y agoAre there any Firefox equivalents to Falcon? I'm very interested in something like this.
- news_to_me 5y agoIf it's a WebExtension, it's usually not too hard to port to Firefox (https://developer.mozilla.org/en-US/docs/Mozilla/Add-ons/WebExtensions https://developer.mozilla.org/en-US/docs/Mozilla/Add-ons/Web...)
- nathan_phoenix 5y agoIn the issues someone says that it works even in FF. You just need to change the extension of the file. Tho I didn't try it yet. https://github.com/lengstrom/falcon/issues/73#issuecomment-629098952 https://github.com/lengstrom/falcon/issues/73#issuecomment-6...
- yunruse 5y agoI love this idea, but the name “digital footprint” sort of implies it’s what effect you’ve had on the Internet for helping keep your online persona under control: your tweets, comments, emails, et cetera. But this is a great idea! Having a search engine for vaguely _anything_ you touch very much does look like it’d increase the signal:noise ratio. It’d be interesting to be able to add whole sites (using, say, DuckDuckGo as an external crawler) to be able to fetch general ideas, such as, say, “Stack Exchange posts marked with these tags”.
- flanbiscuit 5y ago> but the name “digital footprint” sort of implies it’s what effect you’ve had on the Internet for helping keep your online persona under control: your tweets, comments, emails, et cetera. I had the exact same thought when I saw that in the title. That would also be a cool idea to be able to search within your own online accounts. So this is what the project's description of what "digital footprint" means: > Apollo is a search engine and web crawler to digest your digital footprint. What this means is that you choose what to put in it. When you come across something that looks interesting, be it an article, blog post, website, whatever, you manually add it (with built in systems to make doing so easy). If you always want to pull in data from a certain data source, like your notes or something else, you can do that too. This tackles one of the biggest problems of recall in search engines returning a lot of irrelevant information because with Apollo, the signal to noise ratio is very high. You've chosen exactly what to put in it. If I'm interpreting this correctly, this seems like an alternative way of bookmarking with advanced searching because it scrapes the data from the source. Cool idea, means I have to worry less about organizing my bookmarks.
- Minor49er 5y agoThis looks really cool. It's beyond the scope of this project, but I think that having something like this as a browser extension would make it easier to use: instead of manually copying and scraping links, it could index and save pages that you've been on, placing much more significance on anything that you've bookmarked. Granted, this is just an immediate thought. I'm going to give this a proper try once I have some more spare time.
- ya1sec 5y agoGreat thought. I've adopted a similar workflow using the https://www.are.na/ https://www.are.na/ chrome extension to save links to channels. Might be a nice touch to feed channels into the engine using their API
- Minor49er 5y agoThis looks like a fun way to explore topics. I just signed up
- MisterTea 5y ago> I've wasted many an hour combing through Google and my search history to look up a good article, blog post, or just something I've seen before. This is the fault of web browser vendors who have yet to give a damn about book marks. > Apollo is a search engine and web crawler to digest your digital footprint. What this means is that you choose what to put in it. When you come across something that looks interesting, be it an article, blog post, website, whatever, you manually add it (with built in systems to make doing so easy). So it's a searchable database for bookmarks then. > The first thing you might notice is that the design is reminiscent of the old digital computer age, back in the Unix days. This is intentional for many reasons. In addition to paying homage to the greats of the past, this design makes me feel like I'm searching through something that is authentically my own. When I search for stuff, I genuinely feel like I'm travelling through the past. This does not make any sense. It's Unix-like because it feels old? It seems like the author thoroughly misses the point of unix philosophy.
- chris_st 5y ago> So it's a searchable database for bookmarks then. It appears to be that, but it appears also to pull out the content of the web page and index that too, so you can (presumably) find stuff that isn't in the "pure" bookmark, which I think of as a link with maybe a title.
- nextaccountic 5y agoI think browsers should download a full copy of each bookmark (so you can still see it when they are taken down) and make it fully searchable. Actually, I've been trying to find Firefox extensions that give a better interface to bookmarks and there doesn't seem to be one. It's like, people don't use bookmarks anymore and accept that it might as well not exist, and use something else. It's telling that Firefox has two bookmark systems built-in (pocket and regular bookmarks) and they aren't integrated with each other; I suppose that people that use pocket never think about regular bookmarks. edit: but my pet peeve is that it isn't easy to search history for something I saw 10 days ago but I don't remember the exact keywords to search.
- SahAssar 5y agoLooks very much like one of the ideas I've been thinking of building! The way I planned to do it was to use a similar approach to rga for files ( https://github.com/phiresky/ripgrep-all https://github.com/phiresky/ripgrep-all ) and having a webextension to pull all webpages I vist (filtered via something like https://github.com/mozilla/readability https://github.com/mozilla/readability ), dump that into either sqlite with FTS5 or postgres with FTS for search. A good search engine for "my stuff" and "stuff I've seen before" is not available for most people in my experience. Pinboard and similar sites fill some of that role, but only for things that you bookmark (and I'm not sure they do full-text search of the documents). --- Two things I'd mention are: 1. Digital footprint usually means your info on other sites, not just things I've accessed. If I read a blog that is not part of my footprint, but if I leave a comment on that blog that comment is part of it. The term is also mostly used in a tracking and negative context (although there are exceptions), so you might want to change that: https://en.wikipedia.org/wiki/Digital_footprint https://en.wikipedia.org/wiki/Digital_footprint 2. I don't really get what makes it UNIX-style (or what exactly you mean by that? There seems to be many definitions), and the readme does not seem to clarify much besides expecting me to notice it by myself.
- eddieh 5y agoI've been toying with an idea like this too. I set my browser to never delete history items years ago, so I have a huge amount of daily web use that needs to be indexed. The browser's built in history search has saved me a few times, but it is so primitive it hurts.
- MacroChip 5y agoI made https://chrome.google.com/webstore/detail/histree/linpklflmolmnhckgoojppnfhajngaoh?hl=en https://chrome.google.com/webstore/detail/histree/linpklflmo... to put your browsing into a tree view. The search does not search the site content, so it's different from full indexers, but it's a nice enhancement to browser history.
- aero-glide2 5y agoDon't know how to message users on hackernews, so posting as a reply here hope you don't mind. Saw your comment from 5 years ago about wishing Orbiter was open source. https://news.ycombinator.com/item?id=12943028 https://news.ycombinator.com/item?id=12943028 The author has now made it open source! https://www.orbiter-forum.com/threads/orbiter-is-now-open-source.40023/ https://www.orbiter-forum.com/threads/orbiter-is-now-open-so...
- deleted 5y ago[deleted]
- toomanyducks 5y agoIf nothing else, that README is fantastic!
- pantulis 5y agoReminds me a lot of DEVONthink for Mac
- simonw 5y agoMy version of this is https://dogsheep.github.io/ https://dogsheep.github.io/ - the idea is to pull your digital footprint from various different sources (Twitter, Foursquare, GitHub etc) into SQLite database files, then run Datasette on top to explore them. On top of that I built a search engine called Dogsheep Beta which builds a full-text search index across all of the different sources and lets you search in one place: https://github.com/dogsheep/dogsheep-beta https://github.com/dogsheep/dogsheep-beta You can see a live demonstration of that search engine on the Datasette website: https://datasette.io/-/beta?q=dogsheep https://datasette.io/-/beta?q=dogsheep The key difference I see with Apollo is that Dogsheep separates fetching of data from search and indexing, and uses SQLite as the storage format. I'm using a YAML configuration to define how the search index should work: https://github.com/simonw/datasette.io/blob/main/templates/dogsheep-beta.yml https://github.com/simonw/datasette.io/blob/main/templates/d... - it defines SQL queries that can be used to build the index from other tables, plus HTML fragments for how those results should be displayed.
- tomcam 5y agoHoly crap you should submit as a Show HN
- simonw 5y agoIt's failed to make the homepage a few times in the past: https://hn.algolia.com/?q=dogsheep https://hn.algolia.com/?q=dogsheep - the one time it did make it was this one about Dogsheep Photos: https://news.ycombinator.com/item?id=23271053 https://news.ycombinator.com/item?id=23271053
- mosselman 5y agoSimon is not an unknown on HN.
- gizdan 5y agoWow! That's super cool. I will have to check this out at some point. Am I correct in understanding that the pocket tool actually imports the URLs contents? If not, how hard would it be to include the actual content of URLs? Specifically, I'll probably end up using something else (for me NextCloud bookmarks).
- zerop 5y agoHow's it different from instapaper like services. There is also open source alternative of instapaper called wallabag.
- dandanua 5y agoA similar tool – https://github.com/go-shiori/shiori https://github.com/go-shiori/shiori
- encryptluks2 5y agoShiori is "okay" but is not actively being maintained at all. The original author abandoned it and the new maintainer apparently never planned on supporting it.
- soheil 5y agoHas the author tried pressing CMD+Y to view and search browser history?
- soheil 5y agoThere is something really strange about a lot of recent Go projects including this one. I can't put my finger on, but the combination of the author and the type of problem they choose to tackle oftentimes seems baffling to me. Most projects seem to be solving a problem that is often misidentified or otherwise badly solved, but somehow the focus ends up being on the code architecture or the UI design. It's like they're trying to solve a problem just for the sake of writing some code and the correct way to use Go idiomatically or something and don't really care about the problem or how well the solution actually works.
- jrm4 5y agoYeah, as a bit of an old-timer, I'm trying to learn to stop worrying and learn to love watching everybody reinvent wheels? For me it's "why are you people doing that in Javascript?" that continually comes up in my own head, but I suppose I should try to be patient and see if anything comes of it.
- asdff 5y agoI think projects like this are just resume builders. Everyone says "show a project on github," well here is one of these projects. The dev is probably hoping this helps land them a job offer. Its fine if the project is ultimately "lame" in some way, since its not the job description of a developer to make a cool unique app, but to follow orders from the project manager and write code, which is what this project shows this dev can do.
- amirGi 5y agoOP here! Actually, I don't care about this landing me a job lol, I wrote this purely for fun and to (hopefully) be able to use it :)
- amirGi 5y agoOP here, I'm not claiming to be a Go expert, in fact I'm far from it! I used Go as my backend because I didn't want to use Node - it's very likely that I might be using it in ways it might not be intended, please do let me know if so!
- ctocoder 5y agowrote something along the same ilk but got distracted https://github.com/dathan/go-find-hexagonal https://github.com/dathan/go-find-hexagonal
- fidesomnes 5y agoAdding support for transcribed voice notes like from Otter would be nice.
- dpcx 5y agoSimilar also to Promnesia (https://github.com/karlicoss/promnesia https://github.com/karlicoss/promnesia), which includes a browser extension to search the records.
- encryptluks2 5y agoWhy do all these bookmark projects: 1. Rely on JavaScript for the interface. Being built in Go, why not just paginate the results and utilize Bleve or Xapian for search? 2. Store data in a format that is not easily readable by itself. The only exception to this is nb. 3. Suck at CLI tools. I'm looking to rclone, Hugo, kubectl, etc for the right way to build a CLI.
- rhn_mk1 5y agoThis seems similar to recoll augmented with recoll-we. https://addons.mozilla.org/en-US/firefox/addon/recoll-we/ https://addons.mozilla.org/en-US/firefox/addon/recoll-we/
- ryanfox 5y agoI run a similar project: https://apse.io https://apse.io It runs locally on your laptop/desktop, so you don’t need a server to host anything. Also, it can index everything you do, not just web content. It works really well for me!
- cratermoon 5y agoInteresting project but some of what the author writes just sounds flat-out weird. "The first thing you might notice is that the design is reminiscent of the old digital computer age, back in the Unix days." "Apollo's client side is written in Poseidon." I had to look that up: Poseidon is not a language, it's just a javascript framework for event-driven dom updates.
- etherio 5y agoThis is cool! Similar to one of the goals I'm trying to accomplish with Archivy (https://archivy.github.io https://archivy.github.io) with the broader goal of not just storing your digital presence but also acting as a personal knowledge base.
- kordlessagain 5y agoCool! It’s great to see others thinking about this. I’ve been working on https://mitta.us https://mitta.us for a while now and it uses solr, a headless brrowser and google vision to snapshot and index full text. The UI is a bit odd but you can just append mitta.us/ to any URL to save it.
- ThinkBeat 5y agoI use Evernote for this. You can set it ot save a link, a screenshot, or content of the page. You can add tags if you want, and it is also easy to annotate it so you can remember the context better. You can also add links to other post inside Evernote. Pocket is also a great tool I used for many years. Quite similar and different. Both have browser extensions, so it is easy to clip. With Evernote I even have shortcuts defined so I dont have to click for the webpage to be clipped.
- jll29 5y agoMicrosoft Research's Dr. Susan Dumais is the expert on this kind of personal information management. Her landmark system (and associated seminal SIGIR'03 paper) "Stuff I've Seen" tackled re-finding material: http://susandumais.com/UMAP2009-DumaisKeynote_Share.pdf http://susandumais.com/UMAP2009-DumaisKeynote_Share.pdf
- totetsu 5y agothere used to be an actity timeline journal program i ran on ubuntu that let me see which days i accessed which files. It was very useful as a sudent.
- alanh 5y agocode comment in the readme describes the Record as constituting an 'interverted index'. typo for inverted? although it is not obvious to me what would make this an inverted index instead of a normal index
- smusamashah 5y agoThis sounds similar to Monocle https://github.com/thesephist/monocle https://github.com/thesephist/monocle Demo: https://monocle.surge.sh/ https://monocle.surge.sh/ Blog post explaining motivation https://thesephist.com/posts/monocle/ https://thesephist.com/posts/monocle/
- sooheon 5y agoLoved the monocle blog, as well as other posts on that site. [Finda](https://keminglabs.com/finda/ https://keminglabs.com/finda/) was another one I saw in this space.
- anthk 5y agoIs this like recoll, hyperstrayer and so on?