8 ms·
SeenBefore: A search engine for what you have seen before
- adaml_623 14y ago"SeenBefore stores your information securely in the cloud from your work or home computers. So no matter where you read it you can still search for it even when your browsing history has been deleted." Erm... I think it's a good idea but I think many people would need convincing on the security front.
- bambax 14y agoSome of the things I "see", I would prefer they never show up in a Google search. I'm sure there is a configuration setting somewhere to deal with that, but it would be yet another thing to take care of.
- coenhyde 14y agoLooks useful. I've been well aware that google tracks everything I search for but I still don't like it.
- nuttendorfer 14y agoWho am I handing my data over to? I can't find this anywhere on the site.
- chuppo 14y agoAnd would it be possible to configure it to use my own "cloud"?
- vinnyglennon 14y agoDefinitely something we are looking into. Major barrier is the cost for someone keeping a server running 24*7 in cloud(Micro instance on AWS is 175 dollars a year).
- teach 14y agoSome of us already have servers running 24-7 in the cloud. I have two, for example.
- serbrech 14y agoWhat about using your server to run the software, but flat files as storage, I could point it to dropbox for example?
- vidarh 14y agoLots of us here have our servers.... Personally, being able to point it at one of my own servers and/or getting an API, would be fantastic.
- vinnyglennon 14y agoAdded a team page: https://www.seenbefore.com/pages/team https://www.seenbefore.com/pages/team . This had dropped off our list until launch. Sorry.
- ilija139 14y ago"40% of searches online are people simply looking for what they have already seen before." - How did they calculate this statistic?
- xlevus 14y agohttps://duckduckgo.com/?q=random+number+between+1+and+100 https://duckduckgo.com/?q=random+number+between+1+and+100 :P
- gingerjoos 14y agoIf I remember correctly, this was the result of some research done by a startup. Or was it Google? I couldn't tell you because this service didn't exist back when I read that article
- felipeko 14y ago"According to Yahoo, 40% of searches are simply searching for what you saw before." https://www.seenbefore.com/pages/faq#currently_do https://www.seenbefore.com/pages/faq#currently_do They should link to the study.
- vinnyglennon 14y agoLinked: https://www.seenbefore.com/pages/faq#currently_do https://www.seenbefore.com/pages/faq#currently_do . Thank you so much! :)
- adambyrtek 14y agoHow is this different from Google Search History? https://history.google.com https://history.google.com
- gingerjoos 14y agoGSH searches only within your Google search history. This guy searches through your entire browser(s) history.
- adambyrtek 14y agoThanks, I was confused by the fact that this integrates with Google Search.
- lucaspiller 14y agoYES! I've been looking for something like this for ages for stuff I have read on Hacker News.
- pbhjpbhj 14y agoI use a system adapted from http://www.gwern.net/Archiving%20URLs http://www.gwern.net/Archiving%20URLs to archive every page I've bookmarked (using FF) in the previous month. Then I just query with local tools. Not ideal, several flaws, but works well enough for me so far.
- gingerjoos 14y agoBeat me to it! This was something I had been planning to build on my own for a while, but didn't get around to . Congrats! Whenever I have tech discussions with friends I would recall something mentioned in a article I read via HN. But it would take me a whole lot of effort to get that link. Oftentimes I simply couldn't get hold of the link even after an hour of searching. Please do get the Firefox extension out. Would love to use it. Also, please do make sure the extensions/addons are stable. Have been facing problems with Annotary's extensions [1], for instance. By the way, do you have a crawler fetch the link content or do you send it from the user's browser? [1] https://getsatisfaction.com/annotary/topics/unstable_browser_extensions https://getsatisfaction.com/annotary/topics/unstable_browser...
- vinnyglennon 14y agocofounder here. Our first version spidered out for the content but a far more efficient way was to upload compressed version of the data from the user as we can then do hash checks for reference counting. Chrome extension has been used in the wild for last 3 months on 6 continents. Firefox extension too unstable at the moment(also Mozilla ten day review process), but hope to get it out with 1-2 weeks. Would love any feedback, good or bad!
- gingerjoos 14y agoWas mainly concerned about the scalability. For a large number of users, your server would have to handle a large number of concurrent connections while they uploaded their data. If you used spiders, you could push the URLs to a queue and process them at your convenience. How do you deal with 2 users looking at the same URL but seeing different things? example.com/me would be different for user1 and user2. Some pages would be very dynamic, eg. Facebook. And not everyone browses facebook/twitter behind https (which you do not index). Do you not index social networks? I like the fact that the extension requires no user input and works silently in the background. Has some trade-offs, but worth it. Cannot comment on the search quality yet because Chrome is only my secondary browser; not enough history to search for anything meaningful.
- gingerjoos 14y ago
- crntaylor 14y agoObvious point to raise: the reason people regularly delete their browser history is because they watch porn without turning on private browsing. How do you propose to deal with this? You'd need to provide at least the ability to selectively delete portions of the history. But you can selectively delete portions of your browser history too, and people don't - because it would be too easy to miss something. Instead, they just nuke the whole thing. How is your tool different?
- derekorgan 14y agoPorn sites are not recorded
- chris24 14y agoHow does your system define "porn sites"? What about if it was some porn site no one has ever heard of with an innocent-sounding name/domain?
- vidarh 14y agoHe apprently uses a list of 1.7 million sites. But you can also blacklist sites and have any existing entries for it removed: https://www.seenbefore.com/blacklist_items https://www.seenbefore.com/blacklist_items
- rickdale 14y agoI take advantage of my browsers history with porn. When I was in college the first bash script I wrote was to open a movie in my porn collection that I hadn't watched in the longest time. This was great. But now with streaming porn sites I don't have a huge collection and I often watch scenes that I can't seem to find later. There is a lot of porn out there. Sure people clear their browser history because their embarrassed by their porn obsession, but I think this tool could be very useful for pornaholics too.
- beaumartinez 14y ago
- ankimal 14y agoInteresting idea. Some quick questions: - How much data do you store per user? - How do I delete certain results? (preferably after the search comes back) - Another thing to consider is - After how much time does this just become as painful as finding that page through a search engine? - What version of the page gets stored? The latest or the one that I saw? I guess its one step better than Evernoting a page and adding tags myself. Good luck!
- vinnyglennon 14y agoMain issue with Evernoting and Bookmarking is that it requires an effort to say that today, this page is useful and I want to store it. Most pages I want to find are very things I did not think was useful at the time. Each unique page(unique as per the content) is stored per user. Our goal is to build the tools needed to find the information quickly, similiar to what hipmunk.com did for airline search. We have the added dimension of time to use.
- vidarh 14y agoThe main thing I use bookmarks for is categorisation. If you add the ability to tag and/or add notes that becomes part of the search terms, that'd be the killer feature for me - I could throw out my 3500 bookmarks and remove Xmarks (at least if we could get a way of automatically getting our existing bookmarks installed). I'm a paying Xmarks user, but if you were to add a way of tagging sites or adding a note, I'd happily pay for this instead. Just a freeform text field that I could add some keywords into that gets treated as part of the search would actually be sufficient for me.
- vdm 14y agoDup of Archify? https://www.archify.com/ https://www.archify.com/ > 40% of searches online are people simply looking for what they have already seen before. Citation link needed.
- vinnyglennon 14y agoCitation link: http://cond.org/sigir07.pdf http://cond.org/sigir07.pdf [PDF] Information Re-Retrieval: Repeat Queries in Yahoo’s Logs Abstract: "This paper explores repeat search behavior through the analysis of a one-year Web query log of 114 anonymous users and a separate controlled survey of an additional 119 volunteers. Our study demonstrates that as many as 40% of all queries are re-finding queries. Re-finding appears to be an important behavior for search engines to explicitly support, and we explore how this can be done."
- lifeisstillgood 14y agoWow, does 240 people even count as a sample. At Yahoo and Google log sizes its probably the error from cosmic rays in the data center.
- freshhawk 14y agoIf they selected them in a properly random way and had an effect close to 40% then yes, that probably does count as a sample.
- lifeisstillgood 14y agoAs someone who signed up to coursera stats 101, err... Why 40%?
- freshhawk 14y agoI am making some assumptions here absolutely, but because 40% is a large effect you don't need as many samples to be confident. The other way of looking at it is that maybe it's actually 35% or 45% but either way, that's still interesting, even with a rougher approximation of the actual "answer". If, for some reason, you needed to know if it was 40% or 40.01% because that mattered to you then you would absolutely be annoyed at the small sample size. If the finding was 2% then we would care about the uncertainty of +/- 5% since the finding is dwarfed by the error rate. That's a smaller effect size so you would need more samples to separate reality from the noise. I am, by the way, pulling all of these numbers out my ass. Your stats 101 class will teach you the formulas to calculate the actual error bars at work here as well as the assumptions you need to make about the distribution of the data to use those formulas.
- pcl 14y agoI love the date visualization. This is something I think that pretty much all search results could benefit tremendously from.
- elviejo 14y agoI'll take it for a spin... this is something I've wanted for a long time. I was going to hack it by making chrome bookmark every site I visit with a tag:history then when I wanted to search for a site that I've already visited I was going to just search with that tag.
- martythemaniak 14y agoI'm going to give it a spin and let know what I think (it'll take a few weeks of usage), but I can tell you right now that it's definitely solving a real problem I have.
- ThomPete 14y agoI love this idea but I think you will find more traction by turning it into a kind of bookmarking app with less focus on the search engine part.
- alanorourke 14y agoGreat idea. Love it.
- ippisl 14y agoOne feature that could help this: verifying that the account holder is the one using the computer, before showing results. Without this, assuming this plugin is always-on on all the computers one uses, breaking user's privacy just becomes too easy. And there's a lot of data one might want to leave private except porn(and usually don't post them in facebook): medical issues, sexual issues, marriage and some other relationship issues, drugs issues and probably others.
- beaumartinez 14y agoYou could hook it into Chrome's history API[1]. [1] http://developer.chrome.com/extensions/history.html http://developer.chrome.com/extensions/history.html
- arikrak 14y agoIf it would let people search there history and bookmarks, they could start benefiting from it right away.
- vinnyglennon 14y agoCo-founder here. This took us by surprise, we were planning to have Firefox and Safari support done by launch. At this stage, it is priceless to know if we are solving a real problem people have. Also, is this something people would pay for (loops back to if this is enough of a pain point). From the moment we start charging, is the moment we start learning.
- skinnymuch 14y agoI don't spend much online, especially when it comes to recurring fees. However, I use Pinboard enough that it's going to be hard for me to resist not paying the $25 fee they charge for archiving bookmarked pages for the second time. Yeah I could code/hack together something myself and have been thinking of doing it [for fun], but ya know :p So, yeah, count me in as being interested.
- vidarh 14y agoSee comment elsewhere: With tagging or (simple plain text) notes attached, absolutely. Even moreso with a simple API and/or support to push the cached content to my own server. If it could be selectively enabled for private content too, then even better (e.g. there's several extensive private Wiki's I use regularly that are not sensitive enough that I'd worry about getting them indexed, and I'd love to be able to tell you to index them but perhaps disable the caching).
- cnlwsu 14y agoI am going to try this out because it seems like what I spend a large portion of my time doing. The security and privacy of this scares me a lot though.
- webwanderings 14y agoI don't know why I should give you my browser history. I'd like to keep it to myself.
- pilooch 14y agoI agree, sounds like a crazy thing to do when this could easily be achieved locally on my machine. Or am I missing something ?
- freshhawk 14y agoBut if a small piece of software was installed on your machine it wouldn't be "in the cloud". We know that makes everything better. Ok well not application performance ... or cost ... or usability ... but still, "the cloud".
- lesterbuck 14y agoIn a similar vein, Pinboard offers to snapshot and full text index all your bookmarks, for a small annual fee: https://pinboard.in/upgrade/ https://pinboard.in/upgrade/
- espeed 14y agoDoesn't Google already have this? Go to... Show Search Tools -> All Results -> Visited Pages
- eli 14y agoI'm pretty sure that's only for filtering pages visited via a prior google search In theory Chrome lets you search through your history for pages, but it doesn't seem to actually work very well for me.
- bbrian 14y agoThere's also weekly reports that tell you what sites you've been visiting the most, what time of day/what days you visit sites most, and how many pages SeenBefore added to your file.
- akldfgj 14y agoI get that deployment is easier when it is vendor hosted, but this really should be a local app using local storage, withe maybe transient server-side storage for syncing between machiens.
- tungwaiyip 14y agoYes, I have seen it before! I have build a personal search engine MindRetrieve back in 2005. http://mindretrieve.net http://mindretrieve.net Specifically I'm not comfortable for big web company to keep the history of my web activity. So I make it work completely locally. My project did not get much uptake, probably my lackluster marketing and other assorted issues are to blame. So good luck on this one!
- skinnymuch 14y agoToo bad. Looks like a really cool project. Even more impressive when seeing how old it is.
- pseingatl 14y agoMac version? Planned six years ago?
- Osiris 14y agoI think this is a great idea. I've been using Opera, which has a full-text search capability for history, but it's limited to the machine you're using it on. I often find interesting articles on Hacker News while I'm at home that I want to find again when I'm at work. Being able to search by browser history across machines is fantastic for me.
- patriciaorgan 14y agogreat idea and works brilliant!
- willegan 14y agoGreat idea, I assume with Chrome's new incognito browsing, this won't be picked up on seenbefore or am I wrong?
- StavrosK 14y agoHmm, this is similar to http://historio.us http://historio.us, which I built. However, this doesn't require any user interaction, which might work well. Do you store just the URL and depend on Google returning the results? How does it work exactly?
- andy_boot 14y agoI remember thinking that http://historio.us http://historio.us was a neat idea. But Seen Before requires less effort on my part as a user -> I am more likely to use it. I just continue to google as per normal and now I have an extra option on the right to filter results.
- StavrosK 14y agoThis is true. The use cases are a bit different, but I still don't know exactly how this works so I can't say.
- mburns 14y agoFor those using Firefox, there is a similar add-on called RecallMonkey. https://addons.mozilla.org/en-us/firefox/addon/prospector-recall-monkey/ https://addons.mozilla.org/en-us/firefox/addon/prospector-re...
- krassif 14y agoSimilar to my project Peerbelt.com. A notable difference is Peerbelt runs entirely on the client to void privacy concerns. Vinny, let's chat and see if we can collaborate. Cheers, -Krassimir the Peerbelt founder