5 ms·
Hi, I'm Ben, CTO of Academia. Everyone who works at Academia would love it if we were able to make advanced search free. When we first decided to build a prem
by benl 9y ago
Hi, I'm Ben, CTO of Academia.
Everyone who works at Academia would love it if we were able to make advanced search free.
When we first decided to build a premium account, we also made the decision to not take anything out of the free account. Strange as it may seem, the free account never had full-text search because we couldn't justify the cost of building it (full-text search of 20MM PDFs at our traffic levels is expensive to operate). We built it for the premium account because people asked for it in our initial research - and we would love to be able to eventually move it into the free account.
On the team we all agree that we want to keep building premium features in order to make the platform sustainable. The author of this article takes the view that advanced search is not a feature that should be paid-for. My view is that we intend keep building features until we have something that is worthy of his support. The support of the academics who use and enjoy the platform, both in free and paid accounts, is what will keep it around and growing for the long term.
- zajd 9y ago> On the team we all agree that we want to keep building premium features in order to make the platform sustainable. Is there any concern that paywalling features will reduce the number of users on the site?
- benl 9y agoYes, we don't want that to happen. That's one reason (but not the only reason) that we haven't paywalled any of the existing free features.
- merraksh 9y agoI think it would be a little easier to accept (and probably easy to implement) if, in addition to title search, authors and keywords were also searchable. A search for my PhD advisor's last name results in 0 papers, though I uploaded a few papers that he and I coauthored.
- benl 9y agoThat's a good idea. It's already possible to search for author names separately and then find the papers on their profile (if they have one), and it's also possible to search for a research interest tag and find the papers tagged with it. It would make sense to unify those with the title search.
- tomcam 9y ago> he and I coauthored …which means that you wrote them and he, ah, "co-authored" them?
- merraksh 9y agoYes, I did write most of the text, but without his guidance on the research and on what to write about more in depth, there would have been no article. He's been the best advisor. Don't you dare say bad things about him! :-)
- tomcam 9y agoNice to hear!
- maksimum 9y agoIf you have enough experience to guide me wrt what the interesting questions in a field are and how to make potential answers to those questions "meaningful", then I'm happy to write and have you as a co-author.
- anigbrowl 9y agoWhy don't you just figure out how much it's gonna cost you to implement, publish that, and ask for the money? I am 100% on your side but I'm not interested in a bunch of organizational doublespeak, just say you can't afford it and you need $600k or whatever to do this because (budget breakdown). Life's too short to waste time parsing business emo messaging. How much money do you need? Publish your revenue model and stuff so people can help you with it .
- ubernostrum 9y agoThis is honestly a terrible idea. It will get posted to HN, and the comments will be a flood of "why do you need that much to do this, I could do it for $5 using list of this week's vaporware buzzwords that won't survive to the end of the month".
- deleted 9y ago[deleted]
- nojvek 9y agoWhy is accountability a bad idea? Charities do that all the time. Some are well funded exactly for this reason. Most people understand good things cost money when thinking for long term.
- ubernostrum 9y agoAccountability to well-informed people who understand the tradeoffs involved in building sustainable real-world infrastructure is great. HN is not an audience of that kind of people.
- anigbrowl 9y agoIt is when people want to launch things and get lots of buzz, then HN is great. But when people express criticism suddenly the user community is a bunch of angry peasants to be kept at arm's length. I've gotta say I'm getting real tired of entrepreneurs that want to be everybody's friend when they're getting their exciting new venture off the ground but are too cool to discuss the nuances of business with anyone outside the VC bubble whenever they run into a PR problem. That's a general statement, not directed at the academia.edu team. It's a sad reality that a lot of what passes for entrepreneurship today involves telling users, employees, and investors that they're each the most important group so as to make as many people as possible happy, right up to the point where a conflict of interest emerges and then trying to obscure the fact of its existence with platitudes.
- hayd 9y ago> full-text search of 20MM PDFs at our traffic levels is expensive to operate Why not do a free full-text search on the abstract or first couple of pages (or few hundred words)? This might make the premium search upgrade more subtle but I wonder if it could help the majority (?) of legitimate searches get the hit they expected/hoped for. ... Perhaps this would appease the haters?
- benl 9y agoThat's really good idea. As Richard mentions in another thread, search is not actually the primary discovery mechanism on Academia. The social features (the news feed, bookmarks, sharing, recommendations) are the primary discovery mechanism. These are all free and we want to keep them that way.
- dsacco 9y agoThis feature isn't going to be effective for you without some rethinking. People can simply search on Google: site:academia.edu "Potterheads" and retrieve all the results. Full-text search is freely provided by Google. Judging by how many people caught on to using a similar trick with WSJ, it will start to become a very popular way of interacting with the site.
- keiferski 9y agoYou're overestimating the average Google user's google-fu. Most people aren't aware of the 'site:' functionality.
- dsacco 9y agoFair, but I also wouldn't be surprised if someone just builds a browser extension to do this automatically on academia.edu. Also, that sounds like an assumption. Do you have anything concrete to back that up? Assumptions are often broken in painful ways for businesses. You might be surprised how many academics are motivated to learn about Google's advanced search and tell their friends about it if it saves them money.
- xaa 9y agoBioinformatician here. I appreciate that you need to make money somehow. I do find this a little hard to believe, though: > full-text search of 20MM PDFs at our traffic levels is expensive to operate Assuming you convert them to text once, index them, and put them in a standard FTS engine, I'd guess it is on the order of 100GB-1T of text (max), plus some more for the index (basing these estimates with my experience text mining PubMed Central and MEDLINE). So it can all fit on a pretty standard server. Maybe at 100 req/s it would take a few. Yes, you'd want replication. The number of servers required to get good latency FTS is the part of this that I'm least familiar with. Anyone have a ballpark, given these estimates, on what kind of hardware would be required? (I could easily be wrong, and indeed this is very expensive. If so, I'd be curious about ballpark numbers)
- minxomat 9y agoI maintain a P2P on-premise FTS search. Though this one indexes many types of text (plain, HTML, PDF, DOCs). One 8c server (running about 10 workers) can handle 8 to 10qps, depending on the depth required. This is on an index of 20 million documents. If the number of workers is constant, doubling the index will have the qps. 2 million docs take about 50GB of disk space (20 million = 500GB, 1TB with redundancy). It's better to go with SSD arrays here, since random IOPs are much higher than for other workloads. This can skyrocket cost. So for this (our) system, it could be as cheap as $1k for the hardware, e.g. using the Foxconn Purus cloud server: http://www.bargainhardware.co.uk/cheap-e5-2600-lga2011-sixteen-core-cloud-foxconn-server-configure-to-order/ http://www.bargainhardware.co.uk/cheap-e5-2600-lga2011-sixte...
- xaa 9y agoThanks so much. Fascinating. I've made many of these types of app but never to "web scale". Bioinformatics apps are a bit niche, and we take the view that our non-paying academic "customers" can wait however long it takes to finish the query, in the unlikely event that there is high load. Just to make sure I understand correctly, that's about $1-2K of one time cost per 10qps (w/o SSD, and not counting power and maintenance, etc)? When I first saw "cloud server", I thought that was a per-month rental cost, but the link is for actual in-house hardware. If this is even close to correct, my suspicions seem confirmed. Except for one thing. I have no idea how many qps a site like Academia would have. 100qps was completely out of my ass, but it seemed hard to imagine it being any more than 1-2 orders of magnitude higher, at most. Any guess on that?
- rspeer 9y agoHi Ben. Will Academia ever stop impersonating people in e-mail? I receive e-mails all the time purporting to be actually sent by academic colleagues, which are instead form messages sent by Academia.edu. I know that you're not sending these messages with permission. I know that my dead co-author is not giving you permission to send unsettling e-mails from him.
- benl 9y agoHi - Please forward me a copy of one of those emails (to my first name at academia.edu). We only ever send emails from users in response to a request from them to do so (e.g. they send a message or invite a co-author). Please send me the email and I will look into it.
- rspeer 9y agoOkay... this turns out to have been a false accusation. I'm very sorry. The company that does this is ResearchGate, not Academia. Academia's e-mails have appropriate From: lines, and don't appear to be sent from beyond the grave, and you deserve credit for that.
- blusterXY 9y agoHuh? Lucene can index that many documents no problem.