Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
costco
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
costco
11mo ago
You may find these better: https://book.sv/similar?id=211570 > Note 1: If you only provide one or two books, the model doesn't have a lot to work with and may include a handful of somewhat unrelated popular books in
32.
▲
by
costco
11mo ago
This is a result of the use of positional embeddings, which typically results in the final item being weighted very highly. The problem is that this information is shown to be very relevant to the task of predicting the next item interacte
33.
▲
by
costco
11mo ago
Thank you for the kind remarks. I would say your request is reasonable and it’s something I thought about but it’s worth noting that if you scroll down on a Goodreads book page you can see all of the users who gave it 1 star, 2 stars, etc
34.
▲
by
costco
11mo ago
You don't have to import your Goodreads profile. You can type titles and authors in the box and find books to add to the list that way.
35.
▲
by
costco
11mo ago
I would agree the results are generally OK but do not feel magical in most cases (I think in some specific cases they do though). The results can be not great if you add books across many disciplines. For instance if you add "The Ele
36.
▲
by
costco
11mo ago
> Note 1: If you only provide one or two books, the model doesn't have a lot to work with and may include a handful of somewhat unrelated popular books in the results. If you want recommendations based on just one book, click the &q
37.
▲
by
costco
11mo ago
I use `fetch` on relative endpoints so that's odd. There shouldn't be any external API calls on my website other than whatever the Cloudflare captcha uses. I also use HTTPS-only in Chrome and did not experience any issues. I ju
38.
▲
by
costco
11mo ago
I use a Hetzner server with Ryzen 7 3700X and an SSD. I think I could get the model to work with ONNX web but it'd be a 2GB download so the user experience wouldn't be too great. My Meilisearch index is ~40GB but I don't kno
39.
▲
by
costco
11mo ago
It's recommended that you put at least 3 books in. If you would like recommendations just based on one book, click the similar button on the book, it should take you to this page: https://book.sv/similar?id=297914
40.
▲
by
costco
11mo ago
You can import the first or last 64 books of your read, to-read, or currently-reading shelves if you press the "Import Goodreads" button and provide your Goodreads ID.
41.
▲
by
costco
11mo ago
Thank you for the recommendations. I didn't try BERT4Rec because I assumed it would perform the same or worse as what I already had after having read https://dl.acm.org/doi/pdf/10.1145/3699521 . The TIG
42.
▲
by
costco
11mo ago
Worked for me, could be due to server being overwhelmed Here is the URL with your books: https://book.sv/#52752877,46049530,18437030,52480873,3260654...
43.
▲
by
costco
11mo ago
Yes I would say the handling of series is probably the biggest problem. Once my test metrics got to a point I was happy with and my quality spot checks passed (can I follow the models recommendations from one generic history book to Steven
44.
▲
by
costco
11mo ago
I didn't want an agent to get stuck on an infinite loop invoking endpoints that cost GPU resources. Those fears are probably unfounded, so if people really cared I could remove those. /similar is blocked by default because I don
45.
▲
by
costco
11mo ago
If you want recommendations solely based on one book, please try the similar page: https://book.sv/similar?id=13566692 These seem to fit the description you are going for better. The model is trained to predict the next bo
46.
▲
by
costco
11mo ago
Everything (namely Meilisearch, Postgres and the web server in Go) besides the model inference is running on a Hetzner server with a large SSD and an "AMD Ryzen 7 3700X 8-Core Processor." The data.ms directory is about 40GB. Onc
47.
▲
by
costco
11mo ago
Not sure if I can. At the very least book descriptions most likely could not be distributed. There is an academic dataset with around 200M reviews though: https://cseweb.ucsd.edu/~jmcauley/datasets/goodreads.html
48.
▲
by
costco
11mo ago
I think I will expand the input books limit (sadly requires retraining) and or the output books limit of 30.
49.
▲
by
costco
11mo ago
I'm not an expert by any means but as far as sequential recommendations go, aren't SASRec and its derivatives pretty much the name of the game? I probably should have looked into HSTUs more. Also this / sparse transformers
50.
▲
by
costco
11mo ago
It's explicitly trained to predict the next book read in a sequence, which is why you get that behavior. There's probably a better way for me to handle it rather than having 5 books from the same series tend towards the top thoug
51.
▲
by
costco
11mo ago
Thank you for the compliments :) I used 50-100 datacenter proxies. I just logged requests made by the iOS app with Charles and then recreated the headers to the best of my ability though the server did not seem to be very strict at all. W
52.
▲
Show HN: I scraped 3B Goodreads reviews to train a better recommendation model
(book.sv)
606 points
by
costco
11mo ago
|
259 comments
53.
▲
by
costco
1y ago
Finally someone in this thread who gets it. In fact many of the temporary phone number registration services used to make bot accounts have pages on their website where they say "Do you live in country X and have access to large amoun
54.
▲
by
costco
1y ago
I've read that at least for for GSM VoIP gateway setups, people typically rotate SIMs because it looks suspicious for one customer to be making calls all day. In fact there is a whole industry most people have no idea about dedicated
55.
▲
by
costco
1y ago
Related: It’s only about $5k a year to run your own registrar (which is different than a gTLD). You have to sign agreements with Verisign and others to register domains on TLDs that people actually want but it’s not unduly hard and I think
56.
▲
by
costco
1y ago
This is awesome, and the low cost is especially impressive. I rarely have the motivation after working on a side project to actually document all the decisions made along the way, much less in such a thorough way. Regarding your CoreNN li
57.
▲
by
costco
1y ago
If you already have an understanding of how Tor works and want to know how attacks on it work, read these! - https://github.com/mikeperry-tor/vanguards/blob/master/READM... - https://github.co
58.
▲
by
costco
1y ago
This page on the mailing list has links to cases of people who were caught because of an unknown flaw in Tor: https://archive.torproject.org/websites/lists.torproject.org... I can't find a link, but I think peopl
59.
▲
by
costco
1y ago
Were you running specifically a bridge or just a non exit relay? Bridges are generally unlisted and are somewhat expensive to mass scrape (the bridge distributors will require captcha or email or Telegram etc) so they are less likely to sh
60.
▲
by
costco
1y ago
I had used bloom filters in the past without really understanding how they worked. Then one day I decided to implement them just going off the Wikipedia article with the 32-bit MurmurHash function and was surprised at how simple it was. I
More ›