3 ms·
It seems odd to me that so many of these projects are being launched by people who have only just discovered and/or joined HN. I'm worried this is just becoming
by nativeit 1y ago
It seems odd to me that so many of these projects are being launched by people who have only just discovered and/or joined HN. I'm worried this is just becoming LinkedIn for AI opportunists.
- parpfish 1y agoI’ve got a side project that I may (someday) do a show HN with. However, I’d probably make a new account for that because the project is connected to my real name/portfolio and I don’t want that connected with my pseudonymous comments here
- nativeit 1y agoI considered that, but then why would anyone obfuscate this really very reasonable scenario by choosing another ostensibly pseudonymous username?
- deleted 1y ago[deleted]
- fsmv 1y ago[deleted]
- parpfish 1y agoI imagine that this is a common problem and it could be another cool “unlockable” on HN, like the downvotes at 500 karma. Once you get X karma or account age >Y years, you can make one anonymous submissions each quarter that comes from an non-user but still get some sort of “verified” badge that proves it comes from a legit user.
- refulgentis 1y agoYou nailed it IMHO. I quit my job at Google 2 years ago to do LLM stuff, was looking forward to having HN around, but discussions re: LLMs here are a minefield. Why? Everyone knows at least a little, and everyone has a strong opinion on it given the impact of it. People sharing stuff sell it way high, and as with any new thing where people are selling, there's a lot of skeptics. Then, throw in human bias towards disliking what seems like snark / complaining, so stuff with substance gets downvotes. SNR ratio is continually decreasing. Let's dig into why this one is weird: My work inferences using either 3P provider, which do caching, or llama.cpp, in which I do caching. (basically, picture it as there's a super expensive step that you can skip by keeping Map<input string, gpu state>) So I log into HN and see this and say to myself: 3x! throughput increase? This is either really clever or salesmanship, no way an optimization like that has been sitting around on the groud. So I read the GitHub, see it's just "write everyones inputs and outputs to disk, you can then use them to cobble together what the GPU state would be for an incoming request!", and write a mostly-polite comment below flagging "hey, this means writing everything to disk" Then I start replying to you...but then I throw away the comment, because I'm inviting drive-by downvotes. I.e. the minefield describe up top, and if you look like you're being mean, you'll eat downvotes, especially on a weekend. And to your average reader, maybe I just don't understand vLLM, and am taking it out in good hackers just pushing code. Then, when I go back, I immediately see a comment from someone who does use vLLM noting it already does caching. Sigh.
- nativeit 1y agoThanks for sharing. You certainly aren't alone in your sentiments. I am seeing similar trends in arXiv submissions, as it seems it has become something of a means to inflate the value of one's own product(s) with a veneer of academic rigor. There seems to be a S.O.P. emerging for AI tools that follows many of the same trends as the less-than-reputable blockchain/crypto projects.
- Twirrim 1y ago> I am seeing similar trends in arXiv submissions, as it seems it has become something of a means to inflate the value of one's own product(s) with a veneer of academic rigor Unfortunately this isn't new. Almost as long as people have been publishing papers, people have been using them this way. arXiv, arguably, makes it even worse because the papers haven't even gone through the pretense of a peer review, that does serve to filter out at least some of them.
- nativeit 1y agoVery true, the strategy just preys on a long-established logical fallacy, for as long as humans have been vulnerable to an appeal to authority, we will continue to fall for attempts at "research washing". I know I am subconsciously influenced by the aesthetic of a research paper, regardless of its source or status.
- pama 1y agoI had related questions and checked out the project a bit deeper though I havent tested it seriously yet. The project did start work over a year ago based on relevant papers, before vllm or sglang had decent solutions; it might still be adding performance in some workflows though I havent tested it and some of the published measurements in the project are now stale. Caching LLM kv-cache to disk or external memory servers can be very helpful at scale. Cache management and figuring out cache invalidation is hard anyways and I am not sure at what level a tight integration with inference servers or specialized inference popelines can help vs a lose coupling that could advance each component separately. It would be nice if there were decent protocols used by all inference engines to help this decoupling.
- nativeit 1y agoI'll just be unambiguous about this: > Please don't use HN primarily for promotion. It's ok to post your own stuff part of the time, but the primary use of the site should be for curiosity. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- Aurornis 1y agoA couple months ago another project claimed to have sped up llama.cpp (IIRC) on the front page of HN, from another green name account. It gathered hundreds of GitHub stars and was on the front page all day. When some of us finally had time to look at the code we discovered they didn't invent anything new at all. They took some existing command line options for llama.cpp and then changed the wording slightly to make them appear novel. The strangest part was that everyone who pointed it out was downvoted at first. The first comment to catch it was even flagged away! You couldn't see it unless you had showdead turned on. At first glance I don't see this repo as being in the same category, though the "3X throughput increase" claim is very clearly dependent on the level of caching for subsequent responses and the "lossless" claim doesn't hold up as analyzed by another top-level comment. I think AI self-promoters have realized how easy it is to game Hacker News and GitHub stars if you use the right wording. You can make some big claims that are hard to examine in the quick turnaround times of a Hacker News front page cycle.
- bGl2YW5j 1y agoSame. Maintain skepticism.
- cchance 1y agoI mean a lot of people don't comment on HN, and just use it as a site for cool links lol, so you wouldn't see them posting often