Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
VHRanger
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
61.
▲
by
VHRanger
11mo ago
I'm not saying we could detect it from the text alone! The side channel signals (who posted it, where, etc.) are more valuable in tagging than raw text classifier scores. That's why I said our definition of slop can include all ty
62.
▲
by
VHRanger
11mo ago
> A great deal of LLM-generated content shows up in comments on social media. True, but going after classifying the source (user's commenting patterns) is a better signal than the content itself. That said, for us (Kagi) it's a
63.
▲
by
VHRanger
11mo ago
Yes, I'm the ML lead. The current search engine doesn't go after WordPress plugins we consider correlated to bad pages. By far the most efficient method in the search engine for spam is downranking by trackers/javascript weig
64.
▲
by
VHRanger
11mo ago
> May I ask how you plan to deal with YouTube auto-dubbing videos into crappy AI slop? I'm sorry that's a YouTube problem, not a problem with the original content. Sadly we don't have plans to address that at the moment --
65.
▲
by
VHRanger
11mo ago
Perplexity is a term of art in LLM training, yes. A naive way of scoring how AI laden text is would be to run n-1 layers of a model and compare the text to the probability space of tokens from the model. It works somewhat to detect obvious
66.
▲
by
VHRanger
11mo ago
Give it time, the database is just starting. Give it ~2 weeks to start seeing real impact on your results
67.
▲
by
VHRanger
11mo ago
> I wonder where the obstinacy on the part of certain CEOs come from. I can tell you: their board, mostly. Few of whom ever used LLMs seriousl. But they react to wall street and that signal was clear in the last few years
68.
▲
by
VHRanger
11mo ago
Slop is about thoughtless use of a model to generate output. Output from your paper's model would still qualify as slop in our book. Even if your model scored extremely high perplexity on an LLM evaluation we'd likely still tag it
69.
▲
by
VHRanger
11mo ago
We have rules of thumb and we'll have a more technical blog post on this in ~2 weeks. You can break the AI / slop into a 4 corner matrix: 1. Not AI & Not Slop (eg. good!) 2. Not AI & slop (eg. SEO spam -- we already punish
70.
▲
by
VHRanger
11mo ago
Yes, a fun fact about slop text is that it's very low perplexity text (basically: it's statistically likely text from an LLM's point of view) so most algorithms that rank will tend to have a bias towards preferring this text.
71.
▲
by
VHRanger
11mo ago
Hey, Kagi ML lead here. > Kagi pays for hordes of reviewers? Is this another case of outsourcing moderation to sweat shops in poor countries? No, we're simply not paying for review of content at the moment, nor is it planned. We
72.
▲
by
VHRanger
11mo ago
> AI slop eventually will get as good as your average blogger. Even now if you put an effort into prompting and context building, you can achieve 100% human like results. Hey, Kagi ML lead here. For images/videos/sound, not at
73.
▲
by
VHRanger
1y ago
correct! And preferably a good nvme (not a cheap one without a DRAM cache) and good ram.
74.
▲
by
VHRanger
1y ago
American food would be cruel & unusual punishment
75.
▲
by
VHRanger
1y ago
In memory, and if larger than memory it makes .duckdbtmp files to work from
76.
▲
by
VHRanger
1y ago
> Not good. These tools (from search engines to AI) are increasingly part of our brains, and we should have confidentiality in using them. Don't expect that from products with advertising business models
77.
▲
by
VHRanger
1y ago
I don't disagree with you and I agree that the stereotype of autistic people as lacking empathy is harmful and generally untrue. HOWEVER - the person I was replying to was mentioning them being hurt by the behaviors, and that signals a
78.
▲
by
VHRanger
1y ago
Having had personal experience with this problem, there is actually a simple but difficult solution to your problem: Focus in the effect if the person's action on you, and set hard boundaries for yourself. Note that boundaries are in t
79.
▲
by
VHRanger
1y ago
I love termux, but it's really not a replacement for full fat linux. There's tio many incompatibilities with how linux software expects to run for it to work for a dev workflow (unless your workflow is to immediately ssh into some
80.
▲
by
VHRanger
1y ago
huh, it seems like the M4 pro can hit >400GB/s of RAM bandwidth whereas even a 9950x hits only 100GB/s. I'm curious how that is; in practice it "feels" like my 9950x is much more efficient at "move tons of R
81.
▲
by
VHRanger
1y ago
I mean, huge software with a ton of quirks like a AAA video game are arguably not a good benchmark to understand hardware. They're still good benchmarks IMO because they represent a "real workload" but to understand why the 9
82.
▲
by
VHRanger
1y ago
> Can you explain then, how come switching from Intel MBP to Apple Silicon MBP feels like literally everything is 3x faster, the laptop barely heats up at peak load, and you never hear the fans? Going back to my Intel MBP is like going b
83.
▲
by
VHRanger
1y ago
Woah woah woah AMD promises ROCm will stop being a joke very soon! Maybe this year even!
84.
▲
by
VHRanger
1y ago
Most likely not because of NUMA bottlenecks
85.
▲
by
VHRanger
1y ago
> China is strict with people rioting or complaining a little too much about the government, but they don't lock people up for saying general no no words or being too patriotic/nationalistic online. How absurd is this statement
86.
▲
by
VHRanger
1y ago
I was talking about their sonar API and what you get when you go to perplexity.ai and search by default -- it's llama 3.3 70B hosted with cerebras. Which is fast, but way way behind the times -- it's more likely to hallucinate tha
87.
▲
by
VHRanger
1y ago
Fair enough. I encourage you to cycle through the models in the assistant once in a while to see if there's new ones to like. Kimi-k2 for instance was a positive surprise recently and is being added to your standard plan tonight as wel
88.
▲
by
VHRanger
1y ago
> Soon, there will be nothing of value left for Kagi to index. We have something in the works to fight back on that front, do not be worried ;)
89.
▲
by
VHRanger
1y ago
you can also just go straight to the assistant from the query with !ai -- it'll power the answer directly with your default assistant model
90.
▲
by
VHRanger
1y ago
Try adding !ai to the end of queries once in a while and see if it saves you time. It does for a lot of them.
More ›