4 ms·
This is why local AI is so important
by reactordev 5mo ago
This is why local AI is so important
- weird-eye-issue 5mo agoLocal AI models pull in search results just like ChatGPT does ... And they are trained on web data just like any other model...
- FergusArgyll 5mo agoHow does that help if it's using search? You get whatever the search engine outputs
- rplnt 5mo agoThat doesn't solve this particular problem. Your local model was trained on reddit comments written by bots.
- bayindirh 5mo agoIt's already being trained on "public" (ethical or otherwise) data. So, it already has ingested that kind of "optimization" during pre-training and training. I don't think you can fine-tune your way out of it.
- fsflover 5mo agoThis is far from widespread at the moment, so it'll be possible to at least use the current cutting-edge models locally in the future.
- bayindirh 5mo agoFar from widespread? SEO has seeped to all crevices of the internet for the last 20 years.
- fsflover 5mo agoBy this measure, any information you can get whatsoever is biased and there is no reason to trust anything at all.
- latexr 5mo agoThe major difference is that right now when you land on a page you can do your due diligence and decide if you trust the source. You can still be tricked, but it’s harder and you can get better at the detection. With LLMs, everything is given the same importance so you have no idea if the data came from a reputable source or an obvious SEO junk website.
- fsflover 5mo agoAI can also provide the sources. And if you need to be certain, you should ask for that.
- latexr 5mo agoThose sources are often made up (pages never existed) or outright wrong (they don’t say what the LLM claimed). Asking an LLM for sources is about as (in)effective as telling it to not make stuff up.
- fsflover 5mo agoYou seem to be much behind the progress in LLMs. Modern ones provide correct, verifiable citations just fine.
- latexr 5mo agoAll citations are verifiable, even the wrong ones, and they always were. And no, modern ones don’t “do it just fine”, they are still frequently wrong. Either you’ve been incredibly lucky, or have just stopped verifying thoroughly. But if you’re that confident, please do share the exact models which you trust. Proponents always shift goal posts (somehow, every release in perpetuity, those ones are the good ones and everything before were garbage) so let’s avoid the vagueness.
- ToucanLoucan 5mo agoPeople still think these things are smart. That if their word generator eats enough of the Internet, it will somehow give them the real information that's otherwise hidden. Or perhaps a better word; filter the bullshit. To filter bullshit it would first have to understand bullshit, and it doesn't. That's why an LLM will tell you the solution to a problem that doesn't work, and argue with you when you correct it.
- bayindirh 5mo agoThis is what bothers me a lot. For the people who doesn't know how it's made or want to believe, it's a miracle. For me, it's a resource wasting text generator. I'll not lie, I don't use OpenAI, Mistral or Anthropic's models, even for coding. I prefer to read my API docs and cry once. I used Gemini, five or six times in total. Twice I asked a couple of very specific things, and it unearthed them. Since they were not products, but information, that was helpful. Twice, it has given wrong information. When I "told" it, there was another way, it said "of course there are two ways", etc. Tasteless and time wasting. I don't like using an LLM all day long, or offload my thinking to them. It's the ultimate self-poisoning incident. And as you say, these algorithms can't know right/wrong/logical/bullshit, etc. They just spew out text.
- latexr 5mo agoSomething I’ve also seen multiple times is an LLM giving wrong information, I tell it it’s not right, then it tells me I’m “absolutely right” and it provides the exact same answer and tells me that one will work.
- reactordev 5mo agoOh Gemini, how no one uses you enough…
- satvikpendem 5mo agoI was just reading another post yesterday and your comment reminds me of this one [0], same sort of format and experience of the submitted article of the HN post that comment is on. [0] https://news.ycombinator.com/item?id=48211730 https://news.ycombinator.com/item?id=48211730
- jondea 5mo agoIt's less compromised, but it's still basing the answer on compromised queries. This is why I pay for independent reviews (e.g Which) where their incentives are more aligned with yours.
- Schweigerose 5mo agoHow do you make sure that the model you run locally is not tainted? Is there even a way to confirm this without providing the complete training set?
- psb5 5mo agoFwiw I just run kiwix/zeal locally which has old school search index of all articles in wiki/stackoverflow etc. That seems enough for most of my day to day use.
- soloto 5mo agoLocal AI will have the bias that existed at the time of its training, which is different from no bias. For stuff that needs to be current, a local LLM would need to search the net regardless.
- embedding-shape 5mo agoAnd since "no bias" isn't something that actually exists in reality when it comes to language or even anything near humans, "bias in local model I can introspect" will always be miles ahead of "bias I know is there, but cannot introspect".
- soloto 5mo agoAgreed in full.
- rdtsc 5mo agoNot if the models come from Google. The ads will be implicit in the model. X is better that Y an Z would be easy to add to a the training set.
- pautasso 5mo agoDoes this mean the model must be retrained every time a new ad is posted? How much are AI ads going to cost?
- rdtsc 5mo agoYeah, I meant not individual ads but implicit forced/influenced preference for certain brands. Let’s say it always picks Coke vs Pepsi when giving an example of a soft drink. Or picks BMW when asked to pick the best car. Which cloud provider is the best? -Why, GCP of course, etc. Companies then get to bid for a preference “place”. This is more like Google paying to be the search engine default in Firefox.
- deleted 5mo ago[deleted]