Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
idonotknowwhy
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
14 ms
·
61.
▲
by
idonotknowwhy
1y ago
Unsloth. Check their colab notebooks
62.
▲
by
idonotknowwhy
1y ago
You mean conversations? Just the jsonl of the standard hf dataset format to import into other systems?
63.
▲
by
idonotknowwhy
1y ago
I didn't realise the 5070 is slower than the 3090. Thanks. If you want a bit more context, try -ctv q8 -ctk q8 (from memory so look it up) to quant the kv cache. Also an imatrix gguf like iq4xs might be smaller with better quality
64.
▲
by
idonotknowwhy
1y ago
Then i guess vocab is the IPC. 10k mistral tokens are about 8k llama3 tokens
65.
▲
by
idonotknowwhy
1y ago
BetaMax had better fidelity than VHS, but initially could only hold 60 minutes of A/V while VHS could do 120m (ie, almost a full movie without swapping tapes)
66.
▲
by
idonotknowwhy
1y ago
I feel like this is being pushed to get more of the system controlled by the provider's side. After a few years, Anthropic, google, etc might start turning off the api. Similar to how Google made it very difficult to use IMAP / SM
67.
▲
by
idonotknowwhy
1y ago
> Or are you saying you'll retroactively hate a song once you figure out it was AI generated? Can't speak for OP, but for me personally, I don't mind AI augmented content, as long as it's done well. Eg. I recently pla
68.
▲
by
idonotknowwhy
1y ago
Have you tried Deepseek-R1? I run it locally and read the raw thought process, find it very useful (can be ruthless at times) seeing this before it tags on the friendliness. Then you can see it's planning process to tag on the warmth&#
69.
▲
by
idonotknowwhy
2y ago
> It's not even per token. The routing happens once per layer, with the same token bouncing between layers. They don't really "bounce around" though do they (during inference)? That implies the token could bounce back
70.
▲
by
idonotknowwhy
2y ago
> tariff on our closest allies I saw that your president said: > “We imported $3b of Australian beef last year alone… they won’t take any of our beef because they don’t want it to affect their farmers.” Just FYI, the reason we don
71.
▲
by
idonotknowwhy
2y ago
It depends what you're trying to do. Coding - Mistral-Small-2503 or Qwen2.5-32b-Coder Reasoning - QwQ-32b Writing - Gemma-3-27b is good at this. etc
72.
▲
by
idonotknowwhy
2y ago
> Like there must be a size too small to hold "all the information" in. We're already there. If you running a Mistral-Large-2411 and Mistral-Small-2409 locally, you'll find the larger model is able to recall more spe
73.
▲
by
idonotknowwhy
2y ago
Holy shit, did they just make every photo taken from an iPhone (AI enhancement) public domain? And spellchecker for text? What about movies like Deadpool3, where AI wrote part of the script?
74.
▲
by
idonotknowwhy
2y ago
The multimodal models aren't good for this. Refusals aren't the issue (they're fine with BERSERK, though occasionally they'll refuse for copyright). The issue is the tech isn't there yet. You'll want to use cus
75.
▲
by
idonotknowwhy
2y ago
Replace "google" with "unsloth" in the browser address bar if you want to download them without signing up to hf
76.
▲
by
idonotknowwhy
2y ago
Thanks a lot for the v2.5! I'll give that a whirl. Hopefully it's as coherent as v3.5 when quantized so small. > I can't quantify the difference between these and the full FP8 versions of DSR1, but I've been playing w
77.
▲
by
idonotknowwhy
2y ago
Oh right, yeah I've done things like this (phone calls to ChatGPT) or the openwebui Whisper -> LLM -> TTS setup. I thought there might be something more than this by now
78.
▲
by
idonotknowwhy
2y ago
Yeah, I've had similar experiences. I still hesitate if it's a field I don't know too well of course (never trust an LLM), but R1 has been able to solve things I've been stuck on. And watching it's <think><
79.
▲
by
idonotknowwhy
2y ago
> This is missing the most interesting changes in generative AI space over the last 18 months I agree, though personally I'm liking the "big thing" as well. R1 is able to one-shot a lot of work for me, churning away in the
80.
▲
by
idonotknowwhy
2y ago
Yeah I'm surprised by all the negativity as well. I'm listening to the post right now (using xtts-v2 finetuned on a voice I like lol). Sounds like these companies are overvalued / over hyped. Maybe they are / some of the
81.
▲
by
idonotknowwhy
2y ago
Did you just ls my /workspace dir? Lol
82.
▲
by
idonotknowwhy
2y ago
Intel had a great P/E a couple of years ago as well :)
83.
▲
by
idonotknowwhy
2y ago
> Unless something radically changed in the last couple years, I am not sure where you got this from? This was the first thing that stuck out to me when I skimmed the article, and the reason I decided to invest the time reading it all. I
84.
▲
by
idonotknowwhy
2y ago
Deepseek was there side project. They had a lot of GPUs from their crypto mining project. Then Ethereum turned off PoW mining, so they looked into other things to do with their GPUs, and started DeepSeek.
85.
▲
by
idonotknowwhy
2y ago
Most people running local inference do so thorough quants with llamacpp (which runs on everything) or awq/exl2/mlx with vllm/tabbyAPI/lmstudio which are much faster to than using pytorch directly
86.
▲
by
idonotknowwhy
2y ago
I've bought 5 used and they're all perfect. But that's what buyer protection on ebay is for. Had to send back an Epyc mobo with bent pins and ebay handled it fine.
87.
▲
by
idonotknowwhy
2y ago
For me personally, hacking together projects as a hobbiest, 2 reasons : 1. It just works. When i tried to build things on Intel Arcs, i spent way more hours bikeshedding ipex and driver issues than developing 2. LLMs seem to have more cuda
88.
▲
by
idonotknowwhy
2y ago
He could probably do 12b (Nemo) up to 14b (Qwen 2.5) at 4bpw with exllamav2
89.
▲
by
idonotknowwhy
2y ago
What did you switch to? I switched to Manjaro from Arch in 2017 because I don't have time to debug / fix broken updates, and the same Manjaro install has been completely stable since then. Is there another distro which can do this
90.
▲
by
idonotknowwhy
2y ago
It's because oai are going for low latency. I've got pause free working locally by running the tts model after the llm has finished generating a huge chunk of the response. And fine tuning on 15 minutes of the voice you're ta
More ›