Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
tarruda
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
13 ms
·
181.
▲
by
tarruda
2y ago
https://huggingface.co/blog/mlabonne/abliteration
182.
▲
by
tarruda
2y ago
> Download as many LLM models and the latest version of Ollama.app and all its dependencies. I recently purchased a Mac Studio with 128gb RAM for the sole purpose of being able to run 70b models at 8-bit quantization
183.
▲
by
tarruda
2y ago
Small models retain much less of the knowledge they were trained on, especially when quantized. One good use case for 32gb Mac is being able to run 8b models at full precision, something that is not possible with 8-16gb macs
184.
▲
by
tarruda
2y ago
I recommend installing tailscale client on your devices instead of carrying an additional device/router
185.
▲
by
tarruda
2y ago
They didn't add a comparison to Qwen 2.5 3b, which seems to surpass Ministral 3b MMLU, HumanEval, GSM8K: https://qwen2.org/qwen2-5/#qwen25-05b15b3b-performance These benchmarks don't really matter that much,
186.
▲
by
tarruda
2y ago
> need to pay meta if you exceed 700 million monthly users Seems like a good problem to have
187.
▲
by
tarruda
2y ago
Isn't 3b the kind of size you'd expect to be able to run on the edge? What is the point of using 3b via API when you can use larger and more capable models?
188.
▲
by
tarruda
2y ago
> voting in brazil are very trustworthy How can a closed system that cannot be audited be considered trustworthy? After the voting happens, there's no physical proof of the vote. Highly recommend reading this: https://dfa
189.
▲
by
tarruda
2y ago
405b even at low quants would have very low tokens generation speed, so even if you got the 192GB it would probably not be a good experience. I think 405b is the kind of model that only makes sense to run in clusters of A100/H100. IMO
190.
▲
by
tarruda
2y ago
You can also add an extra 50gb of space to pay like $5/month, that way you are paying and it is still an insanely better deal than any of the other cloud providers
191.
▲
by
tarruda
2y ago
AFAIK Nothing seems to beat Oracle cloud: https://www.oracle.com/cloud/costestimator.html For compute: "Each tenancy gets the first 3,000 OCPU hours and 18,000 GB hours per month for free to create Ampere A1 Compu
192.
▲
by
tarruda
2y ago
You can get refurbished Mac Studio M1 Ultra with 128GB VRAM for ~ $3k on ebay. M1 ultra has 800GB/s memory bandwidth, same as the M2 ultra. Not sure if 128GB VRAM is enough for running 405b (maybe at 3-bit quant?), but it seems to offe
193.
▲
by
tarruda
2y ago
Awesome project, thanks for sharing!
194.
▲
by
tarruda
2y ago
Sources?
195.
▲
by
tarruda
2y ago
Seems like Reflection 70b was an attempt to implement the same concept on top of Llama 3 70b
196.
▲
by
tarruda
2y ago
Maybe the benchmark results are different, but it certainly seems like OpenAI is doing the same with it's "thinking" step
197.
▲
by
tarruda
2y ago
Is it possible someone within OpenAI leaked the CoT technique used in O1, and Reflection 70b was an attempt to replicate it?
198.
▲
by
tarruda
2y ago
I would love to see Meta releasing CoT specialized model as a LoRa we can apply to existing 3.1 models
199.
▲
by
tarruda
2y ago
It seems this is the problem with most benchmarks, which is why benchmark performance doesn't mean much these days.
200.
▲
by
tarruda
2y ago
Have you ran the model in full FP16? It is possible a lot of performance is lost when running quantized versions.
201.
▲
by
tarruda
2y ago
Since most LLMs are released as FP16, just the number of parameters is enough to know the total required GPU RAM.
202.
▲
by
tarruda
2y ago
> Having no integer types (ok, this isn't something typescript could just implement) other than BigInt is another big one for me. Is that a performance thing? I believe JavaScript VMs can specialize/optimize Numbers and BigInts
203.
▲
by
tarruda
2y ago
> I'm looking forward to something like this, but local, FOSS and for Linux. This will probably happen soon, but I wonder what are the disk space requirements for saving screenshots of everything you do
204.
▲
by
tarruda
2y ago
I wonder if magic-wormhole could be implemented as a layer on top of Syncthing: - Generate a short code - Use the code as the seed to deterministically generate a Syncthing device key + config Since the Syncthing device key could be genera
205.
▲
by
tarruda
2y ago
How well does it run in AMD GPUs these days compared to Nvidia or Apple silicon? I've been considering buying one of those powerful Ryzen mini PCs to use as an LLM server in my LAN, but I've read before that the AMD backend (ROCm
206.
▲
by
tarruda
2y ago
Does that mean it is possible to use D&D ruleset in a cRPG without paying some kind of license fee?
207.
▲
by
tarruda
2y ago
Is this a problem in distributions that use musl as the system libc (Alpine) ?
208.
▲
by
tarruda
2y ago
Not related to the main thread, but I want to mention that as a display server, Arcan is superior to everything else out there (x11, Wayland and whatever proprietary software is used in Mac/windows). To make things even more impressive
209.
▲
by
tarruda
2y ago
Why is that the case? Are there no good open source alternatives to these tools?
210.
▲
by
tarruda
2y ago
> Get a real job and make money with the skills and knowledge I acquired. Do you mind sharing what kind of job is that?
More ›