Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
andersa
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
18 ms
·
241.
▲
by
andersa
2y ago
> You go on a dating app, a girl responds with her nudes My first assumption would be the "girl" is a bot and the image generated by AI until proven otherwise, i.e. by meeting up at a public location and seeing it's a real
242.
▲
by
andersa
2y ago
I don't understand what would motivate someone to ever send an explicit photo of themselves to a stranger. The whole premise of this scam makes no sense. Like, is this not the most basic of the basics of online safety?
243.
▲
by
andersa
2y ago
That really depends on how much money they are happily paying you to be unproductive.
244.
▲
by
andersa
2y ago
It depends on what you do with it and how much bandwidth it needs between the cards. For LLM inference with tensor parallelism (usually limited by VRAM read bandwidth, but little exchange needed) 2x 4090 will massively outperform a single A
245.
▲
by
andersa
2y ago
Where are you getting an A100 80GB for $10k?
246.
▲
by
andersa
2y ago
That is very interesting if tinygrad can support it! Every other library I've seen had the limitation on dividing the heads, so I'd (perhaps incorrectly) assumed that it's a general problem for inference.
247.
▲
by
andersa
2y ago
It can make a difference when using tensor parallelism to run small batch sizes. Not a huge difference like training because we don't need to update all weights, but still a noticeable one. In the current inference engines there are so
248.
▲
by
andersa
2y ago
2 GPUs works fine too, as long as your model fits. Using different GPUs with same VRAM however, is highly highly sketchy. Sometimes it works, sometimes it doesn't. In any case, it would be limited by the performance of the slower GPU.
249.
▲
by
andersa
2y ago
It should be possible with onboard PCIe switches. You probably don't need the networking or storage to be all that fast while running the job, so it can dedicate almost all of the bandwidth to the GPU. I don't know if there are bo
250.
▲
by
andersa
2y ago
It was probably just before running LLMs with tensor parallelism became interesting. There are plenty of other workloads that can be divided by 6 nicely, it's not an end-all thing.
251.
▲
by
andersa
2y ago
The split is done automatically by the inference engine if you enable tensor parallelism. TensorRT-LLM, vLLM and aphrodite-engine can all do this out of the box. The main thing is just that you need either 4 or 8 GPUs for it to work on curr
252.
▲
by
andersa
2y ago
That calculation is incorrect. You need to fit both the model (140GB) and the KV cache (5GB at 32k tokens FP8 with flash attention 2) * batch size into VRAM. If the goal is to run a FP16 70B model as fast as possible, you would want 8 GPUs
253.
▲
by
andersa
2y ago
Sure, it's also at least an order of magnitude slower in practice, compared to 4x 4090 running at full speed. We're looking at 10 times the memory bandwidth and much greater compute.
254.
▲
by
andersa
2y ago
For example, if you want to run low latency multi-GPU inference with tensor parallelism in TensorRT-LLM, there is a requirement that the number of heads in the model is divisible by the number of GPUs. Most current published models are divi
255.
▲
by
andersa
2y ago
Incredible! I'd been wondering if this was possible. Now the only thing standing in the way of my 4x4090 rig for local LLMs is finding time to build it. With tensor parallelism, this will be both massively cheaper and faster for infere
256.
▲
by
andersa
2y ago
Price?
257.
▲
by
andersa
3y ago
Interesting... you're right, this is giving me lots of wednesdays.
258.
▲
by
andersa
3y ago
The "new" outlook is a disgrace. Using it to access a third party email provider results in Microsoft transferring your credentials to their server so the server can sync all your emails... for reasons, of course.
259.
▲
by
andersa
3y ago
Wait, what? I found thunderbird search to be a super power, so much more reliable (and fast) than outlook.
260.
▲
by
andersa
3y ago
So why don't you stop and think that maybe phones should not do any of these things? Would you like to be thrown in jail because someone's phone's "ai to enhance faces" made it look like you accidentally?
261.
▲
by
andersa
3y ago
That's not what they're doing. This was using one of those generative ai hallucination tools that "add detail" to a video to make it look fancier. Something like that clearly has no place in a courtroom.
262.
▲
by
andersa
3y ago
I'm glad common sense is still alive, this is the correct precedent to set.
263.
▲
by
andersa
3y ago
It's the writing style. Humans simply don't write like that, unless they are writing a research paper.
264.
▲
by
andersa
3y ago
Good point, that would make sense. But same difference. Why is that not how it works?
265.
▲
by
andersa
3y ago
Why is this a thing? Can't packages use specific tags from the git repo? It seems so incredibly stupid to allow this, throwing out all of the "oh but it's open source you can review it" arguments in one go if the source
266.
▲
by
andersa
3y ago
It'd make total sense if it was. This way you get to have the backdoor without your enemies being able to use it against your own companies.
267.
▲
by
andersa
3y ago
Yes, I don't understand why they are using a search benchmark for these... it would be much better to have something like giving it a story up to the context length (from a book? how to find one that's not in the training data?)
268.
▲
by
andersa
3y ago
What makes you think they aren't already?
269.
▲
by
andersa
3y ago
What if we could add giant engines to control the rotation speed and fix this leap day/second nonsense once and for all?
270.
▲
by
andersa
3y ago
Those typically require custom client side code, for a website you have the requirement that a web browser must be able to connect to it using TLS. Or maybe I'm not getting what your suggestion is - Access is supposed to intercept the
More ›