Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
coder543
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
21 ms
·
421.
▲
by
coder543
3y ago
Fascinating, I wouldn’t have expected that!
422.
▲
by
coder543
3y ago
Sound came back later in the video, so I’m pretty sure they just turned the volume way down. The sound was probably very annoying for that period of time, but they actually wanted people to be able to enjoy the beauty of that scene.
423.
▲
by
coder543
3y ago
https://ollama.com/library/gemma/tags You can see the various quantizations here, both for the 2B model and the 7B model. The smallest you can go is the q2_K quantization of the 2B model, which is 1.3GB, but I wou
424.
▲
by
coder543
3y ago
I'm not sure what you mean that it "forgot" about POST? Even as an experienced Go developer, I looked at the code and thought it would probably work for both GET and POST. I couldn't easily see a problem, yet I had not f
425.
▲
by
coder543
3y ago
The first link definitely says Phind-34B on my browser.
426.
▲
by
coder543
3y ago
“RoadTripper”? Or “RoundTripper”?
427.
▲
by
coder543
3y ago
An existing solution: > PMTiles readers use HTTP Range Requests to fetch only the relevant tile or metadata inside a PMTiles archive on-demand. https://docs.protomaps.com/pmtiles/
428.
▲
by
coder543
3y ago
"Hallucination" is a widely understood and accepted term in the LLM industry. If you want to change that, replying to my comment doesn't seem to be the most effective place to start. I don't know that it's the best
429.
▲
by
coder543
3y ago
https://www.clearlypayments.com/blog/interchange-fees-by-cou... "If your business wants to accept credit cards, you’ll need to pay a fee. Interchange makes up the bulk of that cost which merchants pay, roughly 75%
430.
▲
by
coder543
3y ago
You seem to be quoting ChatGPT, and you're not even specifying that it's ChatGPT-4, so I automatically assume ChatGPT-3.5, which hallucinates at an astonishing rate. Regardless, all current LLMs can hallucinate. As ChatGPT's
431.
▲
by
coder543
3y ago
Specifically for visionOS, WorldAnchors: https://developer.apple.com/documentation/arkit/worldanchor And their locations will be remembered (by ID) if an app wants to offer persistent experiences at multiple locat
432.
▲
by
coder543
3y ago
From the article: “I’ve used it for hours at a time without any discomfort, but fatigue does set in, from the weight alone. You never forget that you’re wearing it.” “In terms of resolution, Vision Pro is astonishing. I do not see pixels, e
433.
▲
by
coder543
3y ago
Yep, I seriously considered a Mac Studio a few months ago when I was putting together an “AI server” for home usage, but I had my old 3090 just sitting around, and I was ready to upgrade the CPU on my gaming desktop… so then I had that desk
434.
▲
by
coder543
3y ago
I think that’s too simplified. The best LLMs will still frequently make mistakes. Meta is advertising a HumanEval score of 67.8%. In a third of cases, the code generated still doesn’t satisfactorily solve the problem in that automated bench
435.
▲
by
coder543
3y ago
Fan noise isn’t very much, and you can always limit the max clockspeeds on a GPU (and/or undervolt it) to be quieter and more efficient at a cost of a small amount of performance. The RTX 3090 still seems to be faster than the M3 Max
436.
▲
by
coder543
3y ago
I don’t use local LLMs for CoPilot-like functionality, but I have toyed with the concept. There are a few things to keep in mind: no programmer that I know is sitting there typing code for hours at a time without stopping. There’s a lot mor
437.
▲
by
coder543
3y ago
Already there, it looks like: https://ollama.ai/library/codellama (Look at “tags” to see the different quantizations)
438.
▲
by
coder543
3y ago
Even the M3 Max seems to be slower than my 3090 for LLMs that fit onto the 3090, but it’s hard to find comprehensive numbers. The primary advantage is that you can spec out more memory with the M3 Max to fit larger models, but with the exc
439.
▲
by
coder543
3y ago
> The great thing about gguf is that it will cross to system RAM if there isn't enough VRAM. No… that’s not such a great thing. Helpful in a pinch, but if you’re not running at least 70% of your layers on the GPU, then you barely ge
440.
▲
by
coder543
3y ago
I think your take is a bit optimistic. I like quantization as much as the next person, but even the 2-bit model won’t fit entirely on a 4090: https://huggingface.co/TheBloke/Llama-2-70B-GGUF I would be uncomfortable re
441.
▲
by
coder543
3y ago
The impact on accuracy is somewhere in the single-digit percentages at 4-bit quantization, from what I’ve been able to gather. Very small impact. To draw the analogy out further, if the model was able to get an A on a test before quantizati
442.
▲
by
coder543
3y ago
On a 33B model at q4_0 quantization, I’m seeing about 36 tokens/s on the RTX 3090 with all layers offloaded to the GPU. Mixtral runs at about 43 tokens/s at q3_K_S with all layers offloaded. I normally avoid going below 4-bit quan
443.
▲
by
coder543
3y ago
You don’t stop being andy99 just because you’re a little tired, do you? Being tired makes everyone a little less capable at most things. Sometimes, a lot less capable. In traditional software, the same program compiled for 32-bit and 64-bit
444.
▲
by
coder543
3y ago
Yes. Quantization does not reduce the number of parameters. It does not re-train the model.
445.
▲
by
coder543
3y ago
Quantization is highly effective at reducing memory and storage requirements, and it barely has any impact on quality (unless you take it to the extreme). Approximately no one should ever be running the full fat fp16 models during inference
446.
▲
by
coder543
3y ago
It’s cool that progress is being made on alternative LLM architectures, and I did upvote the link. However, I found this article somewhat frustrating. Showing the quality of the model is only half of the story, but the article suddenly ends
447.
▲
by
coder543
3y ago
From the HN Guidelines: “Please don't use HN primarily for promotion. It's ok to post your own stuff part of the time, but the primary use of the site should be for curiosity.” That user almost exclusively links to what appears to
448.
▲
by
coder543
3y ago
> comes with a heavy runtime (node or python) Ollama does not come with (or require) node or python. It is written in Go. If you are writing a node or python app, then the official clients being announced here could be useful, but they a
449.
▲
by
coder543
3y ago
Ollama is built around llama.cpp, but it automatically handles templating the chat requests to the format each model expects, and it automatically loads and unloads models on demand based on which model an API client is requesting. Ollama a
450.
▲
by
coder543
3y ago
Yes, I need the best quality that is available if I'm going to watch something. It's non-negotiable to me. That is a personal choice, like buying 4K Blu-rays of most things that I like. If Netflix would sell 4K blu-rays of all of
More ›