Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
coder543
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
22 ms
·
481.
▲
by
coder543
3y ago
You can also choose to run at 4-bit quantization, offloading ~27 out of 33 layers to the GPU, and that runs at about 25 tokens/s for me. I think that's about the same speed as you get out of an M1 Max running at 4 bits? Although I
482.
▲
by
coder543
3y ago
Mixtral works great at 3-bit quantization. It fits onto a single RTX 3090 and runs at about 50 tokens/s. The output quality is not "ruined" at all. For the amount of money you're talking about, you could also buy two 309
483.
▲
by
coder543
3y ago
CogVLM is very good in my (brief) testing: https://github.com/THUDM/CogVLM The model weights seem to be under a non-commercial license, not true open source, but it is "open access" as you requested. It would
484.
▲
by
coder543
3y ago
I just need a WiFi 7 Access Point. I don't need the rest of Ubiquiti's stack. Fortunately, Ubiquiti supports (basic) standalone operation for access points: https://help.ui.com/hc/en-us/articles/1259
485.
▲
by
coder543
3y ago
The chart you pointed out is very interesting, but it largely supports my point. The blue line is easiest to read, so let’s look at how the tokens/sec scale for a single user session as the batch size increases. It starts out at about
486.
▲
by
coder543
3y ago
> I wonder are you using a quantized version of Mistral? Yes, we’re comparing phone performance versus datacenter GPUs. That is the discussion point I was responding to originally. That person appeared to be asking when phones are going
487.
▲
by
coder543
3y ago
> These data center targeted GPUs can only output that many tokens per second for large batches. No… my RTX 3090 can output 130 tokens per second with Mistral on batch size 1. A more powerful GPU (with faster memory) should easily be abl
488.
▲
by
coder543
3y ago
4-bit StableLM and 2-bit 7B models do seem to be working more consistently.
489.
▲
by
coder543
3y ago
I just checked and MLC Chat is running the 3-bit quantized version of Mistral-7B. It works fine on the 14 Pro Max (6GB RAM) without crashing, and is able to stay resident in memory on the 15 Pro Max (8GB RAM) when switching with another not
490.
▲
by
coder543
3y ago
EDIT: Attempting to converse with any Q4_K_M 7B parameter model on a 15 Pro Max... the phone just melts down. It feels like it is producing about one token per minute. MLC-Chat can handle 7B parameter models just fine even on a 14 Pro Max,
491.
▲
by
coder543
3y ago
What does snappier even mean in this context? The latency from connecting to a server over most network connections isn’t really noticeable when talking about text generation. If the server with a beefy datacenter-class GPU were running the
492.
▲
by
coder543
3y ago
The linked page claims that LiteLlama scored a zero on the GSM8K benchmark, so let's just say math probably isn't its forte.
493.
▲
by
coder543
3y ago
For example, I asked Mixtral to generate 4 questions and short answers following a prompt format that I provided. Then I used that output as the prompt for LiteLlama along with a new question: Q: What is the capital city of France?
494.
▲
by
coder543
3y ago
This model does not appear to be fine-tuned for chat. I observed the same looping behavior with virtually any direct prompt. If I prime it with a pattern of Q and A with several examples of a good question and a good answer, then a final qu
495.
▲
by
coder543
3y ago
Don’t forget that oxygen is useful for breathing too. Semi-important, I hear. ChatGPT estimates that if we converted all of earth’s atmospheric oxygen into water (via burning with hydrogen), ocean levels would rise about 3.7 meters: https:
496.
▲
by
coder543
3y ago
Teachable Machine has existed for years : https://www.theverge.com/tldr/2017/10/9/16447006/google-teac... Its last real update (AFAIK) was in 2019: https://www.theverge.com/2019
497.
▲
by
coder543
3y ago
Being multimodal doesn’t seem to require much of a size penalty: https://github.com/dlyuangod/TinyGPT-V Even so, Google treats the Gemini Pro Vision model as a separate model from Gemini Pro, so it could have separate
498.
▲
by
coder543
3y ago
Not the person you replied to, but… Microsoft apparently revealed that GPT-3.5 Turbo is 20 billion parameters. Gemini Pro seems to perform only slightly better than GPT-3.5 Turbo according to some benchmarks, and worse in others. If Gemini
499.
▲
by
coder543
3y ago
>> If the federal government were interested in passing a law that required this, I'm sure the Library of Congress could run such a server, but no such law exists. > Require by law all who desire copyright protection to regist
500.
▲
by
coder543
3y ago
Okay, then I set the date on my computer to year the 2124. Now the DRM is unlocked! You can't build a "time lock" with just encryption primitives. Even if you could build a time lock with just encryption primitives, we don
501.
▲
by
coder543
3y ago
I think "visitors/users can edit" (not just the site admin) is also an essential part of the definition. https://en.wikipedia.org/wiki/Wiki https://www.merriam-webster.com/dictionary/
502.
▲
by
coder543
3y ago
It's important to think about threat vectors. A general concept like "the password manager getting compromised" is not really a threat vector, it's more the outcome of a threat vector. How exactly do you think a passwo
503.
▲
by
coder543
3y ago
No, it’s not theater. 2FA was not created as a defense against password manager compromise. That is not its purpose. It protects against password reuse attacks and helps to protect against total compromise of people who have been phished. E
504.
▲
by
coder543
3y ago
I wouldn’t provide the grammar itself directly, since I feel like the models probably haven’t seen much of that kind of grammar during training, but just JSON examples of what success and error look like, as well as an explanation of the ta
505.
▲
by
coder543
3y ago
No, the grammar can do OR statements. You provide two grammars, essentially. You always want to tell the model about the expected response formats, so that it can provide the best response it can, even though you’re forcing it to fit the gr
506.
▲
by
coder543
3y ago
Similarly, I think it is important to provide an “|” grammar that defines an error response, and explain to the model that it should use that format to explain why it cannot complete the requested operation if it runs into something invalid
507.
▲
by
coder543
3y ago
In the context of this thread, I believe even a digital computer would have to be rebuilt if the program is wrong... :P Unless you typically salvage digital computers from the wreckage of a failed rocket test and stick it in the next protot
508.
▲
by
coder543
3y ago
> (1) You pay for Prime. (2) You pay extra for Prime Video. (3) You now pay even more extra for Prime Video to not show you ads. There is no triple dipping occurring here. Prime Video is included with the normal Prime membership under
509.
▲
by
coder543
3y ago
Blu-ray M-DISC has been available since 2013, I believe. One article from 2013 mentioning Blu-ray in the context of M-DISC: https://www.zdnet.com/article/torture-testing-the-1000-year-... So, 10 years would be a good s
510.
▲
by
coder543
3y ago
If you're asking ChatGPT about its own characteristics, you should not believe the responses. Models can't examine themselves, so unless the model was trained on specific information about itself, or unless the info is put into
More ›