Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
waleedk
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
waleedk
3y ago
[Author] Can you cite some examples of this so we're not talking in the void? I've actually tested Llama 2 for summarization but haven't blogged about that yet, and across multiple domains, Llama 2 is pretty good. I do see s
32.
▲
by
waleedk
3y ago
[Author] Can you share more details? How does it fail? What type of domain? Perhaps there is some tweaking to prompts that's required? It would help to understand what you're seeing.
33.
▲
by
waleedk
3y ago
[Author] Can you elaborate on what you mean here? I'd like to understand what you mean by "blind benchmark" and "not public"?
34.
▲
by
waleedk
3y ago
[Contributor to Aviary] Not many are building their own LLMs (though there are some like Bloomberg). But quite a few are experimenting with fine tuning. One of the great things with LLMs is that they can (and are) fine-tuned with small amou
35.
▲
Aviary simplifies OSS LLM eval and deployment
(github.com)
5 points
by
waleedk
3y ago
|
3 comments
36.
▲
by
waleedk
3y ago
[Author] Good luck trying to use clusters of Lambda machines. Lambda labs are cheap for a reason: their API is not very featureful (we looked at them and we saw they didn't even support machine tagging). If you're looking for a bo
37.
▲
by
waleedk
3y ago
[Author] Mosaic must be getting some kind of sweetheart deals on A100 80GB and A100 40GB. The prices they are quoting are not what say the AWS on-demand prices are. They quote $2 per GPU for A100 40GB and $2.50 for A100 80GB. That's li
38.
▲
by
waleedk
3y ago
[Author] TL;DR OS LLM models are coming. Dolly's not that great -- I've hit lots of issues using it to be honest . MosaicML has a nice commercially usable model here: https://www.mosaicml.com/blog/mpt-7b I th
39.
▲
by
waleedk
3y ago
[Author] You approximate the weights using fewer bits. You also switch to ints instead of floats and then do some fancy stuff when multiplying to make it all work together. More detail than you probably wanted: https://huggingfac
40.
▲
by
waleedk
3y ago
[Author] Fair point -- I clarified the language and gave a concrete example. Hope that helps!
41.
▲
by
waleedk
3y ago
[Author] You're welcome -- glad it was useful!
42.
▲
by
waleedk
3y ago
[Author] Completely disagree. Any analysis shows that you see perplexity reduction at 4 bits. Have a look at llama.cpp's results here: https://github.com/ggerganov/llama.cpp#quantization 4 bit has a perplexity sco
43.
▲
by
waleedk
3y ago
[Author] Fair point. Adjusted the language. Nonetheless people do tend to use 16 bit huggingface models, and if you do go to 8 bits and it's wrong, you're never quite sure if it's the quant or the model.