Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
cocogoatmain
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
cocogoatmain
10mo ago
Want to also add that the model doesn’t know how to respond in a user-> assistant style conversation after it’s pretraining, and it’s a pure text predictor (look at the open source base models) There’s also what is being called mid-train
2.
▲
by
cocogoatmain
11mo ago
Provided you had the GPU compute to do so you could train the model to have less refusals, e.g. https://arxiv.org/abs/2407.01376 Quality of response/model performance may change though There’s also nous research’s
3.
▲
by
cocogoatmain
1y ago
128gb unified memory is enough for pretty good models, but honestly for the price of this it is better just go go with a few 3090s or a Mac due to memory bandwidth limitations of this card
4.
▲
by
cocogoatmain
1y ago
I don’t know about the “why” but Alibaba is definitely paying to train open weight models ( https://huggingface.co/Wan-AI )
5.
▲
by
cocogoatmain
1y ago
~~Likely much less than .25/m if that’s mbps. The issue is you’d have no shortage of money at that scale - I run one of the two main Arch Linux package mirrors in my country and while it’s admittedly a quite niche and small distro in c
6.
▲
by
cocogoatmain
1y ago
Assuming you’re a tourist, are you sure it’s not your bank flagging the transaction? When visiting in April I too thought it was an issue with the card when trying to reload on apple wallet, but after two cards declined, found out the tra
7.
▲
by
cocogoatmain
1y ago
Not the original poster but there are some large publicly available dataset such as https://huggingface.co/datasets/allenai/WildChat and https://huggingface.co/datasets/lmsys/lmsys-chat-1