Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
spikedoanz
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
spikedoanz
3y ago
right. I opted for a used 3090 myself and plan to get a 2nd one soon. At current market prices 2x 3090s is cheaper than a single 4090 and provides double vram with more performance. If fine-tuning/lora-ing and energy efficiency is a co
2.
▲
by
spikedoanz
3y ago
If you have a lot of money (but not H100/A100 money), get 4090s as they're currently the best bang for your buck on the CUDA side (according to George Hotz). If broke, get multiple second hand 3090s. https://timdettmers
3.
▲
by
spikedoanz
3y ago
To the best of my knowledge, a combination of special flags and fine tuning is used to let models "know" when to stop generating output: Basically during pre-training/fine-tuning, for every sequence you'd like your model
4.
▲
by
spikedoanz
3y ago
You're correct with interpreting how the model works wrt it returning tokens one at a time. The model returns one token, and the entire context window gets shifted right by one to for account it when generating the next one. As for mod
5.
▲
by
spikedoanz
3y ago
as a side note. I'm aware of the approaches using Langchain/CoT to crawl through the document to search for content. This also seems less than ideal due to a lot of processing at inference time. I'm moreso looking for a zero-
6.
▲
Ask HN: Has anyone tried recursive summarization of documents using LLMs
2 points
by
spikedoanz
3y ago
|
1 comments