Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kcorbitt
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
12 ms
·
91.
▲
by
kcorbitt
3y ago
A sibling commenter mentioned the same thing. I actually did the same analysis on "rust" and "remote work", but cut it from the final post for brevity. I've added an addendum now with sentiment on those topics graph
92.
▲
by
kcorbitt
3y ago
(author here) Interesting question to ask! I actually did the exact same sentiment analysis on Rust and remote work as well, but ended up cutting them from the post in the interest of brevity. My recollection is that the sentiment on those
93.
▲
by
kcorbitt
3y ago
(author here) Yes, I was surprised that AI didn't have more positive sentiment overall! Subjectively, the two flavors of AI-negative sentiment I've seen most commonly on HN are (1) its potential to invade privacy, and (2) its pote
94.
▲
Is AI the next crypto? Insights from HN comments
(openpipe.ai)
237 points
by
kcorbitt
3y ago
|
367 comments
95.
▲
by
kcorbitt
3y ago
If you're looking for a practical guide to getting started with fine tuning, I wrote one a couple of months ago that got pretty popular here on HN. Might be helpful if you're interested in playing around with it! https:/
96.
▲
by
kcorbitt
3y ago
I don't know whether Mistral had benchmark data contamination or not, but it's definitely a stronger model than Llama 2 7b or 13b even on non-benchmark tasks. We proactively tested Mistral across all of our customer-deployed model
97.
▲
by
kcorbitt
3y ago
Eh, OpenAI is too cheap to beat at their own game. But there are a ton of use-cases where a 1 to 7B parameter fine-tuned model will be faster, cheaper and easier to deploy than a prompted or fine-tuned GPT-3.5-sized model. In fact, it might
98.
▲
by
kcorbitt
3y ago
Yeah totally agree. We've found that a ton of OpenAI usage in practice is a variant of either classification or information extraction. This makes sense -- going from a human-native form of information (free text) to a computer-native
99.
▲
by
kcorbitt
3y ago
Yep, the 50x cost reduction is if you self-host a fine-tuned model using the setup demonstrated in in the linked notebooks.
100.
▲
by
kcorbitt
3y ago
I wrote the notebooks in the post with the intention of them being a gentle introduction to fine-tuning. Would love any feedback on open questions you have as you go through them!
101.
▲
by
kcorbitt
3y ago
Depends on your use case. If you're doing pure classification then there are smaller encoder-only models like DeBERTa that might get you better performance with a much smaller model size (so cheaper inference). But if you need text gen
102.
▲
by
kcorbitt
3y ago
There's no rule that your fine-tuning dataset needs to be split into input/output pairs -- you can of course fine-tune a model to just continue a sequence. As a practical matter though, most of the fine-tuning frameworks, includin
103.
▲
by
kcorbitt
3y ago
We ran 5K randomly selected recipes through GPT-4 and extrapolated based on the average cost per query.
104.
▲
by
kcorbitt
3y ago
> What makes sense to fine-tune and what not? In general, fine-tuning helps a model figure out how to do the exact task that is being done in the examples it's given. So fine-tuning it on 1000 examples of an API being used in the wi
105.
▲
by
kcorbitt
3y ago
Yes, if you're just using Llama 2 off the shelf (without fine-tuning) I don't think there are a lot of workloads where it makes sense as a replacement for GPT-3.5. The one exception being for organizations where data security is n
106.
▲
by
kcorbitt
3y ago
Nope, no need for few-shot prompting in most cases once you've fine-tuned on your dataset, so you can save those tokens and get cheaper/faster responses!
107.
▲
by
kcorbitt
3y ago
> My other thoughts to extend this are that you could make it seamless. To start, it'll simply pipe the user's requests to OpenAI or their existing model. So it'd be a drop in replacement. Then, it'll every so often o
108.
▲
by
kcorbitt
3y ago
We're finding that when running Llama-2-7B with vLLM ( https://github.com/vllm-project/vllm ) on an A40 GPU we're getting consistently lower time-to-first-token and lower average token generation time than GPT-
109.
▲
by
kcorbitt
3y ago
Depending on what you're trying to accomplish, I'd highly recommend trying the 7B and 13B models first before jumping to the 70B. They're quite capable and I think lots of folks assume they need to jump to a 70B model when re
110.
▲
Fine-tune your own Llama 2 to replace GPT-3.5/4
955 points
by
kcorbitt
3y ago
|
181 comments
111.
▲
by
kcorbitt
3y ago
A lot of people are using RunPod for experimental/small-scale workloads. They have good network and disk speeds and you can generally find availability for a latest-gen GPU like an L40 or 4090 if your workload can fit on a single GPU.
112.
▲
Show HN: Automatically convert your GPT-3.5 prompt to Llama 2
13 points
by
kcorbitt
3y ago
|
2 comments
113.
▲
by
kcorbitt
3y ago
This is fantastic! I found myself nodding along in many places. I've definitely found in practice that evals are critical to shipping LLM-based apps with confidence. I'm actually working on an open-source tool in this space: http
114.
▲
by
kcorbitt
3y ago
It depends -- do you mean as a general end-user of a chat platform or do you mean to include a model as part of an app or service? As an end user, what I've found works in practice is to use one of the models until it gives me an answe
115.
▲
by
kcorbitt
3y ago
When I was evaluating options a few months ago I found https://github.com/PaddlePaddle/PaddleOCR to be a very strong contender for my use case (reading product labels), but you'll definitely want to put together s
116.
▲
by
kcorbitt
3y ago
Kind of a meta-answer, but my personal learning style is "think of something cool to build, then figure out what I need to know to build it." It just so happens that a lot of the interesting/cool stuff going on right now buil
117.
▲
by
kcorbitt
3y ago
Hmm. I admit that I haven't thought about this deeply, but I'm not sure that's true? It seems to me that you could extend the KV cache either backwards or forwards equally easily.
118.
▲
by
kcorbitt
3y ago
I've thought about building this for a while, glad it's out there! Not only does this guarantee your output is JSON, it lowers your generation cost and latency by filling in many of the repetitive schema tokens without passing the
119.
▲
by
kcorbitt
3y ago
While LLMs can be prompted to write in many different styles, especially if you allow them to edit a text over multiple passes, the default "voice" of ChatGPT is surprisingly recognizable. For instance, the comment I'm replyi
120.
▲
by
kcorbitt
4y ago
Check out https://arxiv.org/abs/2303.11366?s=09 ;) We aren't introspecting previous runs or human examples to optimize workflows right now, but it's a powerful tool and one that I expect we'll employ in
More ›