Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kcorbitt
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
61.
▲
by
kcorbitt
2y ago
This specific variant "LoRA+" described in this paper is even harder to search for. I was doing some research on this technique recently and it turns out that "Lora+" matches with "Lora" in Discord search, whic
62.
▲
by
kcorbitt
2y ago
There are lots of third-party providers that will host your fine-tuned model for you, and just charge per token like OpenAI. Here are some of the providers I've personally used and would vouch for in production, along with their costs
63.
▲
What we've learned in 3 days of Llama 3
(openpipe.ai)
3 points
by
kcorbitt
2y ago
|
0 comments
64.
▲
by
kcorbitt
3y ago
Btw, if you've tried fine-tuning OpenAI models before January and came away unimpressed with the quality of the finished model, it's worth trying again. They made some unannounced changes in the last few months that make the fine-
65.
▲
by
kcorbitt
3y ago
Just to correct the record a bit, while you certainly can spend $1K+ to fine-tune a model we've found that the average user on openpipe.ai spends less than $30 per fine-tune. It actually takes quite a lot of data (likely more than th
66.
▲
by
kcorbitt
3y ago
IMO it's possible to over-generalize from this datapoint (lol). While it's true that creating a general "finance" model that's stronger than GPT-4 is hard, training a task-specific model is much easier. Eg. "a
67.
▲
by
kcorbitt
3y ago
In my experience this kind of mass exodus is more associated with failures in leadership/management than an inevitable result of success. For example, OpenAI has been far more successful than Stability and its senior employees obviousl
68.
▲
by
kcorbitt
3y ago
I don't know... I lived in Barcelona for two years recently and had pretty bad luck. In addition to one-off incidents, some consistent noise issues we had: - The Pakistani restaurant below our apartment would regularly host receptions&
69.
▲
by
kcorbitt
3y ago
So sounds like the real news is that Microsoft basically acquihired the Inflection founding team?
70.
▲
by
kcorbitt
3y ago
That's interesting! Would that still involve each worker node needing to have Nodejs installed to run the process that actually reads from the queue? That's doable, but makes the deployment story a little more annoying/compli
71.
▲
by
kcorbitt
3y ago
Boxes-wise, I'd like a management interface at least as good as the one Sidekiq had in Rails for years. Would also need some hard numbers around performance and probably a bit more battle-testing before using this in our current produc
72.
▲
by
kcorbitt
3y ago
I love your vision and am excited to see the execution! I've been looking for exactly this product (postgres-backed task queue with workers in multiple languages and decent built-in observability) for like... 3 years. Every 6 months
73.
▲
by
kcorbitt
3y ago
I once interviewed a candidate for an engineering position in Seattle. It quickly turned out that he had fabricated his entire work history and education. My first clue was that he claimed two years of work experience at Costco HQ in Issaqu
74.
▲
Mixtral Curious? Comparing Mistral 7B and Mixtral for fine-tuning
(openpipe.ai)
1 points
by
kcorbitt
3y ago
|
0 comments
75.
▲
by
kcorbitt
3y ago
I don't think they've released a fine-tuning API, but we'll definitely support it once they do!
76.
▲
by
kcorbitt
3y ago
Obviously talking my own book here, but we've helped dozens of customers make the transition from prompted GPT-4 or GPT-3.5 to their own fine-tuned models at OpenPipe. The most common reaction I get is "wow, I didn't expect t
77.
▲
by
kcorbitt
3y ago
This actually already exists! We did a writeup of the relevant optimizations here: https://openpipe.ai/blog/s-lora
78.
▲
by
kcorbitt
3y ago
OpenPipe | ONSITE (Seattle) | Full Time Hi HN! I'm Kyle, the founder of OpenPipe, which you might recognize from some of our recent HN threads: - Mistral 7B Fine-Tune Optimized https://news.ycombinator.com/item?id=38712
79.
▲
S-LoRA: Serving Thousands of Models from One GPU for Fun and Profit
(openpipe.ai)
1 points
by
kcorbitt
3y ago
|
0 comments
80.
▲
by
kcorbitt
3y ago
OpenPipe | ONSITE (San Francisco or Seattle) | Full Time Hi HN! I'm Kyle, the founder of OpenPipe, which you might recognize from some of our recent HN threads: - Mistral 7B Fine-Tune Optimized https://news.ycombinator.com&#
81.
▲
by
kcorbitt
3y ago
(Post author here). Totally fair concern. I'll find some representative examples on a sample task we've done some fine-tuning on and add them to the post. EDIT: Ok so the prompt and outputs are long enough that adding them to the
82.
▲
by
kcorbitt
3y ago
Hey, I'm the post author. This is a totally fair point! I do think though that depending on your specific requirements open-source models can be a 10x+ improvement. For example, we serve Mistral 7B for less than 1/10th the cost
83.
▲
by
kcorbitt
3y ago
A good way to build intuition for how much text fits in a token is by pasting a block of text into a tokenizer playground, like this one: https://huggingface.co/spaces/Xenova/the-tokenizer-playgroun...
84.
▲
by
kcorbitt
3y ago
For short context tasks looks like it's slightly stronger than Llama 7B and slightly weaker than Mistral 7B. Really impressive showing for a completely new architecture. I've also heard that it was trained on far fewer tokens than
85.
▲
by
kcorbitt
3y ago
It'll be on Huggingface soon. This is how they dropped their original 7B model as well. It's a marketing thing, but it works!
86.
▲
by
kcorbitt
3y ago
No public statement from Mistral yet. What we know: - Mixture of Experts architecture. - 8x 7B parameters experts (potentially trained starting with their base 7B model?). - 96GB of weights. You won't be able to run this on your home G
87.
▲
by
kcorbitt
3y ago
The problem is that most non-OpenAI models haven't actually been fine-tuned with function calling in mind, and getting a model to output function-calling-like syntax without having been trained on it is quite unreliable. There are a fe
88.
▲
by
kcorbitt
3y ago
For a previous startup a couple of years ago I tried to get our infra onboarded with terraform, hit a bunch of blockers, switched to Pulumi, and have used it for every project big or small since. Hits a really nice sweet spot for me -- yes,
89.
▲
by
kcorbitt
3y ago
If they lose all the employees and then voluntarily give up their Microsoft funding the only asset they'll have left are the movie rights. Which, to be fair, seem to be getting more valuable by the day!
90.
▲
by
kcorbitt
3y ago
Awesome work! Here's a recent paper released yesterday, also focused on efficiently serving many LoRAs simultaneously: https://arxiv.org/abs/2311.03285 Really looking forward to these innovations becoming more wid
More ›