Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kcorbitt
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
Analyzing OpenAI's Reinforcement Fine-Tuning: Less Data, Better Results
(openpipe.ai)
4 points
by
kcorbitt
2y ago
|
0 comments
32.
▲
by
kcorbitt
2y ago
No, generally speaking OpenAI doesn't re-use training data between customers. It's worth it to them anyway because they learn what does/doesn't work on different tasks Of course, it isn't your IP free and clear eith
33.
▲
by
kcorbitt
2y ago
Ohh, I really like that as a potential proxy metric!
34.
▲
by
kcorbitt
2y ago
This is a fair point. The reason why I think "correlation" is a better metric than "predicts the exact correct score" is because of how I'll be using this model in the next post. Broadly, the main use case for this
35.
▲
by
kcorbitt
2y ago
Yes! The architecture is almost identical. The only difference is in the final layer. In an LLM used for text generation, the final layer has a separate output for every potential token the model could produce, and we decide which token to
36.
▲
by
kcorbitt
2y ago
I hadn't heard of isotonicregression before but I like it! > it's good to create two models, one for likelihood of zero karma, and another expected karma, conditional on it being non-zero. Another way to do this is to keep a si
37.
▲
by
kcorbitt
2y ago
> For me the objective of "most up votes" is not fully correlated with where I get the most value on HN. Most of the time, the most up voted I would have found them anyway on other platforms. Yes, this is a fantastic point. I&#
38.
▲
by
kcorbitt
2y ago
I link to it from the post, but all the code is open source! You can find the specific training script here: https://github.com/OpenPipe/best-hn/blob/main/stories_train_... And all the graphs for the blo
39.
▲
by
kcorbitt
2y ago
Yep that makes sense. Would be interesting to do a follow-up that explicitly includes these variables and see if it meaningfully improves the results.
40.
▲
by
kcorbitt
2y ago
Nope. I was actually planning on asking dang if he has any insights there. If he sees this thread hopefully he can chime in!
41.
▲
by
kcorbitt
2y ago
Haha great question. Since it's only trained on on-platform HN content and not external links, this post is a little bit out of distribution for it unfortunately. I'm thinking about scraping a corpus of external links and running
42.
▲
by
kcorbitt
2y ago
Yep, that's why I included the post date in the information available to the model; in theory (if it's smart enough) it should be able to take that into account. That said I didn't include time-of-day; it would be interesting
43.
▲
by
kcorbitt
2y ago
Hey all, this project was a labor of love I worked on in my spare time over the last couple of weeks. Happy to answer any questions!
44.
▲
Using reinforcement learning and $4.80 of GPU time to find the best HN post
(openpipe.ai)
217 points
by
kcorbitt
2y ago
|
95 comments
45.
▲
by
kcorbitt
2y ago
Also my dad wrote large parts of the Windows 95 kernel so I guess I've always had a soft spot for Windows, even if I haven't used it in 10 years. :)
46.
▲
by
kcorbitt
2y ago
Nostalgia and vibes!
47.
▲
by
kcorbitt
2y ago
On the one hand very true, but on the other hand if you're a dev any python or nodejs package you install and run could do the same thing and the world mostly continues working.
48.
▲
by
kcorbitt
2y ago
(author here) yes it often confidently declares success when it clearly hasn't performed the task, and should have enough information from the screenshots to know that. I'm somewhat surprised by this failure mode; 3.5 Sonnet is pr
49.
▲
Show HN: Agent.exe, a cross-platform app to let 3.5 Sonnet control your machine
(github.com)
406 points
by
kcorbitt
2y ago
|
232 comments
50.
▲
by
kcorbitt
2y ago
I saw this when it was making the rounds on X a few days ago. Fair warning: it seems like at least some sections are AI-generated, and there isn't much insight to be gained from reading the actual sections compared to eg. reading the r
51.
▲
DPO fine-tuning outperforms SFT
(openpipe.ai)
1 points
by
kcorbitt
2y ago
|
0 comments
52.
▲
by
kcorbitt
2y ago
Funnily enough, web scraping was actually the motivating use-case that started my co-founder and I building what is now openpipe.ai. GPT-4 is really good at it, but extremely expensive. But it's actually pretty easy to distill its skil
53.
▲
by
kcorbitt
2y ago
Big congrats on the official launch! Slightly tooting my own horn here, but at OpenPipe we've got a collaboration set up with Traceloop. That means you can record your production traces in Traceloop then export them to OpenPipe where y
54.
▲
by
kcorbitt
2y ago
(Disclaimer: founder of OpenPipe). Thanks for the shout-out. Note that we're actively working on improved evaluations that will let you add more specific criteria as well as more evaluation types, like comparing field values to that of
55.
▲
by
kcorbitt
2y ago
(Disclaimer: I'm the founder of OpenPipe, one of the fine-tuning services OP tried and ultimately the one that produced the highest performing model, it appears.) Data extraction is a use case that fine-tuned models are fantastic at,
56.
▲
by
kcorbitt
2y ago
Hi, I'm the post author. At the end of the article I share how you can access all three MoA models we've prepared via our public chat completions API!
57.
▲
OpenPipe Mixture of Agents: Outperform GPT-4 at 1/25th the Cost
(openpipe.ai)
13 points
by
kcorbitt
2y ago
|
2 comments
58.
▲
by
kcorbitt
2y ago
I actually did a blog post a few months ago where I analyzed HN commenter sentiment across AI, blockchain, remote work and Rust. The final graph at the very end of the post is the relevant one on this topic! https://openpipe.ai&#
59.
▲
by
kcorbitt
2y ago
Could you or anyone else with experience with Temporal share how hard it is to self-host in practice? Like, is this more like Redis (self-hosting is trivial) or Supabase (nominally self-hostable, but if you try to do it you'll quickly
60.
▲
by
kcorbitt
2y ago
OpenPipe | Bellevue, WA | ONSITE | Full Time | Founding Engineer Hey HN! I'm the founder of OpenPipe, the easiest way to deploy fine-tuned models to production. We're looking for a founding engineer who is comfortable wearing many
More ›