Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
nihit-desai
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
Show HN: Open-source LLM for data labeling
(huggingface.co)
6 points
by
nihit-desai
2y ago
|
0 comments
2.
▲
Show HN: Refuel-LLM, a large language model for data annotation and enrichment
(app.refuel.ai)
9 points
by
nihit-desai
3y ago
|
0 comments
3.
▲
Correcting and Improving LLM Predictions Without Labels
(hazyresearch.stanford.edu)
1 points
by
nihit-desai
3y ago
|
0 comments
4.
▲
Never ask an LLM to generate a confidence score
(refuel.ai)
17 points
by
nihit-desai
3y ago
|
0 comments
5.
▲
by
nihit-desai
3y ago
function calling, as I understand it, makes LLM outputs easier to consume by downstream APIs/functions ( https://openai.com/blog/function-calling-and-other-api-updat... ). Autolabel is quite orthogonal to this - it&
6.
▲
by
nihit-desai
3y ago
Yep! I totally understand the concerns around not being able to share data externally - the library currently supports open source, self-hosted LLMs through huggingface pipelines ( https://github.com/refuel-ai/autolabel
7.
▲
by
nihit-desai
3y ago
Hi! The earlier post was a report summarizing LLM labeling benchmarking results. This post shares the open source library. Neither is intended to be an ad. Our hope with sharing these is to demonstrate how LLMs can be used for data labeling
8.
▲
Show HN: Autolabel, a Python library to label and enrich text data with LLMs
(github.com)
153 points
by
nihit-desai
3y ago
|
22 comments
9.
▲
by
nihit-desai
3y ago
>> don't trust that there was no funny business going on in generating the results for this blog All the datasets and labeling configs used for these experiments are available in our Github repo ( https://github.com
10.
▲
by
nihit-desai
3y ago
I mean, sure. For ground truth, we are using the labels that are part of the original dataset: * https://huggingface.co/datasets/banking77 * https://huggingface.co/datasets/lex_glue/viewer&#
11.
▲
by
nihit-desai
3y ago
Hmm, I'm not suggesting training a smaller model from scratch - in most cases you'd want to finetune a pretrained model (aka, transfer learning) for your specific usecase/problem domain. The need for labeled data for any kind
12.
▲
by
nihit-desai
3y ago
Partially agree, but it's a continuous value rather than a boolean. We've seen LLM performance largely follow this story: https://twitter.com/karpathy/status/1655994367033884672/phot... From benchma
13.
▲
by
nihit-desai
3y ago
Good question - one followup question there is value for who? If it is to train the LLM that is labeling, then I agree. If it is to train a smaller downstream model (e.g. finetune a pretrained BERT model) then the value is as good as coming
14.
▲
by
nihit-desai
3y ago
Hi, one of the authors here. Good question! For this benchmarking, we evaluated performance on popular open source text datasets across a few different NLP tasks (details in the report). For each of these datasets, we specify task guideline
15.
▲
LLMs can label data as well as human annotators, but 20 times faster
(refuel.ai)
55 points
by
nihit-desai
3y ago
|
31 comments
16.
▲
AI Canon
(a16z.com)
518 points
by
nihit-desai
3y ago
|
219 comments
17.
▲
by
nihit-desai
3y ago
A comprehensive list of GPU options and pricing from cloud vendors. Very useful if you're looking to train or deploy large machine learning/deep learning models.
18.
▲
Cloud GPU Resources and Pricing
(fullstackdeeplearning.com)
184 points
by
nihit-desai
3y ago
|
59 comments
19.
▲
by
nihit-desai
3y ago
Very neat! I was looking for something exactly like this for a library I'm building - will try it out
20.
▲
by
nihit-desai
4y ago
This upcoming course covers topics such as bootstrapping datasets and labels, model experimentation, model evaluation, deployment and observability. The format is 4 weeks of project-driven learning with a peer cohort of motivated, interesti
21.
▲
Show HN: New course on real-world ML systems
(corise.com)
13 points
by
nihit-desai
4y ago
|
1 comments