3 ms·
You can ignore the jaccard similarity field. That was just to monitor the text->cleaned text conversion to make sure it didn’t stray too far from the original w
by deepsquirrelnet 1y ago
You can ignore the jaccard similarity field. That was just to monitor the text->cleaned text conversion to make sure it didn’t stray too far from the original while it was fixing whitespace OCR issues.
I didn’t use embeddings. Nreimers account on huggingface has the minilm models which are BERT-like, but trained using distillation. https://huggingface.co/nreimers/MiniLMv2-L6-H384-distilled-from-BERT-Large https://huggingface.co/nreimers/MiniLMv2-L6-H384-distilled-f... Is the one I started from.
You can then just load that and train it on your data using a standard transformers classification pipeline. ChatGPT can zero shot that part reasonably well if you gave it this description.
From there you should check out the GRPO trainer in TRL. It has taken me a bit of time to learn how to use it effectively. There’s a TON of parameters in the configuration, and occasionally I have to hunt down arxiv papers to understand them.
- veggieroll 1y agoAh! That makes more sense. Thank you for clarifying. Great to see you got so much activity on your HN thread here.