27 ms·
What is interesting to me about this approach is that they're fine tuning an existing model. If they could train a model using input/output pairs from scratch a
by parentheses 3y ago
What is interesting to me about this approach is that they're fine tuning an existing model. If they could train a model using input/output pairs from scratch and entirely on the optimization task, I could see a smaller LLM performing drastically better than the compiled output - according to whatever the loss function is implemented to optimize for.