3 ms·
Thanks for the explanation and the link. I learned something today. I'm definitely not an expert, but to me, RL and other techniques looks like a guide or a co
by criticalfault 1y ago
Thanks for the explanation and the link. I learned something today.
I'm definitely not an expert, but to me, RL and other techniques looks like a guide or a constraint on the still 'next token prediction' concept. What I do not get is - is this all about training? Or is this about inference.
In any case, this is still an eye opener and I need to study this a bit more.
When talking inference, models from huggingface are composed of what then? Because they can do angentic stuff, no?