5 ms·
(1) This is like a new prolog in which the rules are linguistic rules (rationales), and the distilling step-by-step is a way of learning the rules. Call the new
by worldofideas123 3y ago
(1) This is like a new prolog in which the rules are linguistic rules (rationales), and the distilling step-by-step is a way of learning the rules. Call the new prolog lprorat (learning programming with rationales).
(2) Hypothesis: The success of distilling chain of thought learning is that it provides a pseudo large context window. The input pair (label,rationale) allows the student LLM to mimic having a large context window (the context window of a human agent that informs the system by providing the rationale data). So it reduces the embedding distance between words for which entailment is difficult to obtain with short context windows.
(3) The above hypothesis suggests to perform the training process in two phases (a) and (b):
(a) A small model is used as a teacher, this allows the learning model to connect d-distant words. (b) A larger LLM is used as a teacher, the training process continues with new data using the rationale provided by a stronger larger LLM.
In phase (a) the learning system learns to infer d-distant words relations, in phase (b) the system is already prepared to relate (2*d)-distant words because it have developed the necessary skills phase (a).
(4) The two-phase approach can be directly generalized to a multi-phase approach using three meta-parameters: the number of teachers, the relative power of each LLM teacher, and the percentage of data in the training set for each teacher.
(5) Need of pruning: From a geometrical point of view, each phase allows the system to strengthen relations between distant words, increasing latent spaces density promotes part of the latent space getting disconnected, that is creating outliers. Pruning should reduce the risk of outliers that promote hallucinations. Outliers are the germ for establishing relation between very distant words but the rationale data perform that function well, so removing outliers don't pause the learning pace.
(6) It is clear that the rationale data enhances the attention mechanism. This suggests modifying the attention mechanism by giving appropriate weights to the attention with respect to the rationale and self attention for the tag.
(7) It would be desirable for the authors to release those small and powerful LLMs. Therefore giving individuals the power to enter into the SOTA realm and allowing us to create new low latency and cost applications with enhanced privacy.
Edited several times for grammar and new ideas.
- spuz 3y agoWhen you say prolog, are you referring to the programming language, or something else?
- worldofideas123 3y agoI was referring to the prolog programming language. What would happen if we replace the rationale with a concrete program in prolog: Example tag = daughter(X), rationale = father(_,X),female(X); mother(_,X),female(X).