3 ms·
Hi we are CLLM authors and thanks for sharing your experience and insights! I can see this drawing skill refining process echoes with the training process in CL
by snyhlxde 2y ago
Hi we are CLLM authors and thanks for sharing your experience and insights! I can see this drawing skill refining process echoes with the training process in CLLM, the only thing is at this point stressor in CLLM training is not getting progressively demanding.
For example, while drawing, you can set very specific time limit on how long you are allowed to draw in each trial and make the time progressively shorter. In CLLM, maybe we can make this the learning process more and more difficult by mapping more and more distant states in Jacobi trajectory to its final state.
We are using the term "consistency" because we draw parallelism between consistency LLM and the consistency model in diffusion image generation where the training processes are analogous.
- Quarrel 2y agoIs it just me, or does this read like it was written by an LLM ... ?!
- snyhlxde 2y agolol I take that as a compliment. Good try but sadly no LLM in this writing :)
- jasonjmcghee 2y agoIt's just much more formal than people generally speak on HN.
- boroboro4 2y agoDo you use same dataset to train / eval the model? Was the model used for example trained on GSM8K dataset for example?
- snyhlxde 2y agoYes, we consider both domain-specific applications (spider for text2SQL, gsm8k for math, codesearchnet for python) as well as open-domain conversational applications (ShareGPT). We use test set from each application to evaluate CLLMs’ performance in our paper. On the other hand, technically CLLM works on any kind of queries. But the speedup might vary. Feel free to try out our codebase for your use cases!