3 ms·
Gemini Distillation Service
- deleted 2mo ago[deleted]
- deleted 2mo ago[deleted]
- moralestapia 2mo agoFine-tuning and distillation used to mean two different things. Now not so much anymore, I see.
- bizzletk 2mo agoCan you explain the difference?
- nerdsniper 2mo agoFine-tuning is done against a dataset. Distillation is done against a model.
- verdverm 2mo agoDistillation is used to build part of a data set for fine-tuning (loosely interpreted). Advanced model traces are useless if you don't have a base model that is good enough to be improved by them.
- porridgeraisin 2mo agoDistillation originally meant matching the distribution of the student model to the teacher model using something like a KL divergence. When you instead fine-tune the student on the samples from the teacher, which is what people mean by distillation today, you are in effect doing a monte-carlo version of the same thing. While in theory this is higher variance, given modern setups where the student and teacher are both large and are RLd heavily (leading to a sharp teacher distribution), and given that you typically use lots and lots of samples, it ends up OK.
- sapienskid 2mo agoHow to get access?
- rs38 2mo ago404?
- verdverm 2mo agojust heard on Risky Business pod that they apparently have taken it down, I distinctly remember reading it two days ago, came here to find the link, found your 404?
- verdverm 2mo agoarchived link: https://web.archive.org/web/20260728173925/https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/tuning/distillation https://web.archive.org/web/20260728173925/https://docs.clou...