3 ms·Show HN: Minimal LLM Post-Training Experiments on an 8GB GPU (SFT, DPO, GRPO)21 points by popopanda 2mo agopopopanda 2mo ago[flagged]