Training LLMs on an 8GB GPU is now a weekend project
Someone just released a minimal codebase for supervised fine-tuning, direct preference optimisation, and group relative policy optimisation on consumer hardware.
Read the note 1 min read
From the notebook
1 note on this topic.