Tag: #grpo
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 3 posts
How Reasoning Models Are Built — From Chain-of-Thought to GRPO, in Twelve Papers
o1, R1 and QwQ are not new architectures. They are long output optimised toward a gradeable objective. This traces the lineage through twelve papers — Chain-of-Thought (2022), STaR, PRM, DeepSeek-R1's GRPO, s1's budget f
2026-09-06 · 12 min read #ai#llm#reasoning#reinforcement-learning#grpoLLM Fine-tuning Frameworks 2026 — A Deep Dive into Axolotl, Unsloth, LLaMA-Factory, TRL, PEFT, and TorchTune ♪ Listenable
A complete map of the 2026 LLM fine-tuning ecosystem. Open-source frameworks like Axolotl, Unsloth, LLaMA-Factory, TRL, PEFT, and TorchTune. LLM Foundry (MosaicML, acquired by Databricks). Cloud fine-tuning APIs from Mod
2026-05-16 · 28 min read #llm#finetuning#axolotl#unsloth#llama-factoryAI Safety & Alignment 2026 Deep Dive - Constitutional AI · RLHF · DPO · GRPO · Mechanistic Interpretability · AISI Evals · Red Team ♪ Listenable
A single-shot map of AI safety and alignment as of 2026. Starts from conceptual roots like outer/inner alignment and mesa-optimization, walks through training-time alignment (RLHF, DPO, GRPO, Constitutional AI), frontier
2026-05-16 · 20 min read #ai-safety#ai-alignment#constitutional-ai#rlhf#dpo