SFT vs DPO vs GRPO
Choose the right post-training stage from the supervision you actually have.
By the LLMix engineering practice · Updated
- Method selection
- SFT
- DPO
- GRPO
Guides
Practical decision frameworks, implementation patterns, and failure analyses for teams customizing models and training specialized agents.
Choose the right post-training stage from the supervision you actually have.
By the LLMix engineering practice · Updated
The practical architecture: state, tools, transitions, resets, verifiers, rewards, and rollout integration.
By the LLMix engineering practice · Updated
A triage order for runs that do not learn: rewards, variance, rollouts, KL, truncation, staleness, leakage.
By the LLMix engineering practice · Updated
Guides are written and maintained by LLMix with AI-assisted drafting, and revised as the underlying methods and tooling change. Updated dates change only for material revisions.
LLMix scopes bespoke post-training programs around one capability: data, environments, rewards, optimization, and validated release.