Skip to main content
LLMix

Guides

Post-Training Engineering Guides

Practical decision frameworks, implementation patterns, and failure analyses for teams customizing models and training specialized agents.

Start here

SFT vs DPO vs GRPO

Choose the right post-training stage from the supervision you actually have.

By the LLMix engineering practice · Updated

  • Method selection
  • SFT
  • DPO
  • GRPO
Read the guide

Why GRPO Training Fails

A triage order for runs that do not learn: rewards, variance, rollouts, KL, truncation, staleness, leakage.

By the LLMix engineering practice · Updated

  • GRPO
  • Failure diagnosis
  • RLVR
Read the guide

Guides are written and maintained by LLMix with AI-assisted drafting, and revised as the underlying methods and tooling change. Updated dates change only for material revisions.

Turn the framework into a program.

LLMix scopes bespoke post-training programs around one capability: data, environments, rewards, optimization, and validated release.