Microsoft’s Apache-2.0 ZeRO optimizer for multi-GPU, multi-node LLM training.
Catalog snapshot
The ZeRO-2/3 configs are the usual starting points; multi-GPU expected.
Fetched 1
DeepSpeed is Microsoft’s Apache-2.0 training optimizer: ZeRO memory sharding, multi-node scaling, and inference kernels that let fixed hardware train far larger models. A CLI library for serious training runs, not casual tuning.
ZeRO sharding · multi-node training · inference kernel optimizations · parameter offload