The method to train large AI models on local hardware is generally called QLoRA (Low Rank Adaptation over Quantized base model). In the last few years it's usually done with HuggingFace Transformers (which is the basis of training frameworks such as Unsloth and Axolotl) and bnb 4-bit base model. However, bnb does not yet support MoE models, so the local training of recent MoE models seemed stall for some time.
Since Transformers 5.18, it's started to support GGUF, and it supports recent models with MoE and sparse attentions (and it's not slow, already faster than llama.cpp on Mac). GGUF is a versatile container format. It can be smaller than 4 bpw with surprisingly good quantization quality, so there are new possibilities for local training.
I've shown that we can train Qwen3.8-Flash-Next (125B-A6B + 51B engram) in 40 GiB VRAM without CPU offload. I've optimized it on Strix Halo and it trains at 200 token/s. There is still room to optimize, compared to > 1600 token/s prompt processing we've achieved, and the common sense that LoRA training (with gradient checkpointing) takes 4-5x work of prompt processing. CPU/disk offload (like Strata) and multi-GPU also need more work that I'm not currently focusing on.
On Strix Halo we can also train DeepSeek-V4-Flash (284B-A13B) in 90 GiB VRAM at 100 token/s, but I think it's less practical than Qwen3.8FN for local use.