The biggest model burned on a chip has 17b parameters.
How many parameters does MiMo-V2.6-Pro have? MiMo-V2.6-Pro has 1.0 trillion parameters (42 billion active).
What are the active parameters of MiMo-V2.6-Pro? MiMo-V2.6-Pro is a Mixture of Experts (MoE) model with 1.0 trillion total parameters, but only 42 billion active parameters are used during inference.
So my question is, what would the process be to burn such a "good enough" model on a chip and make it hyper fast and low energy?