Compute — Windows ARM64 / Copilot+ PC
Snapdragon X-series laptops and desktops, where memory headroom allows large targets and multi-model setups.Speculative decoding with MTP
Accelerate Gemma-4-26B decoding with a Multi-Token Prediction draft model, in both
geniex infer and the local server.Before you start
Every tutorial assumes you have:- The CLI installed and on your
PATH— see Install. - A supported Snapdragon device, or a remote session on Qualcomm Developer Cloud — see Platforms & runtimes.
- Enough free disk for the models involved. Tutorials list sizes up front;
geniex listshows what you already have cached.
Was this page helpful?