--n-cpu-ffn sounds interesting for dense models, I might have to try that maybe with Gemma 4 12b in my pitiful 8GB VRAM lol