ROSE: Reordered SparseGPT
A new method for more surgical one-shot pruning of Large Language Models, enabling massive models to run on modest hardware. By reordering the structure of pruning, it better preserves emergent reasoning pathways while achieving aggressive sparsity. The result is generative AI that can operate on consumer GPUs and mobile accelerators without the heavy artifacts of crude quantization.