Lower bit widths and fewer nonzero weights promise smaller models and fewer arithmetic operations. Those reductions do not guarantee lower latency: packing, dequantization, sparse metadata, irregular ...