On quantisation, a thing I wish more evaluations reported: which tasks degrade, not just how much average quality drops. A 2% average drop can mean everything got slightly worse, or it can mean nothing changed except that multi-step arithmetic fell off a cliff. Those are completely different products for anyone building on top of it. Average-quality reporting hides the shape of the degradation, and the shape is the part that determines whether the quantised model is usable for your workload.
