Internship: Efficient Quantization of Large Language Models
King Abdullah University of Science and Technology (KAUST)
Saudi Arabia
Summary
Advances in extreme LLM quantization exploring linearity-based theory and alternatives to straight-through estimation to reduce global error, enabling precise low-precision inference and efficient large-scale model deployment.