Adding logit penalty for "wait", "maybe" and "perhaps" to Qwen models improves their accuracy
Adding logit penalty for "wait", "maybe" and "perhaps" to Qwen models improves their accuracy
Deep Dive
Meta came out with a banger paper https://arxiv.org/pdf/2606.00206 , but it did not look at various quantizations supported in llama.cpp. So I did a run on 50 random MATH-500 questions ( https://huggingface.co/datasets/HuggingFaceH4/MATH-500 ) and ran it on various quantizations of https://huggingfa