Image credit: UnsplashMin-p is a decoding rule that scales the token-selection threshold with the probability of the model’s most likely next token. We study this approach across reasoning and creative-writing tasks, examining how confidence-dependent truncation changes generation as temperature varies.
Presented at ICLR 2025 (Oral). First released in July 2024; the latest arXiv revision is from November 2025.