Natural Language Processing

Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs

A decoding method that scales its token cutoff with the model's confidence, giving a simple way to control sampling as temperature changes.

LiveBench: A Challenging, Contamination-Limited LLM Benchmark

An evolving language-model benchmark with recent questions and objective scoring, designed to limit test contamination.

OpenDebateEvidence: A Massive-Scale Argument Mining and Summarization Dataset

We present OpenDebateEvidence, a massive dataset containing 3.5 million documents from competitive debate, enabling advancement in argument mining and summarization through large language model fine-tuning.