A decoding method that scales its token cutoff with the model's confidence, giving a simple way to control sampling as temperature changes.
An evolving language-model benchmark with recent questions and objective scoring, designed to limit test contamination.
We present OpenDebateEvidence, a massive dataset containing 3.5 million documents from competitive debate, enabling advancement in argument mining and summarization through large language model fine-tuning.