Reducing verbatim loops during generation by penalizing tokens that continue a previously seen sequence.
Encouraging varied text by occasionally excluding dominant choices when several plausible next tokens are available.
Detecting and reducing repetitive language patterns through sampling and targeted fine-tuning.
Reassessing hallucination detection with evaluation metrics that better reflect meaning.
An evolving language-model benchmark with recent questions and objective scoring, designed to limit test contamination.