Accounting for ambiguous ground truth when comparing AI systems with experts.
Detecting and reducing repetitive language patterns through sampling and targeted fine-tuning.
Reassessing hallucination detection with evaluation metrics that better reflect meaning.
Improving the reliability of fine-tuned foundation models with priors that account for uncertainty.
Adjusting dropout during inference to estimate uncertainty without additional labels or training.