Abstract

From statistics, scoring rules evaluate probabilistic forecasts of an unknown state against the realized state and are a fundamental building block in the incentivized elicitation of information. Proper scoring rules are ones where the forecaster is incentivized to report their true belief about the unknown state.

The talk will review the classical and recent work on (statistical) scoring rules. This review will include the geometry of proper scoring rules (classical) and the optimization of scoring rules (to induce more effort of the forecaster to learn about the state).

The talk will generalize these statistical scoring rules to give an algorithm for scoring elicited text against ground truth text using domain-knowledge-free queries to a large language model (specifically ChatGPT).

A key application for text scoring is in peer grading where students grade each other's work. In peer grading, a challenge is in incentivizing high quality peer reviews. The textual scoring rule can be used to grade a textual peer review by comparing it to the textual instructor review. This example parallels super alignment, specifically, the more intelligent students are scored using the less intelligent domain-knowledge-free queries that are sent to a large language model.

Video Recording