Dynamic Inference with Repeated AI Measurements in German Newspaper Discourse

Models German newspaper discourse from repeated LLM measurements while separating response stability from semantic validity.

Together with Andreas Reich, I study how repeated large language model classifications can measure political discourse. The project focuses on German newspaper coverage from 1990 to 2025. Its current application examines whether inflation-related articles explicitly attribute price pressure to German federal political actors or policies.

Repeated prompts can reduce call-level randomness, but they cannot remove systematic classification error. We therefore distinguish response stability from semantic validity. A small human validation sample anchors the substantive construct, while repeated model calls identify dependence among responses to the same text.

The current paper develops a continuous-time latent-state model for the activation and persistence of discourse. It targets activation rates, decay rates, and spell durations, using the German newspaper corpus as its empirical application. The paper remains a design scaffold; licensed newspaper snippets and archive links are excluded from the public repository.