Schedule
Program (Location: Room C206; see also the ESSLLI 2026 Schedule):
Monday, August 3
11:00-11:10: Introduction. Louise McNally, Gemma Boleda (U. Pompeu Fabra)
11:10-11:30: Why the small mug? Modifier-informed reasoning in reference interpretation. Kelly Cheuk, Hannah Rohde, Chris Cummins (U. Edinburgh)
11:30-12:30: Invited talk: Frege in the (Residual) Stream: Sense and Reference inside Language and Vision-Language Models. Denis Paperno (Utrecht U.)
Abstract: Often the same referent can be described in multiple ways. In Sinn und Bedeutung, Gottlob Frege proposed an influential logical framework for thinking about those types of situations. In his example, expressions "the morning star" and "the evening star" refer to the same celestial body (Venus), but are not fully interchangeable; Frege explains this by the difference in their sense despite referential identity. More recent work has focused on the distinctions between referentially compatible expressions in other contexts than Frege — more pragmatic than semantic in nature — but his framework is still helpful. For example, Dale and Reiter (1995)'s well-known algorithm for Referring Expression Generation in context in fact precedes any actual referring expression generation, and is a procedure for sense selection (which properties to name, at what level of granularity, under which classification of the referent).
Modern neural models perform tasks like referent description as apparent black boxes, mapping inputs to text outputs. Do they operate with sense and reference of expressions internally? I will survey interpretability research that begins to answer this. Against this background, I will argue that interpretability methods allow us to address empirically the interplay of sense and reference inside generative models, e.g. when a model produces a referring expression, in what order the referent and the sense that is chosen to describe it are identified.
This perspective sheds light on the central theme of the workshop: to what extent our theoretical distinctions (sense vs. reference, taxonomic granularity, cross-classifiability, etc.) are sufficient for characterizing what happens during the actual choice of a referring expression, and whether an examination of model internals can help us revise our assumptions.
Tuesday, August 4
11:00-11:45: The role of subject discourse-accessibility, specificity, and concreteness in multiple center-embedding. Emily Davis, Robert Kluender (UCSD)
11:45-12:30: Effect of surprisal on the form and prosody of referring expressions. Ivan Rygaev, Güliz Günes, Asya Achimova (U. Tübingen)
Wednesday, August 5
11:00-11:20: Accessibility-precision tradeoffs in grounded referring expression choice: Insights from monolingual and bilingual speakers. Junyi Chen, Martin Zettersten, Anne L. Beatty-Martínez (UCSD)
11:20-12:20: Invited talk: From Winograd's SHRDLU to Vision-Language Models: Why Aren't We There Yet? Raffaella Bernardi, Free U. of Bozen-Bolzano
Abstract: In 1972, Terry Winograd introduced SHRDLU, a system that could "answer questions, execute commands, and accept information in an interactive English dialogue." More than fifty years later, these abilities remain at the heart of research on grounded language understanding.
In 2021, Sandro Pezzelle and I revisited Winograd's vision in our position paper Linguistic Issues Behind Visual Question Answering, using what we called Winograd's desiderata—the linguistic phenomena and reasoning abilities demonstrated by SHRDLU—to assess the state of the art in Visual Question Answering. Since then, Vision-Language Models have made remarkable progress, raising the question of how much closer we have come to realizing Winograd's original vision.
In this talk, I will revisit the open challenges identified in our 2021 paper and discuss recent advances in visually grounded language understanding, with a particular focus on reference expression generation and interpretation, visually grounded negation, and pragmatic reasoning. I will argue that taking visual dialogue not as another benchmark, but as the target communicative ability and studying the learning trajectory that leads to it may help bridge the gap between current Vision-Language Models and genuinely interactive conversational agents.
12:20-12:30: Poster flash talks 1-3 (see list under "Poster session")
Thursday, August 6
11:00-11:10: Poster flash talks 3-6 (see list under "Poster session")
11:20-12:30: Poster session (All posters). List of posters:
Mandarin demonstratives and bare NPs: more than anaphoricity and uniqueness -- Evidence from natural discourse data. Jennifer Yao (Poly. U. CPCE)[Cancelled]- Anaphoric relations between narrative layers. Katja Jasinskaja, Klaus von Heusinger (U. Köln)
- Partner effects on referential expression production in autistic adults. Yage G. Xin, Rachel Dudley (UCSD)
- Croatian jedan and neki as (non)specificity markers. Rita Rumboldt (FU-Berlin)
- Multidimensional indexicality of second-person pronouns: The indexical meanings of T/V address forms. Yizhuo Zhang (U. Groningen)
- Referring expressions and global anchoring: Subjective predication, third-person extension, and honorific agreement in Korean. Chungmin Lee (Seoul National U.)
Friday, August 7
11:00-11:20: Recognitional demonstratives in Chinese: A pragmatic analysis of nàxiē in the CCL Corpus. Lijun Li, Chunxu Chen (U. Freiburg)
11:20-12:20: Invited talk: How grammar and predictability affect pronoun production and interpretation: evidence from interactive tasks. Laia Mayol (U. Pompeu Fabra)
Abstract: When do speakers choose a personal pronoun (e.g., she) instead of a fuller referring expression (e.g., Rosalía or the singer in white), and how do listeners interpret these expressions? Previous research has identified two key influences on reference production and interpretation: contextual cues that shape referent predictability and the structural properties of the antecedent. While structural properties are known to affect both production and interpretation, it remains unclear whether predictability plays a comparable role for speakers and listeners. Moreover, most previous studies have examined production and interpretation separately, typically using non-interactive tasks. We address these questions in two interactive experiments designed to promote efficient communication. In a forced-choice task, referent predictability influenced both speakers' choice of pronouns and listeners' interpretation. However, this effect did not replicate in a free-production task, suggesting that the role of predictability may depend on task demands.
12:20-12:30: Closing comments