CLARIN-CH Day 2026

30 October 2026

University of Bern

Introduction

After the first two editions of the CLARIN-CH Day, which addressed “ORD: Challenges and Opportunities” (2024) and “Towards a CLARIN-CH Ecosystem of Federated Infrastructure for Language Data” (2025), the 2026 edition is dedicated to community building and knowledge exchange among researchers, practitioners, and institutions working with language resources and language technology in Switzerland.

The theme of this year’s edition is “New developments, challenges, and opportunities for open language resources and language technology”.

The event brings together researchers and experts to discuss new developments in the field. It showcases the variety of areas covered by the CLARIN-CH consortium: basic research (linguistics, pragmatics, argumentation), applied research (language education, translation and interpretation), language preservation efforts (digitalisation of Rumantsch, national dictionaries), corpus linguistics and NLP tools and methods, the use of AI and LLMs in academia, research infrastructures, and the ORD paradigm.

Highlights

Keynote Talk by Prof. Dr Theresa Heyd

University of Heidelberg

Abstract: Understanding sociolinguistic affect in digital discourse data

Studying the connection between language and emotion is not a new topic, and in the past decade, it has often been informed by quantitative and computational approaches such as sentiment analysis. At the same time, the interdisciplinary field of affect studies (Ahmed 2004, Wetherell 2012 and others) has developed an understanding of affect as a cultural, political and stylistic practice, leading away from empiricist concepts of discrete emotions as countable and quantifiable. This is particularly relevant for the interactional and fluid linguistic practices that we encounter in digital contexts. How can we do linguistic research on the affective dimensions of digital linguistic practice? What are meaningful and insightful data that we can work with to understand networked and sometimes short-lived articulations of digital affect, from ick to cringe to aura? And what are some of the procedural and ethical hurdles that may be specifically relevant when engaging with such contexts? In my talk, I will present findings from my work on the sociolinguistics of digital affect (Heyd in press). I will first discuss some underpinnings and implications of understanding sociolinguistic affect in digital discourse data through the lens of affect theory, and present selected case studies and data points. The second part of my talk will reflect on the methodological aspects of working with such affective data, some of which are messy, hard-to-locate and come with ethical and platform-specific challenges.
  • Thematic Sessions on key topics such as multilingual resources, FAIR data principles, and ethical considerations in language data management.
  • Networking Opportunities for participants to connect and exchange.
  • Showcases & Demos of innovative language resources, tools, and infrastructure developed within CLARIN-CH and its partner institutions.
  • Round Table with contributions from the CLARIN-CH Research Data Management Working Group, focusing on ethical challenges, anonymisation, and Data Management Plans (DMPs).

Program

TimeProgramme
9:30 – 10:00Welcome Coffee
10:00 – 11:00Keynote Talk by Prof. Dr. Theresa Heyd: Understanding sociolinguistic affect in digital discourse data
11:00 – 11:30

Thematic Session 1: AI and Language Technologies

  • Integrating Generative AI into Corpus-Assisted Analysis of Public Discourse | Dolores Lemmenmeier (ZHAW) (Abstract, S. 3)
  • Whose dialect, whose variety? AI transcription as a hidden technology of language planning | Francine González Borrell and Yvette Bürki (UNIBE) (Abstract, S. 3)
  • Towards a Fine-Grained Persian Dataset for Sexualized and Gender-Based Hostile Language Detection | Hanieh Habibi, Amir H. Payberah and Davide Picca (UNIL) (Abstract, S. 4)
11:30 – 12:00

Thematic Session 2: Corpora and Discourse Analysis

  • Building a Corpus of Russian Military Telegram Blogs: Methodology, Challenges, and Ethical Considerations | Olga Bikkulova (UNIBE) (Abstract, S. 5)
  • A multi-level approach to annotating deniability in political and legal discourse | Bruna Paz Schmid, Steve Oswald and Lou Odermatt (UNIFR) (Abstract, S. 6)
  • Digital stancetaking towards gender diversity on the platform gutefrage.net | Aline Siegenthaler (UNIFR) (Abstract, S. 6)
12:00 – 13:30Lunch
13:30 – 14:00

Showcases & Demos

  • AViRI – Towards a Shared Infrastructure for Audiovisual Data and Metadata | Josephine Diecke, Simon Spiegel and Teodora Vukovic (UZH) (Abstract, S. 14)
  • Audio-video data processing pipelines | Masoumeth Chapariniya and Teodora Vukovic (UZH)
  • MAVA and DAVA: building a cluster of infrastructure for processing multimodal audio-visual data | Teodora Vukovic (UZH)
14:00 – 14:30

Thematic Session 3: Open Research Data & FAIR

  • Metamorphosis of a DMP: Navigating Consent and De-identification Across Three Stages of an MA Project | Oscar Jordan (UNIL) (Abstract, S. 7)
  • Language Data Across Disciplines: Survey Insights on ORD Practices, FAIR Awareness, and Training Needs | Julia Krasselt (ZHAW), Elsa Liste Lamas (ZHAW), Maike Fischer (ZHAW), Cristina Grisot (UZH) and Joanna Blochowiak (UZH) (Abstract, S. 8)
  • From Infrastructure to Impact: Understanding Adoption in European Digital Research Infrastructures | Beliz Gökmen, Andrew A. Clark, Gorka Fraga Gonzalez and Steven Moran (UZH) (Abstract, S. 9)
14:30 – 15:00

Thematic Session 4: Language Resources and Infrastructures

  • Making the Ephemeral Durable: The Swiss Sticker Corpus (SSC) as Open Research Data | Yvette Buerki, Kellie Gonçalves and Janis Schneider (UNIBE) (Abstract, S. 9)
  • VotingBooklets: A Multilingual Parallel Corpus of Swiss Federal Referendum Texts | Elina Stüssi and Jannis Vamvas (UZH) (Abstract, S. 10)
  • Database-first lexical infrastructure for Proto-Albanian research: interoperability, FAIR data, and reusable workflows for Albanian–Romanian cognacy and the Proto-Albanian–Thracian substrate segment | Edmond Cane (Luarasi University) (Abstract, S. 10)
15:00 – 15:30Coffee Break
15:30 – 16:00

Thematic Session 5: Multilingualism & Linguistic Diversity

  • The importance of multilingual and cross-linguistic corpora for the study of language contact | Olivier Winistörfer (UZH), Maxim Makartsev (University of Oldenburg) and Anastasia Escher (ETH Zurich) (Abstract, S. 11)
  • Linguistic data structures and standardization are not theory-neutral | Sandra Auderset (UNIBE) and Tiago Tresoldi (Abstract, S. 12)
  • At-Issueness Sensitivity in L2 Misinformation Detection: The Role of Cognitive Load | Alan Lombardini (UNIFR), Giulia Giunta (UNINE), Diana Mazzarella (UNINE), Didier Maillat (UNIFR) (Abstract, S. 13)
16:00 – 16:30

Poster Session and Networking

  • RefCo2 from the fields to the archive: multi-perspective testing of a corpus-quality platform for reusable language data | Jocelyn Aznar (UNIBE), Miranda Dickerman (UZH) and Jonathan Reich (UNIBE) (Abstract, S. 14)
  • The “Person Marking in South-Central Trans-Himalayan” (PMST) database | Linda Konnerth, Sandra Auderset, Jonathan Reich, and Susie Kanshouwa (UNIBE) (Abstract, S. 15)
  • Infrastructural support for data analysis and reuse in Interactional Linguistics: potential of LCP-Videoscope | Johanna Miecznikowski (USI), Elena Battaglia (University of Bologna), Jérôme Jacquin (UNIL), Thomas Schmidt (University of Duisburg-Essen / linguisticbits), Teodora Vuković (UZH) and Jeremy Zehr (UZH) (Abstract, S. 16)
  • Beyond Clinical Pain Descriptors: Detecting Figurative Pain Descriptions in Patient Narratives | Dolores Lemmenmeier (ZHAW), Nikolai Shurakov (UZH), Yvonne Ilg (UZH), Lucien Baumgartner (UZH) and Sabina Maria Räz (USZ) (Abstract, S. 17)
  • FAIRifying a Multilingual European Parliament Corpus for the Study of Gender-Inclusive Lexical Innovation | Nikos Tsourakis, Aurélie Picton, Julie Humbert-Droz, Charlotte Jeanneret and Anne Kosakevitch (UNIGE) (Abstract, S. 18)
  • CLARIN – a pan-European distributed digital infrastructure for language data and technology | Cristina Grisot (UZH) 
  • Two communities, one shared practice: CLARIN-CH RDM Working Group and SRDSN L-RDM Node | Joanna Blochowiak and Cristina Grisot (UZH) (Abstract, S. 19)
  • LiRI and the Lia Rumantscha work together to build new infrastructure and collect data to promote Romansh | Ignacio Pérez Prat (Lia Rumantscha), Jannis Vamvas, Nikolina Rajović and Igor Mustač (UZH) (Abstract, S. 19)
  • LiRI and Linguist List | Steven Moran and Valeriia Vyshnevetska (UZH) (Abstract, S. 20)
  • Swissxox RAG: Retrieval-First AI for Large Multilingual News Archives | Igor Mustač and Nikolina Rajović (UZH) (Abstract, S. 20)
16:30 – 17:30Round Table on DMPs for language data and Conclusive Remarks
17:30

End of the event

Download the full programme here.

You can find all abstracts collected in the Booklet of Abstracts.

Registration

Registration for CLARIN-CH Day 2026 is now open. The event is free and open to everyone with an interest in language data, research infrastructure, and Open Science. No prior knowledge or CLARIN-CH affiliation is required.

Coffee breaks and lunch are included. Participants cover their own travel and accommodation costs.

Join us at CLARIN-CH Day 2026

Location

University of Bern

Organising committee

  • Sandrine Zufferey (UniBE)
  • Joanna Blochowiak (UZH, CLARIN-CH)
  • Cristina Grisot (UZH, CLARIN-CH National Coordinator)

Scientific Committee: Members of the CLARIN-CH Consortium

Edition
0

This is the third edition of the CLARIN-CH Day. If you want to read more on the previous editions, including programmes, books of abstracts, and event recaps, you can find information here: