Hauptinhalt

Aktivitäten

2026

  • VarDial

    13th Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial)

    Datum: 29. März 2026
    Ort: Rabat, Marokko

    Description

    VarDial is a well-established series of workshops promoting a forum for scholars working on a range of topics related to the study of diatopic language variation from a computational perspective.

    The workshop deals with computational methods and language resources for closely related languages, language varieties, and dialects. We welcome papers dealing with one or more of the following topics:

    • Language resources and tools for similar languages, varieties and dialects;
    • Evaluation of language resources and tools applied to non-dominant language varieties;
    • Cross-lingual transfer and adaptation of models to similar languages, varieties and dialects;
    • Automatic identification of lexical variation;
    • Automatic classification of language varieties;
    • Machine translation between closely-related languages, language varieties and dialects;
    • Corpus-driven studies in dialectology and language variation;
    • Computational approaches to mutual intelligibility between dialects and similar languages;
    • Text similarity and adaptation between language varieties;
    • Linguistic issues in the adaptation of language resources and tools (e.g., cognate detection, semantic discrepancies, lexical gaps, false friends);
    • Studies focusing on related creole languages and their lexifier languages;
    • Studies focusing on diachronic language variation (e.g. phylogenetic methods, historical dialects).

    Organizers

    Yves Scherrer
    University of Oslo (Norway)

    Noëmi Aepli
    University of Pennsylvania (USA)

    Verena Blaschke
    LMU Munich and Munich Center for Machine Learning (Germany)

    Tommi Jauhiainen
    University of Helsinki (Finland)

    Nikola Ljubešić
    Jožef Stefan Institute and University of Ljubljana (Slovenia) 

    Preslav Nakov
    Mohamed bin Zayed University of Artificial Intelligence (UAE) 

    Jörg Tiedemann 
    University of Helsinki (Finland)

    Marcos Zampieri 
    George Mason University (USA)

  • ICLaVE13

    Session "Embracing Variability in Natural Language Processing" im Rahmen der ICLaVE13

    Datum: 30. Juni 2026
    Ort: Lausanne, Schweiz

    Description

    The 13th International Conference on Language Variation in Europe (ICLaVE | 13) is co-hosted by the University of Lausanne and the University of Bern.

    The theme of the conference is language, im/mobilities, and belonging. With this theme, we wish to explore the crucial role that linguistic variation plays in the construction, regimentation, and evaluation of communal belonging.

    Embracing Variability in Natural Language Processing

    For the longest time, Natural Language Processing (NLP) has not engaged with variation in language in a systematic way – one that matches both the variability of language use in real life and the many different forms of linguistic research on the topic of language variation and change. Often, input variation in language technology is seen as some kind of “noise” to be normalized during processing. Only recently, the NLP community has developed more interest in analyzing, modeling and visualizing varied language material (Zampieri et al. 2020, Joshi et al. 2025). At the same time, even though dialectologists have successfully applied large-scale computational processing and analysis methods in the subfield of dialectometry (Wieling & Nerbonne 2015), sociolinguistics only recently started to systematically make use of computational methods of language processing and data analysis at a larger level (Purschke & Hovy 2019). In this context, the term “computational sociolinguistics” (CSLX; Nguyen et al. 2016) has been coined.

    With the shift from traditional NLP approaches to (generative) large language models (LLMs), new challenges and opportunities emerge, given the inherent variability of language in all domains and social contexts: language models needs to be able to tackle input variation in text data to represent actual language use; under-resourced varieties need to be adequately represented in large-scale models; evaluation benchmarks need to reflect linguistic variability without unwanted biases. On the other hand, the natural variability of language production offers many promising starting points to expand the focus of NLP research. Such models also hold significant potential for uncovering linguistic variation across different dimensions.

    The question of variability in NLP is becoming even more pertinent as generative LLMs are increasingly adopted by the general population. There are indications suggesting that NLP models can even shape linguistic variation and drive language change among humans (Yakura et al. 2024). Linguistic variation is intrinsically tied to regional identity and belonging to certain social groups, and it is crucial to ensure that language technology represents and maintains these connections in the most responsible way. LLMs have been found to show biases towards standard language varieties and reinforcing tendencies towards standardization, potentially impacting people’s own perceptions about the status of their language and identity (Smith et al. 2024). Enabling models to generate text in minoritized language varieties also raises several risks and concerns, such as stereotypization or cultural appropriation, without this being of obvious benefit to the speakers of such varieties (Smith et al. 2024, Blaschke et al. 2024). Engaging with NLP thus creates new - much needed - bridges between linguistic theory, overarching sociological questions such as belonging and identity, and technological practice.

    Against this backdrop, we propose a panel consisting of a diverse group of researchers in terms of disciplinary background, countries, target languages and varieties, gender and career stage, embracing variation in NLP also on a thematic and structural level. The main goals of the panel are to systematically take stock of the available resources for small languages and non-standardized varieties, and to contribute best practice examples for developing variation-friendly NLP resources.

    Session Chair:

    Yves Scherrer

    Barbara Plank

    Christoph Purschke

    Alfred Lameli

2025

  • Scoping Workshop

    Gründung des Netzwerks Regionale Sprache und Künstliche Intelligenz im Rahmen eines Scoping Workshops

    Datum: August 2025
    Ort: Hannover

    Beschreibung

    Gefördert von der VolkswagenStiftung trafen sich 35 Wissenschaftler*innen im August 2025 auf Schloss Herrenhausen in Hannover. Bei der Zusammenkunft handelte es sich um einen dreitägigen Scoping Workshop, also ein Arbeitstreffen, bei dem Perspektiven und Aufgaben für die Weiterentwicklung einer Fachdisziplin ausgearbeitet werden. Im Fokus des Workshops stand das Thema „Künstliche Intelligenz und regionale Identität – Potenziale der Dialektologie im Zeitalter digitaler Transformation“.

    Aus diesem Arbeitsprozess ist nun das Positionspapier „Regionale Sprache und Künstliche Intelligenz im Zeitalter der digitalen Transformation“ entstanden, das in der Zeitschrift für Dialektologie und Linguistik (ZDL) erschienen ist. Das Positionspapier beschäftigt sich damit, wie Künstliche Intelligenz und Dialektforschung Hand in Hand gehen können. Es zeigt den aktuellen Stand der Einbindung von KI-Verfahren in die Forschungspraxis auf, benennt notwendige Voraussetzungen und lotet die Potenziale der Weiterentwicklung aus. Betrachtet werden jedoch auch mögliche Risiken. Daraus werden schließlich Empfehlungen für unterschiedliche Akteure abgeleitet, wie etwa die Fachgesellschaften (z. B. Internationale Gesellschaft für Dialektologie des Deutschen) und Fördermittelgeber (z. B. Deutsche Forschungsgemeinschaft).

    Die gemeinsame Arbeit im Rahmen des Scoping Workshops hat zur Gründung des Netzwerks Regionale Sprache und Künstliche Intelligenz geführt, das sich als offener Forschungsverbund an der Schnittstelle von Regionalsprachenforschung, Computerlinguistik und KI-Forschung versteht. Das Netzwerk ist über das Forschungszentrum Deutscher Sprachatlas erreichbar.

    Gerlach, Lisa. 2026. Vom Dorfplatz zum Chatbot – Regionale Sprache im Zeitalter Künstlicher Intelligenz. Sprachspuren: Berichte aus dem Deutschen Sprachatlas 6(3). In: Sprachspuren: Berichte aus dem Deutschen Sprachatlas 6(3). https://doi.org/10.57712/2026-03

    Mitglieder

    siehe Startseite