Natural Language Understanding

The Natural Language Understanding group works at the intersection of machine learning and natural language processing, with an emphasis on representation learning for the meaning of language, attention-based deep learning models, and structured prediction. 

Introduction

Understanding language is essential to understand intelligence.  The empirical success of transformer language models at understanding language shows the effectiveness of attention mechanisms in deep learning architectures for AI.  The NLU group focuses on how to understand and improve attention-based representation learning from the perspectives of both graph embeddings and information theory.  We target various language tasks, such as reasoning, knowledge extraction, controlled generation, privatisation, and linguistic structure prediction. 

Graph-to-Graph Transformer:  We have developed attention mechanisms to input, embed, and predict the nodes and edges of a graph.  Currently we are investigating knowledge graph embeddings, and unsupervised induction of linguistic structures. 

Nonparametric Variational Information Bottleneck:  We have developed attention mechanisms which infer probability distributions over attention-based representations.  Currently we are using NVIB to add a controllable memory to LLMs, induce interpretable abstract representations, privatise text embeddings, and improve abstract reasoning abilities

Alumni

ABBASI, Ali
BEHJATI, Melika
BHATT, Chidansh
COMAN, Andrei
ESPINOSA MENA, Jose Rafael
FEHR, Fabio
GASNIER, Catherine
HABIBI, Maryam
HAJLAOUI, Najeh
HONNET, Pierre-Edouard
KARIMI MAHABADI, Rabeeh
LE, Quoc Anh
LI, Kexing
LISON, Pierre
LIYANAPATHIRANA, Jeevanthi Uthpala
LOAICIGA, Sharid
LUONG, Ngoc-Quang
MAHDABI, Parvaz
MAI, Florian
MARFURT, Andreas
MATENA, Lukas
MEYER, Braida
MEYER, Thomas
MICULICICH, Lesly
MOHAMMADSHAHI, Alireza
MRINI, Khalil
PAPPAS, Nikolaos
PATEL, Kumar
PILAULT, Jonathan
POPESCU-BELIS, Andrei
PU, Xiao
RAZGHANDI, Ali
REKABSAZ, Navid
YAZDANI, Majid

Ongoing projects

BALM

We address the controllability of large language models (LLMs) by giving them interpretable beliefs and programmable knowledge through leveraging the PI's work on understanding and improving transformer embeddings. Transformers' empirical success comes from the attention function's ability to induce graphs of relations from text. Our recent work has extended this ability to knowledge graphs, and to inducing the nodes of the graph as well, known as entity induction, with the first variational-Bayesian generalisation of the attention mechanism. This project will further develop this information-theoretic understanding of transformer embeddings and its sparsity-inducing regulariser, for learning graphs of higher-level abstract entities. The resulting Bayesian beliefs over generalised transformer embeddings of texts and graphs will give us the more interpretable, more programmable and more learnable abstract representations which are the core of this proposed project.

To leverage and extend these fundamental advances in representation learning, we will develop LLM architectures with a memory. Motivated by the success of Retrieval Augmented LLMs, our Belief Augmented Language Models (BALMs) will move knowledge extracted from training data out of large uninterpretable weight matrices into our interpretable Bayesian beliefs over large transformer embeddings. These beliefs will then be: augmented with human-editable knowledge graphs and selected new texts, refined with control objectives and multi-hop reasoning, and combined with inference of concensus beliefs and opinion summarisation. BALMs will be developed both to evaluate these beliefs and as a chat interface for specifying, accessing and editing the beliefs themselves, including the collaborative specification of shared beliefs. These fundamental advances in deep learning theory and architectures will allow us to control what an LLM says by controlling what it believes, thereby unlocking the power of AI for society.

BOVINE

In recent years, attention-based models like Transformers have radically improved the performance of natural language understanding (NLU), demonstrating the appropriateness of attention-based representation for language. In (Henderson, 2020) we show that these representation share many characteristics with those found in traditional computational linguistics (e.g. graph structure), except that they do not automatically learn multiple levels of representation nor their entities (morphemes, phrases, discourse entities, etc). Motivated by this challenge of entity induction, our recent work has discovered a very non-traditional perspective, which characterises attention-based models like Transformers as doing nonparametric Bayesian inference. Given an input text, our Nonparametric Variational Information Bottleneck (NVIB) Transformer infers distributions over nonparametric mixture distributions (Henderson and Fehr, 2023). We have even shown that pretrained Transformers can be converted into equivalent NVIB Transformers, and regularised post-training (Fehr and Henderson, 2023).

This reinterpretation of Transformers, combined with their unprecedented empirical success, leads us to postulate the hypothesis that natural language understanding is nonparametric variational Bayesian inference over mixture distributions. This claim of the adequacy of NVIB leads to two fundamental challenges which are not currently being addressed, each with an associated technological aim:

1. How can NVIB support inducing graph-structured representations at multiple levels of representation?
   Making deep learning representations interpretable.

2. How can NVIB enable controlling the information in representations?
   Making deep learning representations controllable.

For the first challenge, we will extend our previous structure processing methods (Mohammadshahi and Henderson, 2020, 2021, 2023; Miculicich and Henderson, 2022), developed for set-of-vector representations, to mixture-of-component distributions. And we will focus on unsupervised learning methods, rather than our previous supervised learning methods. To extend these models to multiple levels, we will take the approach of embedding all levels in one big mixture of non-homogeneous components, which are computed with iterative refinement. This extends our previous work on iterative graph refinement (Mohammadshahi and Henderson, 2021; Miculicich and Henderson, 2022), adding the induction of the nodes of the graph and the induction of multiple levels of representation. Learning representations which are interpretable as linguistic structures will be a testbed for the general aim of deep learning of interpretable representations.

For the second challenge, we will leverage the information theory behind NVIB to model both inferring implicit information and removing private information. We will investigate the use of KL divergence as a measure of entailment in semantic inference. We will apply the framework of Rényi differential privacy (Mironov, 2017) to provide privacy guarantees by adding noise which removes targeted information from Transformer embeddings. This method extends differential privacy to anything that can be embedded with a Transformer (especially text), with many important applications. These methods address the general aim of controlling the information in deep learning representations.

Addressing these challenges will lead to fundamental advances in machine learning, including novel deep learning architectures and fundamental insights into Transformers and their pretraining. We will do both intrinsic and extrinsic evaluations of our induced representations, expecting to show improvements on core NLP tasks, including privacy-preserving sharing of textual data. Given the current level of interest in the AI research and development community for Transformers and variational Bayesian methods, we expect the proposed research to have a profound impact on the field.

EVOLANG-2

Language is what sets humans apart from all other species. Despite much effort, however, its evolutionary origins have remained obscure. At the same time, the role of language is currently undergoing radical changes, with cultural, psychological and evolutionary ramifications barely understood. New digital channels, ubiquitous online knowledge bases, and continued advancement of artificial intelligence are reshaping our communicative environment and modifying the way we learn and use language. An indepth exploration of the origins and future of language is urgently needed, propel ling language science to the forefront of societal and economic challenges. Our project explores the evolutionary origins and future development of linguistic communication with an unprecedented transdisciplinary research programme. We conceptualise language as a system of components with distinct evolutionary trajectories and adopt a large-scale comparative framework to study these trajecto ries in nature and function along three thematic axes: 1. The Dynamic Structures of Language: How and why have the structures of language and their temporal dynamics evolved? How will these structures interact with new technologies and means of communication? 2. The Biological Substrates of Language: What are the biological mechanisms that make language possible? Can and should we intervene on language functions with neurotechnology? 3. The Social Cognition of Language: What are the social cognitive mechanisms that underlie linguistic communication, both phylogenetically and ontogenetically? How did these mechanisms evolve and how will they change with artificial communicators? Our long-term vision is a fully-fledged phylogeny of language, tracing the evolution of the core components from earlier forms of communication and cognition. This phylogeny is expected to shed light on the drivers of linguistic and cognitive complexity and on the extent to which the digital age might create new niches in which human communication evolves. We posit that only a radically evolutionary perspective can lead to a sufficiently deep understanding of the current changes in human communication, to a capacity to predict how human communication might develop in the future, and to competence in taking ethically responsible decisions in relation to technological developments. We will tackle these questions by leveraging cutting-edge developments in language science, neuroscience, computer science, and evolutionary biology. In language science, the digital revolution makes it possible to model ontogeny and diachrony species-wide and cross-culturally, refocusing the question of language origins from static to dynamic traits. In neuroscience, we can describe and model language and speech processing with unprecedented biological plausibility, allowing intervention on language functions with sophisticated neuroengineering tools. In computer science, machine intelligence systems allow language analysis and processing with striking efficiency and accuracy. In evolutionary studies, comparative research with primates and other animals living in natural conditions has revolutionised current theories of animal cognition, with direct implications for the origins of language. Our proposal for a National Centre of Competence in Research brings together a large group of scientists in Switzerland with the shared goals of (1) understanding the evolution of language-ready brains and their neural mechanisms, the social conditions these mechanisms require, and the dynamic structures they produce; (2) designing new applications and neuroengineering techniques for language learning, assessment, recovery, translation, and disorder remediation; and (3) engaging the public in scientifically informed debates on the prospects, challenges, and ethics of digital communication and neurotechnological interventions. Language and its dynamic diversity is part of the national identity of multilingual Switzerland, making this project ideally suited to inspire and engage a large general public with cutting-edge research.

Past projects

AROLES

The AROLES project aims at transferring scientific know-how in the domains of audio-visual processing and multimedia retrieval from the IM2 NCCR to the Klewel SME dedicated to lecture capture and web-based broadcasting. The transferred know-how and technology are intended to add functionalities for multimedia recommendation and skimming to the Klewel portal. This will increase its attractiveness for its end-users, and as a consequence will make Klewel’s business proposition more competitive for its customers, who are the creators of the multimedia content, such as conference or event organizers. More specifically, the AROLES project has three technical goals. The first goal is to develop a system for recommendation of lectures and snippets from a large multimedia archive, in Klewel’s context, using IM2 know-how in multimodal signal processing and semantic relatedness measures. Recommendations will be based on a mixture of content-similarity with the currently viewed recording and of collaborative filtering i.e. similarity of viewing patterns from Klewel users who viewed the current recording. The second goal is to extend this technology to the recommendation of short snippets, rather than entire lectures. Finally, the third goal is to use the previous achievements and additional multimodal processing to extract short multimedia summaries of the lectures, intended for users who only want to skim through a lecture. The algorithms will be developed by a post-doc at Idiap (IM2 NCCR) for 24 months full-time, while the data preparation and access, the user interface, and the evaluation will be done by a Klewel engineer supported at 75%.

COMTIS

Machine translation (MT) has made significant progress in the past decade, but its focus has remained on the translation of sentences considered individually. However, in order to ensure overall coherence throughout a translated text, an MT system must also consider and render correctly the items that depend on intersentential relations. The perceived coherence of a translated text, and therefore its overall quality, are mainly influenced by the following markers: pronouns, verb tense/mode/aspect, discourse connectives, and politeness/style/register. None of these markers can be reliably translated on a pure sentence-by-sentence basis. This project aims at extending the current statistical MT (SMT) approach by modeling these intersentential dependencies (ISDs), along the following five themes. Linguistic analysis, corpus data, annotation and test suites, automatic identification of intersentential dependencies, statistical machine translation for ISD-labeled texts, evaluation methods for MT coherence and their application

DOMAT

Statistical and neural machine translation systems (in short, SMT and NMT) have reached significant quality and speed levels. Such systems use large amounts of monolingual and bilingual data to train their models and tune their meta-parameters. As a result, translating a sentence from a language unknown to a user can be done with acceptable quality, and hence a clearly perceived utility. However, the translation of complete texts is still far from publishable, and requires substantial post-editing by humans. One reason for this difference is that certain linguistic constraints cannot be reliably translated using only local information, especially when they apply across different sentences. The homogeneous models used by SMT or NMT systems - very large translation tables or connection weights - are a strength for robust and quick sentence translation, but impose a strong limitation when constraints of different ranges must be taken into account to translate a document. In recent years, I have pioneered a method to address document-level problems that degrade MT quality, drawing on specific linguistic knowledge as required by each problem. These methods improved the translation of discourse connectives or verb tenses based on sentence-level or document-level semantic features, or constrained the choice of referring expressions such as pronouns and noun phrases. For implementation, the methods took advantage of existing approaches, such as factored models, to integrate linguistic knowledge with SMT. However, integrating several solutions for dealing with document-level constraints into a unified system is not tractable with current approaches, due to the fact that knowledge sources are heterogeneous and are not tightly coupled with the MT systems. Moreover, the quality improvements brought by leveraging several distinct knowledge sources may not add up because of interaction between them. Finally, the need for computing all features for all words or sentences of a document raises strong efficiency issues. In the DOMAT project, we aim to design a novel approach for providing on-demand linguistic knowledge to statistical or neural MT systems. Both types of systems will be considered to provide comparison terms, as they both have strengths and weaknesses. The linguistic knowledge will be learned by specific processing modules, which will extract and output features in a format that is usable by SMT and NMT systems. To make this architecture operational, we will explore strategies to trigger the modules, for instance based on quality estimation or translation confidence. To populate the architecture, we will build several modules to extract document-level features that are relevant to translation, principally document structure (discourse relations) and coreference (including pronominal anaphora). The starting points for these modules will be our previous achievements in document-level SMT. The DOMAT project will mainly support two PhD theses at Idiap/EPFL, one on designing and comparing statistical and neural architectures for integrating and triggering on-demand knowledge sources in MT, and the other one designing such knowledge sources, which learn specific text-level constraints and output suitable data structures for NMT. The solutions developed in DOMAT will make tractable the demands for adequate, 07.04.2017 08:34:45 Page - 5 - fluent and efficient translation of large documents, and will result in a principled approach for learning high-level linguistic knowledge to improve translation quality.

EVOLANG

Language is what sets humans apart from all other species. Despite much effort, however, its evolutionary origins have remained obscure. At the same time, the role of language is currently undergoing radical changes, with cultural, psychological and evolutionary ramifications barely understood. New digital channels, ubiquitous online knowledge bases, and continued advancement of artificial intelli gence are reshaping our communicative environment and modifying the way we learn and use lan guage. An indepth exploration of the origins and future of language is urgently needed, propel ling language science to the forefront of societal and economic challenges. Our project explores the evolutionary origins and future development of linguistic communication with an unprecedented transdisciplinary research programme. We conceptualise language as a system of components with distinct evolutionary trajectories and adopt a large-scale comparative framework to study these trajecto ries in nature and function along three thematic axes: 1. The Dynamic Structures of Language: How and why have the structures of language and their temporal dynamics evolved? How will these structures interact with new technologies and means of communication? 2. The Biological Substrates of Language: What are the biological mechanisms that make language possible? Can and should we intervene on language functions with neurotechnology? 3. The Social Cognition of Language: What are the social cognitive mechanisms that underlie linguistic communication, both phylogenetically and ontogenetically? How did these mechanisms evolve and how will they change with artificial communicators? Our long-term vision is a fully-fledged phylogeny of language, tracing the evolution of the core components from earlier forms of communication and cognition. This phylogeny is expected to shed light on the drivers of linguistic and cognitive complexity and on the extent to which the digital age might create new niches in which human communication evolves. We posit that only a radically evolutionary perspective can lead to a sufficiently deep understanding of the current changes in human communication, to a capacity to predict how human communication might develop in the future, and to competence in taking ethically responsible decisions in relation to technological developments. We will tackle these questions by leveraging cutting-edge developments in language science, neuroscience, computer science, and evolutionary biology. In language science, the digital revolution makes it possible to model ontogeny and diachrony species-wide and cross-culturally, refocusing the question of language origins from static to dynamic traits. In neuroscience, we can describe and model language and speech processing with unprecedented biological plausibility, allowing intervention on language functions with sophisticated neuroengineering tools. In computer science, machine intelligence systems allow language analysis and processing with striking efficiency and accuracy. In evolutionary studies, comparative research with primates and other animals living in natural conditions has revolutionised current theories of animal cognition, with direct implications for the origins of language. Our proposal for a National Centre of Competence in Research brings together a large group of scientists in Switzerland with the shared goals of (1) understanding the evolution of language-ready brains and their neural mechanisms, the social conditions these mechanisms require, and the dynamic structures they produce; (2) designing new applications and neuroengineering techniques for language learning, assessment, recovery, translation, and disorder remediation; and (3) engaging the public in scientifically informed debates on the prospects, challenges, and ethics of digital communication and neurotechnological interventions. Language and its dynamic diversity is part of the national identity of multilingual Switzerland, making this project ideally suited to inspire and engage a large general public with cutting-edge research.

Latest publications

Fast-and-frugal text-graph transformers are effective link predictors
Coman Andrei Catalin, Theodoropoulos Christos, Moens Marie-Francine, Henderson James
Findings of the association for computational linguistics
2025
Nonparametric variational regularisation of pretrained transformers
Fehr Fabio, Henderson James
First conference on language modelling
2024
Nonparametric variational regularisation of pretrained transformers
Fehr Fabio, Henderson James
ArXiv
2023
A VAE for transformers with nonparametric variational information bottleneck
Henderson James, Fehr Fabio
The eleventh international conference on learning representations
2023
Recursive non-autoregressive graph-to-graph transformer for dependency parsing with iterative refinement
Mohammadshahi Alireza, Henderson James
Transactions of the Association for Computational Linguistics (2021)
2021
Recursive non-autoregressive graph-to-graph transformer for dependency parsing with iterative refinement
Mohammadshahi Alireza, Henderson James
Transactions of the Association for Computational Linguistics(under submission)
2020