Perception & Activity Understanding

We develop scene understanding and behavioral foundation models to decode human activities and behaviors from real-world sensors, spanning computer vision, deep learning, multimodal signal processing, and the social sciences. Our research focuses on recognizing and interpreting non-verbal behaviors, from gaze and attention to gestures and social dynamics, with applications in health, robotics, and industrial systems. 

Introduction

The group develops scene understanding and behavioral foundation models using methods from machine learning, computer vision, multimodal signal processing, and the social sciences to understand activities and behaviors from real-world sensors and data, with an emphasis on those related to humans. While detection, tracking, or pose estimation were previous research focuses, the group now focuses on recognition and analysis of non-verbal behaviors, and the temporal interpretation of lower-level signals in higher level constructs including gestures, activities, communication and social behaviors, social relationships, personality traits, or medical indices. In particular, the group has strong expertise in designing models for decoding gaze and attention, which are central to cognitive processes involving intentions, actions, and communication, providing a window into social cognition.  Applications include surveillance, animal life monitoring, human behavior analysis, mental health assessment, human interactions, and social robotics. More broadly, the group is also interested in developing image processing systems designed to address real-world challenges, including industrial applications. 

 

Alumni

ALI, Abid
AMINIAN, Bozorgmehr
BA, Silèye
BENGIO, Samy
CAN, Gulcan
CAO, Yuanzhouhan
CHAKRABORTY, Shayok
CHAVANE, Cécile
CHEN, Yiqiang
CHEN, Le
CHEN, Cheng
CHIAPPA, Silvia
CHOCKALINGAM, Thiyagarajan
CHOUDHARY, Anand
CHUTISILP, Naravich
CROVETTO, Gianna Larissa
CUTILLO, Leucio Antonio
DE CAMPOS, Ruben
DE OLIVEIRA PRESTES, Lukas
DESPRÈS, Nicolas
DE WAHA-BAILLONVILLE, Gilles
DIMITRAKAKIS, Christos
DORSAZ, Ludmilla
DROZDZAL, Michal
DUVAL, Matthieu
EMONET, Rémi
EUTAMENE, Camil
FARKHONDEH, Arya
FORMAZ, Magali
FUCHS, Michael
FUNES MORA, Kenneth Alberto
GAY, Paul
GERVAISE, Lara
GEVERS, Louis
GILARDI, Nicolas
GIRAUD, Raphaël
GRÄSSLE, Lukas
GUPTA, Anshul
HEILI, Alexandre
HU, Xu
HU, Hainan
JIN, Wanjun
KANEVSKI, Mikhael
KHALIDOV, Vasil
LAURINEN, Perttu
LE, Nam
LE, Quan
LEFÈVRE, Stéphanie
LHUISSIER, Miguel
LIU, Gang
LIU, Chongyang
LOPEZ-MENDEZ, Adolfo
MAHMOUDIAN, Navid
MARTINEZ GONZALEZ, Angel
MUGNIER, Tom
NAKATANI, Chihiro
OERTEL, Catharine
OKAL, Billy
PANNATIER, Martin
RACCA, Mattia
RICCI, Elisa
ROMAN RANGEL, Edgar Francisco
SCHEFFLER, Carl
SCHIFFERLE, Patrick
SHAIK, Abu
SHARIF, Mohamed Amjed
SHEIKHI, Samira
SIEGFRIED, Rémy
SOUSA EWERTON, Marco
STEL, Lucas
TAVENARD, Romain
VARADARAJAN, Jagannadan
VARADARAJAN, Karthik Mahesh
VERZAT, Philippine
WANG, Anlan
WU, Di
YAO, Jian
YONEZAWA, Tomoko
YU, Yu
ZHANG, Xiaocheng

Ongoing projects

GAMOE

Several key aspects for joint attention modeling remain unexplored and underdeveloped. Addressing these gaps constitutes the primary focus of the present proposal. More precisely, we will work on:

  • Multimodal integration of interactional cues. So far only visual and pose cues were included in the models, without incorporating speaking status information and transcripts.
  • Object-level modeling: we will include referential objects enabling more precise modeling of joint attention and word–referent mapping.
  • Tool refinement: Enhance annotation tools to be scalable and user-friendly for behavioral research.

The key goals and innovation of the present project lie in the multimodal integration of language and interaction cues with automated verification of objects’ presence, while measuring how attention is directed toward these objects during communication. This approach is novel in combining gaze, language, and referential object information, and will be validated in naturalistic, out-of-the-lab field settings, providing a scalable and ecologically valid framework for modeling joint attention and word-referent mapping. Through these advancements, we aim to set new standards in predictive modeling of joint attention and develop scalable tools for the community to advance the study of early language development. The project mainly brings together the expertise in gaze modeling (Idiap) and early child language development (UZH).

TESSELLARIUS

Le projet propose de développer une installation événementielle incluant vision, IA et robotique pour la création automatique d’une mosaïque à partir de fragments aux formes et couleurs variées, issus du recyclage. L’installation est destinée à une exposition publique, évoquant à la fois l’origine romaine de Martigny (Martigny-la-Romaine, Octodure), tout en montrant les avancées de l’IA et de la robotique dans les domaines créatifs et artistiques, avec des IA à la fois physique (robots) et computationnelle (vision, image processing, analyse et classification automatique de formes, optimisation géométriques de placements).


Très utilisée pendant l'Antiquité romaine, la mosaïque est un art décoratif dans lequel on utilise des fragments de pierre, d'émail, de verre, ou de céramique pour former des motifs ou des figures en assemblant et orientant les fragments de manière judicieuse et artistique selon les éléments et contours essentiels de la figure à représenter. Notre projet propose de revisiter cet art du fragment provenant de la Rome antique, en l’adaptant aux nouvelles technologies de l’IA et de la robotique pour trier, analyser et classifier automatiquement les fragments mis à disposition, planifier le choix d’utilisation de ces fragments, leurs placements et leurs orientations sur la fresque à réaliser. Ces

fragments, qui peuvent être de différentes formes (en plus de la variété de couleurs et d’orientations possible), seront automatiquement assemblés pour représenter une image souhaitée, en discussion avec la ville de Martigny, définie selon leur souhait et l’événement ciblé (par exemple, hommage à Léonard Gianadda, fresque du Château de la Bâtiaz, Fondation Barry du Grand-Saint-Bernard, dessin correspondant au thème d’un événement organisé par la ville comme le carnaval ou la foire du Valais, etc.).


The project proposes to develop an event-based installation incorporating vision, AI, and robotics for the automatic creation of a mosaic

from recycled fragments of varying shapes and colors. The installation is intended for a public exhibition, evoking both the Roman origins of Martigny (Martigny-la-Romaine, Octodure) and showcasing the advancements of AI and robotics in creative and artistic domains, employing both physical (robots) and computational AI (vision, image processing, automatic shape analysis and classification, and geometric optimization of placements).


Widely used during Roman Antiquity, mosaic is a decorative art form in which fragments of stone, enamel, glass, or ceramic are used to

create patterns or figures by judiciously and artistically assembling and orienting the fragments according to the essential elements and contours of the figure to be represented. Our project proposes to revisit this art of the fragment from ancient Rome, adapting it to new AI and robotics technologies to automatically sort, analyze, and classify the available fragments, and plan their use, placement, and orientation within the artwork to be created. These fragments, which can be of various shapes (in addition to the variety of possible colors and orientations), will be automatically assembled to represent a desired image, in consultation with the city of Martigny, defined according to their wishes and to the targeted event (for example, a tribute to Léonard Gianadda, a canvas representing Le Château de la Bâtiaz, the Barry Foundation of the Grand-St-Bernard or a design corresponding to the theme of an event organised by the city such as the carnival or the Foire du Valais, etc).


TUNASBE

Understanding and interpreting people’s states, activities, and behaviors stands as a foundational objective in the field of artificial intelligence. Indeed, the ability to discern non-verbal behaviors, interpret social signals, and infer the mental states of others are crucial components of social intelligence, enabling individuals to predict and understand the behavior of others, anticipate their reactions, gauge their interest, and engage in more meaningful and effective social exchanges.

In particular, gaze and attention form central elements driving many cognitive processes related to intentions, actions, and communications, with many connections with other behavioral cues like head and hand gestures, speaking status, facial expressions, or interactions. While past research efforts attempted to model these cues, the associated tasks were often treated independently which forgoes any possible relationships between them, or in specific settings.

In light of this, the goal of this research project is to develop holistic and comprehensive models for human attention and social behavior understanding in the wild, where videos feature a wide diversity of scenes, environments, people, objects, and activities.  This implies the integration of social prediction tasks at two different levels: (i) subject-level, where we seek to model head dynamics and facial behavior (e.g. facial expressions, gaze estimation, head/hand gestures), and (ii) scene-level,  by analyzing people’s states and behavior to infer gaze following patterns (where and at what a person looks at), attentiveness, interactions, and social communication, which requires context modeling.

The core idea is to leverage a multi-task co-training framework to model these tasks jointly, thereby accounting for intra-task and inter-task dynamics. The unified model is poised to bring the following benefits: (a) efficiency owing to a single model supporting multiple tasks, (b) improved performance by virtue of multi-task supervision, and (c) strong person/scene representations that can transfer well to person-centric downstream tasks.

Achieving the desired result hinges upon addressing multiple challenges related to the specific tasks themselves, or the overarching unification objective. These can be framed as a set of research questions that will drive our investigations: (a) How can we leverage vision-language models to incorporate a semantic layer into the gaze following task? (b) How should we model head-face-gaze dynamics to infer head gestures and gaze directions? (c)  How can we represent person-centric information given head and body streams? (d) How can we exploit the graph message passing framework or cross-attention models to design fusion mechanisms of subject-level and scene-level information? (e) How to curate data for the tasks of interest, and capitalize on label propagation to learn from a combination of heterogeneous datasets and annotations?

Given the unification end-goal, our approach is to leverage a consistent token-based representation in all our tasks, thereby ensuring compatibility between the different architectural components developed separately. We envision the final architecture to include scene and person transformer encoders followed by a fusion module to allow the exchange of information between each person and the scene or other people. Finally, the updated person tokens will serve as input queries to a transformer decoder before feeding into task-specific prediction heads.

By investigating unified computational models of social behavior and communication in the wild, including how visual attention is influenced by and coordinated with other cues, we expect to impact the research community by advancing the state-of-the-art in human-centric computer vision, providing novel tasks and benchmarks, and contribute to other disciplines that rely on such models to automatically code these behaviors for subsequent large-scale analyses.

Past projects

3D2CUT

The main objective of this project is to verify the feasibility of using innovative artificial intelligence algorithms to analyze vine based on vineyard images. The aim is to automatically extract the essential components and use them to recommend appropriate pruning.

ADVANCE

The project provides an augmented dialogue tool that exploits verbal and non-verbal indices to improve interviews’ quality. Its main business application is to support HR interviews.

AI4AUTISM2

Nowadays, 1 in 59 children is diagnosed with autism spectrum disorders (ASD), which makes this condition one of the most prevalent neurodevelopmental disorders. The hereby project is grounded on the recognition that, on the one hand, early diagnosis at scale of autism in young children requires the development of tools for digital phenotyping and automated screening, through computer vision and Internet of Things sensing. On the other hand, current gold-standard approaches in autism are not intended to provide a precise quantitative estimate of ASD symptoms in children. We therefore aim to examine the potential of digital sensing to provide automated measures of the extended autism phenotype, for the purpose of stratifying autism subtypes in ways that would allow for precision medicine. Recent developments in digital sensing, big data and machine-learning have offered unforeseen opportunities for seamless sensing of body movement, social scene capture, and measure of object manipulation. Together, these tools are key for modeling social interactions and offer avenues for both improving screening and fine-grained characterization of autistic symptoms in young children. Despite considerable efforts invested to explore such automatic behavioral analysis, most studies in ASD digital phenotyping have been conducted on modest samples sizes, used mono-modal approaches, were focused on eliciting very specific behaviors by largely controlled prompts, and have suffered from technical difficulties in behavior sensing (view points, children population, image resolution for gaze). To address these limitations, we propose an interdisciplinary project combining the skills of experts in clinical research, engineering and computational social sciences in order to address these clinical, scientific, and technical challenges. It is grounded on the Geneva Autism Cohort consisting of young children with ASD and their age-matched typically developing peers, extensively assessed with gold standard standardized clinical and cognitive assessments, as well as neuroscience tools. Further, our preliminary results demonstrated that relying on a substantial dataset it is feasible to successfully train a deep neural network directly from on a global scene representation (people poses) to predict ASD with above 80% accuracy. This Sinergia proposal is set to stretch a giant leap forward, by investigating three key research directions. First, from a clinical research perspective, we will design digital tools for screening and automated profiling of autism phenotype. We will test these tools in a structured setting with well-established clinical protocol, as well as in a less structured environment (free play in day-care centers). Second, with Internet of Things (IoT) sensors, we will investigate the motor skills of very young children, through the integration of inertial and low-cost UWB indoor localization data. Additionally, we will develop a solution for the longitudinal monitoring of fine-grained motor skills development. Last but not least, our project is rooted in modern computational perception and machine learning. We will investigate novel deep learning and computer vision techniques by leveraging the availability of large behavioral and clinical annotation data. At the core of this effort, we will develop multimodal machine-learning methods and models for the analysis of motor and gaze coordination patterns which are at the core of ASD, and for ASD diagnosis and profiling with a focus towards interpretable models.

ASLEEP

Latest publications

Learning ego-exo visual representations for conversational gaze estimation
Gupta Anshul, Qian Yijun, Gao Ruohan, Ananthabhotla Ishwarya, Odobez Jean-Marc, Krishna Ithapu Vamsi, Murdock Calvin
Conference on computer vision and pattern recognition workshops
2026
Loose social-interaction recognition in real-world therapy scenarios
Ali Abid, Dai Rui, Marisetty Ashish, Astruc Guillaume, Thonnat Monique, Odobez Jean-Marc, Thümmler Suzanne, Bremond Francois
IEEE/CVF winter conference on applications of computer vision
2025
MTGS: a novel framework for multi-person temporal gaze following and social gaze prediction
Gupta Anshul, Tafasca Samy, Farkhondeh Arya, Vuillecard Pierre, Odobez Jean-Marc
38th conf. on neural information processing system
2024
Towards smart pruning: ViNet, a deep-learning approach for grapevine structure estimation
Gentilhomme Théophile, Villamizar Michael, Corre Jérome, Odobez Jean-Marc
Computers and Electronics in Agriculture
2023