AmbiGest: A Dataset of Social Gestures with Inter-Class Similarity and Intra-Class Variability

Accepted to the LFA Workshop at FG 2026

Hajra Anwar Beg, Mohamed Daoudi, Angela Bartolo
Univ. Lille / IMT Nord Europe   |   SCALab, Univ. Lille, CNRS

Abstract

We propose AmbiGest, a dataset of dyadic social gestures designed for action recognition, interaction understanding, and reaction generation. AmbiGest contains 117 short clips organized into four communicative categories: greetings, calling, refusal, and thanks.

The dataset is designed around two central challenges: intra-class variability, where the same communicative intent can be performed in different styles, and inter-class similarity, where different intents may share visually similar motion patterns.

To preserve participant anonymity, raw RGB videos are not released. Instead, AmbiGest provides extracted 3D SMPL-X representations together with hierarchical annotations.

Dataset Overview

AmbiGest focuses on subtle everyday social gestures rather than large-scale action categories. It is designed to evaluate whether models can reason about communicative intent when different gestures look similar or when the same intention is expressed in different ways.

  • 117 dyadic gesture sequences.
  • 4 communicative categories: greetings, calling, refusal, and thanks.
  • 3 actor pairs for studying interpersonal variability.
  • Hierarchical annotations: class, subclass, semantic group, actor-pair ID, take index, and natural language description.
Overview of the AmbiGest dataset

Key Design

AmbiGest deliberately combines two properties that make social gesture understanding challenging:

  • Intra-class variability: the same intent can be performed in different styles, amplitudes, and timings across actor pairs.
  • Inter-class similarity: different intents can share similar motion patterns, such as a wave used for greeting and a wave used to call someone.

This design encourages models to move beyond coarse motion cues and to reason about social meaning.

Video Examples

The examples below illustrate the two main challenges targeted by AmbiGest: intra-class variability, where the same intention appears in different styles, and inter-class similarity, where different intentions may look visually close.

Intra-class Variability: Calling

The same communicative intent, calling someone, can be expressed with different gesture styles.

Calling
Hand beckon

Calling
Finger beckon

Calling
Hands around mouth

Intra-class Variability: Greetings

Greetings can appear in different social forms, from formal handshakes to casual or enthusiastic gestures.

Greetings
Formal handshake

Greetings
Enthusiastic handshake

Greetings
Elbow bump

Inter-class Similarity: Wave-like Gestures

Different communicative intents can share similar motion patterns. A wave, for example, can be used to greet someone or to call someone.

Greetings
One-hand wave

Calling
Wave to call

Inter-class Similarity: Ambiguous Hand and Arm Gestures

Different communicative intents can rely on visually similar hand or arm movements. These examples show gestures that all involve salient upper-body motion, but correspond to different social meanings: calling, greeting, and refusal.

Calling
Raise hand then beckon

Greetings
One-hand wave

Refusal
Head shake with one finger

Motion Representations

AmbiGest provides two complementary 3D reconstruction outputs:

  • PromptHMR: preserves the global placement and relative spatial positioning of the two subjects, making it useful for interaction-level spatial reasoning.
  • SMPLest-X: provides more expressive whole-body motion and finer hand articulation, but with less reliable relative placement between the two subjects.

Releasing both representations allows researchers to select the representation that best matches their task or to compare performance across reconstruction pipelines.

Baseline Experiments

We evaluate gesture recognition using a leave-one-couple-out protocol, where two actor pairs are used for training and the remaining pair is used for testing. Two temporal models are compared: a TCN and a Transformer encoder.

The results show that interaction context is important for resolving ambiguity. With PromptHMR features, actor-only input is limited, while adding the reactor and translation improves performance, reaching 91.3% accuracy with the Transformer. With SMPLest-X features, the best Transformer configuration reaches 89.0%.

Citation

@inproceedings{beg2026ambigest,
  title={AmbiGest: A Dataset of Social Gestures with Inter-Class Similarity and Intra-Class Variability},
  author={Anwar Beg, Hajra and Daoudi, Mohamed and Bartolo, Angela},
  booktitle={2026 International Conference on Automatic Face and Gesture Recognition (FG), LFA Workshop},
  year={2026}
}

Acknowledgements

This project is supported by the France 2030 program, the FR CNRS 2052 Sciences et Cultures du Visuel, and the ANR Investissements d’Avenir program Equipex+ Continuum.