Portrait of Mokhtar Z. Alaya

Mokhtar Z. Alaya

Maître de Conférences

LMAC – UTC

Papers

Publishing high-quality papers has always been my constant goal.

Filter by theme:
Filter by type:
Sort by year:

2026

Does Normalization Choice Matter for Causal Large Time-Series Models?

Samy-Melwan Vilhes, Gilles Gasso, Mokhtar Z. Alaya

Conference "Spotlight" Workshop on Time Series in the Age of Large Models, ICLR 2026

Theme: Deep Learning for Time Series

ICLR Code OpenReview

Figure from the paper

Large models for time-series forecasting have been emerged as a promising paradigm for training models on heterogeneous collections of signals. These models typically rely on causal autoregressive architectures, where each observation is sequentially predicted from past. In practice, real-world time-series exhibit non-stationarities, which significantly influence predictive performance. To mitigate this, normalization is commonly employed. However, in efficient causal settings it might induce information leakage from future observations during training. Recent alternatives, including causal normalization and statistics computed from initial observations, have been proposed to address this issue, but their practical implications remain insufficiently understood. In this work, we evaluate normalization strategies for transformer-based large time-series models trained with patching and efficient causal strategy. We showcase that normalization choice significantly influences both training convergence and forecasting performance.

Composites Multi-source Sensor Data Fusion Framework for Structural Health Monitoring of Polymer-Matrix Composites (PMC) Based on Latent-Space Clustering Using a Convolutional Autoencoder

Loan Dolbachian, Mokhtar Z. Alaya, Walid Harizi, Imen Gnaba, Zoheir Aboura, and Salim Bouzebda

Journal Mechanics of Advanced Materials and Structures, 2026

Theme: Interdisciplinary Research (Artificial Intelligence for Mechanics, AI4Meca)

Figure from the paper

This study presents an integrated framework combining the data fusion and clustering techniques with Neural Networks (NNs) for the processing and classification of heterogeneous experimental data. The proposed approach establishes a synergistic link between materials science and data science, enabling enhanced interpretation of the complex multi-sensor datasets. A convolutional autoencoder was employed to cluster many data acquired from the load–unload tensile tests conducted on smart polymer matrix composite (PMC) specimens embedding two Lead Zirconate Titanate (PZT) and one Polyvinylidene Fluoride (PVDF) transducers. The PZT sensors provided high-sensitivity electrical data reflecting the internal mechanical responses, while the PVDF transducer recorded complementary strain-related variations—together forming a robust dataset for Structural Health Monitoring (SHM). These internal measurements were complemented by external tools, including Digital Image Correlation (DIC) and Acoustic Emission (AE) systems, which supported both validation and benchmarking of the neural framework. The findings demonstrate that the proposed architecture effectively integrates and classifies diverse sensor signals, underscoring the necessity of clear, sensor-specific data structuring for reliable SHM applications.

Optimal Transport Convergence of Conditional Distribution Estimation for Single-Indexed Locally Stationary Functional Time Series

Jan N. G. Tinio, Mokhtar Z. Alaya, Salim Bouzebda

Journal Results in Applied Mathematics, 2026

Theme: Optimal Transport for Machine Learning

ScienceDirect

Figure from the paper

Modeling nonstationary functional time series has become increasingly important in applications where both temporal dynamics and functional structure play central roles. The framework of locally stationary functional time series (LSFTS) provides a flexible representation for such data by allowing their probabilistic features to evolve smoothly over time. At the same time, the infinite-dimensional nature of functional predictors motivates the use of functional single index models (FSIM) as an effective dimension-reduction strategy. In this paper, we develop a methodology for estimating the conditional distribution of LSFTS within the FSIM framework. We introduce a Nadaraya–Watson (NW) kernel estimator that simultaneously smooths across time and along a data-driven single index. Using tools from optimal transport, we derive uniform convergence rates under the 1-Wasserstein distance, accounting for small ball probabilities and weak dependence through beta-mixing conditions. We further propose a practical procedure for estimating both the single index and the bandwidth via cross-validation. Extensive simulations confirm the theoretical convergence behavior, and applications to real datasets — including financial and physiological time series — demonstrate the effectiveness of the approach for distributional forecasting. An additional case study on electricity price data illustrates how the proposed estimator can deliver competitive predictive accuracy in complex nonstationary environments.

Bounds in Wasserstein Distance for Locally Stationary Functional Time Series

Jan N. G. Tinio, Mokhtar Z. Alaya, Salim Bouzebda

Journal Computational Statistics, 2026

Theme: Optimal Transport for Machine Learning

Springer Code

Figure from the paper

We investigate nonparametric estimation of conditional distributions for locally stationary functional time series (LSFTS) within a general semi-metric framework. The response variable is scalar, while the covariate process evolves in an infinite-dimensional space and exhibits smooth temporal nonstationarity. We propose a Nadaraya–Watson-type estimator formulated as a locally weighted empirical measure that jointly localizes in rescaled time and functional covariate space. Under mild structural assumptions encompassing small-ball probability conditions, weak beta-mixing dependence, and minimal smoothness of the conditional distribution function, we establish explicit convergence rates for the estimator in Wasserstein distance. The analysis reveals how temporal localization, functional complexity, and dependence interact to govern distributional convergence in nonstationary functional settings. Extensions to higher-order Wasserstein metrics are derived under boundedness conditions. Comprehensive numerical experiments, including locally stationary functional autoregressive models and real-world datasets, corroborate the theoretical findings and illustrate the practical effectiveness of the proposed methodology for distributional inference beyond conditional means.

Sparsified-Learning for High-Dimensional Heavy-Tailed Locally Stationary Time Series, Concentration and Oracle Inequalities

Yingjie Wang, Mokhtar Z. Alaya, Salim Bouzebda, Xinsheng Liu

Preprint arXiv, 2026

Theme: Structured Statistical Learning

arXiv

Sparse learning is ubiquitous in many machine learning tasks. It aims to regularize the goodness-of-fit objective by adding a penalty term to encode structural constraints on the model parameters. In this paper, we develop a flexible sparse learning framework tailored to high-dimensional heavy-tailed locally stationary time series (LSTS). The data-generating mechanism incorporates a regression function that changes smoothly over time and is observed under noise belonging to the class of sub-Weibull and regularly varying distributions. We introduce a sparsity-inducing penalized estimation procedure that combines additive modeling with kernel smoothing and define an additive kernel-smoothing hypothesis class. In the presence of locally stationary dynamics, we assume exponentially decaying β-mixing coefficients to derive concentration inequalities for kernel-weighted sums of locally stationary processes with heavy-tailed noise. We further establish nonasymptotic prediction-error bounds, yielding both slow and fast convergence rates under different sparsity structures, including Lasso and total variation penalization with the least-squares loss. To support our theoretical results, we conduct numerical experiments on simulated LSTS with sub-Weibull and Pareto noise, highlighting how tail behavior affects prediction error across different covariate-dimensions as the sample size increases.

A Unified Kantorovich Duality for Multimarginal Optimal Transport

Yehya Cheryala, Mokhtar Z. Alaya, Salim Bouzebda

Preprint arXiv, 2026

Theme: Optimal Transport for Machine Learning

arXiv

Multimarginal optimal transport (MOT) has gained increasing attention in recent years, notably due to its relevance in machine learning and statistics, where one seeks to jointly compare and align multiple probability distributions. This paper presents a unified and complete Kantorovich duality theory for MOT problem on general Polish product spaces with bounded continuous cost function. For marginal compact spaces, the duality identity is derived through a convex-analytic reformulation, that identifies the dual problem as a Fenchel-Rockafellar conjugate. We obtain dual attainment and show that optimal potentials may always be chosen in the class of \(c\)-conjugate families, thereby extending classical two-marginal conjugacy principle into a genuinely multimarginal setting. In non-compact setting, where direct compactness arguments are unavailable, we recover duality via a truncation-tightness procedure based on weak compactness of multimarginal transference plans and boundedness of the cost. We prove that the dual value is preserved under restriction to compact subsets and that admissible dual families can be regularized into uniformly bounded \(c\)-conjugate potentials. The argument relies on a refined use of \(c\)-splitting sets and their equivalence with multimarginal \(c\)-cyclical monotonicity. We then obtain dual attainment and exact primal-dual equality for MOT on arbitrary Polish spaces, together with a canonical representation of optimal dual potentials by \(c\)-conjugacy. These results provide a structural foundation for further developments in probabilistic and statistical analysis of MOT, including stability, differentiability, and asymptotic theory under marginal perturbations.

Optimal Transport Guarantees to Nonparametric Regression for Locally Stationary Time Series

Jan N. Tinio, Mokhtar Z. Alaya, Salim Bouzebda

Conference AISTATS, 2026

Theme: Optimal Transport for Machine Learning

PMLR Code

Figure from the paper

Locally stationary time series (LSTS) represent an essential modeling paradigm for capturing the nuanced dynamics inherent in time series data, whose statistical characteristics, including mean and variance, evolve smoothly over time. In this paper, we propose a conditional probability distribution estimator for LSTS through Nadaraya–Watson (NW) kernel smoothing. NW estimator leverages local kernel smoothing to approximate the conditional distribution of a response variable given its covariates. Under mild conditions, we establish optimal transport convergence guarantees to the proposed NW-based conditional probability estimator. These guarantees are initially proven in the univariate setting using the Wasserstein distance, and subsequently in a multivariate setting employing the sliced Wasserstein distance. To corroborate our theoretical findings, we conduct a wide range of numerical experiments to assess the convergence rates and showcase the practical relevance of the estimator in capturing intricate temporal dependencies in complex nonstationary phenomena.

Unmixing Mean Embeddings for Domain Adaptation with Target Label Proportion

Alain Rakotomamonjy, Maxime Berar, Mokhtar Z. Alaya

Conference AISTATS, 2026

Theme: Optimal Transport for Machine Learning

HAL PDF

Figure from the paper

We introduce a novel approach to domain adaptation within the context of Learning from Label Proportions (LLP). We address the challenging scenario where labeled samples are available in the source domain, but only bags of unlabeled samples with their corresponding label proportions are accessible in the target domain. Our proposed method, bagMME (Bag Matching Mean Embeddings), tackles the distributional shift between domains by focusing on matching class-conditional distributions. A key contribution of bagMME is a simple yet effective unmixing strategy that leverages the target label proportions to estimate the target class-conditional mean embeddings. These estimated target means are then aligned with their corresponding source class-conditional means, thereby reducing the domain discrepancy. We theoretically demonstrate the soundness of our approach and its effectiveness in mitigating distributional shifts. Extensive experiments on various computer vision datasets showcase the superior performance of bagMME compared to state-of-the-art baselines. Our results highlight the critical role of incorporating target label proportions into the learning process for improved generalization on the target domain.

Application of Convolution Neural Network for Unfolding Simulated Neutron Spectra of an Activation Spectrometer

Rodayna Hmede, Thibaut Vinchon, Quentin Ducasse, Mokhtar Z. Alaya, Wilfried Monange, Jeremy Bez

Journal IEEE Transactions on Nuclear Science, 2026

Theme: Interdisciplinary Research (Neural Network for Unfolding Spectra)

IEEE Xplore

Figure from the paper

Neutron studies are of significant interest in fields such as radiation protection, nuclear reactor physics, and criticality safety, where accurate determination of neutron field is essential. Determining the neutron field typically involves unfolding detector signals, such as those obtained from the activation and counting neutron spectrometer (SNAC) detector. Traditional methods, as Bayesian approaches, have been widely integrated for this neutron spectrum unfolding. However, these methods rely on an initial solution estimation, introducing biases or uncertainties. Recent studies in artificial intelligence (AI) have demonstrated its potential to address challenges as hysteresis regression. This work is based on our novel convolutional neural network (CNN) architecture to overcome the hysteresis problem in neutron spectrum unfolding. The CNN model predicts the neutron spectrum directly from detector counts, eliminating the need for prior solution predictions. The proposed architecture was trained on a large simulation dataset and validated through a combination of Serpent simulations of various Californium-252 (Cf) spectra and Monte Carlo N-Particles (MCNPs) simulations of the Silene reactor. These two complementary simulation approaches are used to evaluate the CNN’s evaluation in realistic neutron environments. The results demonstrate the model’s high efficiency and accuracy, as evidenced by key performance metrics and the quality of the predicted spectrum (SQ). This approach represents a significant step forward in optimizing and validating AI-based methods for neutron field, especially in criticality dosimetry and radiation protection applications. Preliminary comparisons with Bayesian unfolding codes already indicate than CNN-based predictions can capture fine spectral features. A benchmark against MAXED, GRAVEL, and Nubay remains a key perspective, together with validation campaigns on neutron facilities.

2025

PatchTrAD: A Patch-Based Transformer focusing on Patch-Wise Reconstruction Error for Time Series Anomaly Detection

Samy-Melwan Vilhes, Gilles Gasso, Mokhtar Z. Alaya

Conference 33rd European Signal Processing Conference, EUSIPCO 2025

Theme: Deep Learning for Time Series

EUSIPCO Code

Figure from the paper

Time series anomaly detection (TSAD) focuses on identifying whether observations in streaming data deviate significantly from normal patterns. With the prevalence of connected devices, anomaly detection on time series has become paramount, as it enables real-time monitoring and early detection of irregular behaviors across various application domains. In this work, we introduce PatchTrAD, a Patch-based Transformer model for time series anomaly detection. Our approach leverages a Transformer encoder along with the use of patches under a reconstruction-based framework for anomaly detection. Empirical evaluations on multiple benchmark datasets show that PatchTrAD is on par, in terms of detection performance, with state-of-the-art deep learning models for anomaly detection while being time efficient during inference.

Adversarial Semi-Supervised Domain Adaptation for Semantic Segmentation: A New Role for Labeled Target Samples

Marwa Kechaou, Mokhtar Z. Alaya, Romain Hérault, Gilles Gasso

Journal Computer Vision and Image Understanding, 2025

Theme: Optimal Transport for Machine Learning

ScienceDirect

Figure from the paper

Adversarial learning baselines for domain adaptation (DA) approaches in the context of semantic segmentation are under explored in semi-supervised framework. These baselines involve solely the available labeled target samples in the supervision loss. In this work, we propose to enhance their usefulness on both semantic segmentation and the single domain classifier neural networks. We design new training objective losses for cases when labeled target data behave as source samples or as real target samples. The underlying rationale is that considering the set of labeled target samples as part of source domain helps reducing the domain discrepancy and, hence, improves the contribution of the adversarial loss. To support our approach, we consider a complementary method that mixes source and labeled target data, then applies the same adaptation process. We further propose an unsupervised selection procedure using entropy to optimize the choice of labeled target samples for adaptation. We illustrate our findings through extensive experiments on the benchmarks GTA5, SYNTHIA, and Cityscapes. The empirical evaluation highlights competitive performance of our proposed approach.

2024

Gaussian-Smoothed Sliced Probability Divergences

Mokhtar Z. Alaya, Alain Rakotomamonjy, Maxime Bérar, Gilles Gasso

Journal Transactions on Machine Learning Research, 2024

Theme: Optimal Transport for Machine Learning

OpenReview

Figure from the paper

Gaussian smoothed sliced Wasserstein distance has been recently introduced for comparing probability distributions, while preserving privacy on the data. It has been shown that it provides performances similar to its non-smoothed (non-private) counterpart. However, the computational and statistical properties of such a metric have not yet been well-established. This work investigates the theoretical properties of this distance as well as those of generalized versions denoted as Gaussian-smoothed sliced divergences \(\textrm{GSD}_p\). We first show that smoothing and slicing preserve the metric property and the weak topology. To study the sample complexity of such divergences, we then introduce \(\hat{\hat\mu}_{n}\) the double empirical distribution for the smoothed-projected \(\mu\). The distribution \(\hat{\hat\mu}_{n}\) is a result of a double sampling process: one from sampling according to the origin distribution \(\mu\) and the second according to the convolution of the projection of \(\mu\) on the unit sphere and the Gaussian smoothing. We particularly focus on the Gaussian smoothed sliced Wasserstein distance \(\textrm{GSW}_p\) and prove that it converges with a rate \(O(n^{-1/{2p}})\). We also derive other properties, including continuity, of different divergences with respect to the smoothing parameter. We support our theoretical findings with empirical studies in the context of privacy-preserving domain adaptation.

2023

Neutron Spectrum Unfolding using two Architectures of Convolutional Neural Networks

Maha Bouhadida, Asmae Mazzi, Mariya Brovchenko, Thibaut Vinchon, Mokhtar Z. Alaya, Wilfried Monange, François Trompier

Journal Nuclear Engineering and Technology, 2023

Theme: Interdisciplinary Research (Neural Network for Unfolding Spectra)

ScienceDirect

Figure from the paper

We deploy artificial neural networks to unfold neutron spectra from measured energy-integrated quantities. These neutron spectra represent an important parameter allowing to compute the absorbed dose and the kerma to serve radiation protection in addition to nuclear safety. The built architectures are inspired from convolutional neural networks. The first architecture is made up of residual transposed convolution's blocks while the second is a modified version of the U-net architecture. A large and balanced dataset is simulated following “realistic” physical constraints to train the architectures in an efficient way. Results show a high accuracy prediction of neutron spectra ranging from thermal up to fast spectrum. The dataset processing, the attention paid to performances' metrics and the hyper-optimization are behind the architectures' robustness.

2022

Binacox: Automatic Cut-Points Detection in High-Dimensional Cox Model, with Applications to Genetic Data

Simon Bussy, Mokhtar Z. Alaya, Anne-Sophie Jannot, Agathe Guilloux

Journal Biometrics, 2022

Theme: Structured Statistical Learning

PubMed Code

Figure from the paper

We introduce the binacox, a prognostic method to deal with the problem of detecting multiple cut-points per features in a multivariate setting where a large number of continuous features are available. The method is based on the Cox model and combines one-hot encoding with the binarsity penalty, which uses total-variation regularization together with an extra linear constraint, and enables feature selection. Nonasymptotic oracle inequalities for prediction and estimation with a fast rate of convergence are established. The statistical performance of the method is examined in an extensive Monte Carlo simulation study, and then illustrated on three publicly available genetic cancer datasets. On these high-dimensional datasets, our proposed method significantly outperforms state-of-the-art survival models regarding risk prediction in terms of the C-index, with a computing time orders of magnitude faster. In addition, it provides powerful interpretability from a clinical perspective by automatically pinpointing significant cut-points in relevant variables.

Theoretical Guarantees for Bridging Metric Measure Embedding and Optimal Transport

Mokhtar Z. Alaya, Maxime Bérar, Gilles Gasso, Alain Rakotomamonjy

Journal Neurocomputing, 2022

Theme: Optimal Transport for Machine Learning

ScienceDirect

Figure from the paper

We propose a novel approach for comparing distributions whose supports do not necessarily lie on the same metric space. Unlike Gromov-Wasserstein (GW) distance which compares pairwise distances of elements from each distribution, we consider a method allowing to embed the metric measure spaces in a common Euclidean space and compute an optimal transport (OT) on the embedded distributions. This leads to what we call a sub-embedding robust Wasserstein (SERW) distance. Under some conditions, SERW is a distance that considers an OT distance of the (low-distorted) embedded distributions using a common metric. In addition to this novel proposal that generalizes several recent OT works, our contributions stands on several theoretical analyses: (i) we characterize the embedding spaces to define SERW distance for distribution alignment; (ii) we prove that SERW mimics almost the same properties of GW distance, and we give a cost relation between GW and SERW. The paper also provides some numerical illustrations of how SERW behaves on matching problems.

2021

Optimal Transport for Conditional Domain Matching and Label Shift

Alain Rakotomamonjy, Rémi Flamary, Gilles Gasso, Mokhtar Z. Alaya, Maxime Berar, Nicolas Courty

Journal Machine Learning, 2021

Theme: Optimal Transport for Machine Learning

Springer Code

Figure from the paper

We address the problem of unsupervised domain adaptation under the setting of generalized target shift (joint class-conditional and label shifts). For this framework, we theoretically show that, for good generalization, it is necessary to learn a latent representation in which both marginals and class-conditional distributions are aligned across domains. For this sake, we propose a learning problem that minimizes importance weighted loss in the source domain and a Wasserstein distance between weighted marginals. For a proper weighting, we provide an estimator of target label proportion by blending mixture estimation and optimal matching by optimal transport. This estimation comes with theoretical guarantees of correctness under mild assumptions. Our experimental results show that our method performs better on average than competitors across a range domain adaptation problems including digits, VisDA and Office.

Heterogeneous Wasserstein Discrepancy for Incomparable Distributions

Mokhtar Z. Alaya, Maxime Bérar, Gilles Gasso, Alain Rakotomamonjy

Preprint arXiv, 2021

Theme: Optimal Transport for Machine Learning

arXiv

Figure from the paper

Optimal Transport (OT) metrics allow for defining discrepancies between two probability measures. Wasserstein distance is for longer the celebrated OT-distance frequently-used in the literature, which seeks probability distributions to be supported on the same metric space. Because of its high computational complexity, several approximate Wasserstein distances have been proposed based on entropy regularization or on slicing, and one-dimensional Wasserstein computation. In this paper, we propose a novel extension of Wasserstein distance to compare two incomparable distributions, that hinges on the idea of distributional slicing, embeddings, and on computing the closed-form Wasserstein distance between the sliced distributions. We provide a theoretical analysis of this new divergence, called heterogeneous Wasserstein discrepancy (HWD), and we show that it preserves several interesting properties including rotation-invariance. We show that the embeddings involved in HWD can be efficiently learned. Finally, we provide a large set of experiments illustrating the behavior of HWD as a divergence in the context of generative modeling and in query framework.

POT: Python Optimal Transport

Rémi Flamary, Nicolas Courty, Alexandre Gramfort, Mokhtar Z. Alaya, Aurélie Boisbunon, Stanislas Chambon, Laetitia Chapel, Adrien Corenflos, Kilian Fatras, Nemo Fournier, Léo Gautheron, Nathalie T.H. Gayraud, Hicham Janati, Alain Rakotomamonjy, Ievgen Redko, Antoine Rolet, Antony Schutz, Vivien Seguy, Danica J. Sutherland, Romain Tavenard, Alexander Tong, Titouan Vayer

Journal Journal of Machine Learning Research, 2021

Theme: Optimal Transport for Machine Learning

POT website JMLR Code

Figure from the paper

Optimal transport has recently been reintroduced to the machine learning community thanks in part to novel efficient optimization procedures allowing for medium to large scale applications. We propose a Python toolbox that implements several key optimal transport ideas for the machine learning community. The toolbox contains implementations of a number of founding works of OT for machine learning such as Sinkhorn algorithm and Wasserstein barycenters, but also provides generic solvers that can be used for conducting novel fundamental research. This toolbox, named POT for Python Optimal Transport, is open source with an MIT license.

2020

Partial Gromov-Wasserstein with Applications on Positive-Unlabeled Learning

Laetitia Chapel, Mokhtar Z. Alaya, Gilles Gasso

Conference NeurIPS, 2020

Theme: Optimal Transport for Machine Learning

NeurIPS Code

Figure from the paper

Classical optimal transport problem seeks a transportation map that preserves the total mass between two probability distributions, requiring their masses to be equal. This may be too restrictive in some applications such as color or shape matching, since the distributions may have arbitrary masses and/or only a fraction of the total mass has to be transported. In this paper, we address the partial Wasserstein and Gromov-Wasserstein problems and propose exact algorithms to solve them. We showcase the new formulation in a positive-unlabeled (PU) learning application. To the best of our knowledge, this is the first application of optimal transport in this context and we first highlight that partial Wasserstein-based metrics prove effective in usual PU learning settings. We then demonstrate that partial Gromov-Wasserstein metrics are efficient in scenarii in which the samples from the positive and the unlabeled datasets come from different domains or have different features.

Open Set Domain Adaptation using Optimal Transport

Marwa Kechaou, Romain Hérault, Mokhtar Z. Alaya, Gilles Gasso

Conference ECML-PKDD, 2020

Theme: Optimal Transport for Machine Learning

Springer

Figure from the paper

We present a 2-step optimal transport approach that performs a mapping from a source distribution to a target distribution. Here, the target has the particularity to present new classes not present in the source domain. The first step of the approach aims at rejecting the samples issued from these new classes using an optimal transport plan. The second step solves the target (class ratio) shift still as an optimal transport problem. We develop a dual approach to solve the optimization problem involved at each step and we prove that our results outperform recent state-of-the-art performances. We further apply the approach to the setting where the source and target distributions present both a label-shift and an increasing covariate (features) shift to show its robustness.

2019

Screening Sinkhorn Algorithm for Regularized Optimal Transport

Mokhtar Z. Alaya, Maxime Bérar, Gilles Gasso, Alain Rakotomamonjy

Conference NeurIPS, 2019

Theme: Optimal Transport for Machine Learning

NeurIPS Code

Figure from the paper

We introduce in this paper a novel strategy for efficiently approximating the Sinkhorn distance between two discrete measures. After identifying neglectable components of the dual solution of the regularized Sinkhorn problem, we propose to screen those components by directly setting them at that value before entering the Sinkhorn problem. This allows us to solve a smaller Sinkhorn problem while ensuring approximation with provable guarantees. More formally, the approach is based on a new formulation of dual of Sinkhorn divergence problem and on the KKT optimality conditions of this problem, which enable identification of dual components to be screened. This new analysis leads to the Screenkhorn algorithm. We illustrate the efficiency of Screenkhorn on complex tasks such as dimensionality reduction and domain adaptation involving regularized optimal transport.

Collective Matrix Completion

Mokhtar Z. Alaya, Olga Klopp

Journal Journal of Machine Learning Research, 2019

Theme: Structured Statistical Learning

JMLR Code

Figure from the paper

Matrix completion aims to reconstruct a data matrix based on observations of a small number of its entries. Usually in matrix completion a single matrix is considered, which can be, for example, a rating matrix in recommendation system. However, in practical situations, data is often obtained from multiple sources which results in a collection of matrices rather than a single one. In this work, we consider the problem of collective matrix completion with multiple and heterogeneous matrices, which can be count, binary, continuous, etc. We first investigate the setting where, for each source, the matrix entries are sampled from an exponential family distribution. Then, we relax the assumption of exponential family distribution for the noise and we investigate the distribution-free case. In this setting, we do not assume any specific model for the observations. The estimation procedures are based on minimizing the sum of a goodness-of-fit term and the nuclear norm penalization of the whole collective matrix. We prove that the proposed estimators achieve fast rates of convergence under the two considered settings and we corroborate our results with numerical experiments.

Binarsity: a Penalization for One-Hot Encoded Features in Linear Supervised Learning

Mokhtar Z. Alaya, Simon Bussy, Stéphane Gaïffas, Agathe Guilloux

Journal Journal of Machine Learning Research, 2019

Theme: Structured Statistical Learning

JMLR Code

Figure from the paper

This paper deals with the problem of large-scale linear supervised learning in settings where a large number of continuous features are available. We propose to combine the well-known trick of one-hot encoding of continuous features with a new penalization called binarsity. In each group of binary features coming from the one-hot encoding of a single raw continuous feature, this penalization uses total-variation regularization together with an extra linear constraint. This induces two interesting properties on the model weights of the one-hot encoded features: they are piecewise constant, and are eventually block sparse. Non-asymptotic oracle inequalities for generalized linear models are proposed. Moreover, under a sparse additive model assumption, we prove that our procedure matches the state-of-the-art in this setting. Numerical experiments illustrate the good performances of our approach on several datasets. It is also noteworthy that our method has a numerical complexity comparable to standard \(\ell_1\) penalization.

2017

High-Dimensional Time-Varying Aalen and Cox Models

Mokhtar Z. Alaya, Sarah Lemler, Agathe Guilloux, Thibault Allart

Preprint arXiv, 2017

Theme: Structured Statistical Learning

HAL

We consider the problem of estimating the intensity of a counting process in high-dimensional time-varying Aalen and Cox models. We introduce a covariate-specific weighted total-variation penalization, using data-driven weights that correctly scale the penalization along the observation interval. We provide theoretical guaranties for the convergence of our estimators and present a proximal algorithm to solve the convex studied problems. The practical use and effectiveness of the proposed method are demonstrated by simulation studies and real data example.

2015

Learning the Intensity of Time Events with Change-Points

Mokhtar Z. Alaya, Stéphane Gaïffas, Agathe Guilloux

Journal IEEE Transactions on Information Theory, 2015

Theme: Structured Statistical Learning

IEEE Xplore

Figure from the paper

We consider the problem of learning the inhomogeneous intensity of a counting process, under a sparse segmentation assumption. We introduce a weighted total-variation penalization, using data-driven weights that correctly scale the penalization along the observation interval. We prove that this leads to a sharp tuning of the convex relaxation of the segmentation prior, by stating oracle inequalities with fast rates of convergence, and consistency for change-points detection. This provides first theoretical guarantees for segmentation with a convex proxy beyond the standard i.i.d signal + white noise setting. We introduce a fast algorithm to solve this convex problem. Numerical experiments illustrate our approach on simulated and on a high-frequency genomics dataset.