Portrait of Mokhtar Z. Alaya

Mokhtar Z. Alaya

Maître de Conférences

LMAC – UTC

Research Summary

My research focuses on the theoretical, algorithmic, and applied aspects of statistical machine learning, with interests spanning sparse inference, matrix completion, and survival analysis. A central theme of my work is optimal transport: I develop theory and algorithms for efficient Sinkhorn computations, sliced divergences, partial transport, and Gromov–Wasserstein problems, with applications to domain adaptation. Together with my collaborators, I investigate optimal transport approaches to nonparametric estimation for locally stationary time series. More recently, my research has expanded to deep learning for time series and interdisciplinary applications in mechanics, chemistry, and neutron physics.

Research Interests

  • Statistical Machine Learning Theory & Applications
  • Machine Learning (ML) and Deep Learning (DL)
  • Optimal Transport for ML/DL
  • High-dimensional Statistics
  • Matrix Completion
  • Data Science

Optimal Transport for Machine Learning

Optimal transport (OT) based data analysis has proven a significant usefulness to achieve many central tasks in machine learning, statistics, computer vision, among many others. This success is due to the natural geometric comparison framework offered by OT toolboxes. In a nutshell, OT is mathematical tool to compare distributions by computing a transportation mass plan from a source to a target distribution. Distances based on OT are referred to as the Wasserstein distance and have been successfully employed in a wide variety of machine learning.

Optimal matching of two clouds

Each point of μ travels in a straight line to its match in ν, under the matching that minimizes the total squared distance (W2 is in units of the panel width).
Optimal transport between two densities on the lineThe source density mu, with two modes, is on the left and the target density nu on top. The square shows the entropic transport plan computed with Sinkhorn's algorithm: its mass lies along an increasing curve, the exact optimal map T, drawn dashed. In the animation, grains of mass leave mu, move across to the map T, then up to nu, where they arrive distributed as nu.μνTπεsourcetargetentropic plan (Sinkhorn)exact map
Transporting μ onto ν. The entropic plan computed by Sinkhorn’s algorithm gathers its mass along the exact optimal map T (dashed), which on the line is the monotone rearrangement T = Fν−1 ∘ Fμ: grains of equal mass leave μ at its quantiles, turn on T and land on the quantiles of ν.

Structured Statistical Learning

Availability of massive data in high-dimension, namely when the number of features (covariates) is much larger than the number of observations, arises in diverse fields of sciences, ranging from computational biology and health studies to financial engineering and risk management, to name a few. These data have presented serious challenges to existing learning methods and reshaped statistical thinking and data analysis. To address the curse of dimensionality problem, sparse inference is now an ubiquitous technique for dimension reduction and variable selection. Sparse solution generally helps in better interpretation of the model and more importantly leads to better generalization on unseen data. A fundamental step in sparsity is to do careful variable selection based on the idea of adding a penalty term on the model complexity to some goodness-of-fit.

Total-variation denoising

A total-variation penalty λ turns noisy observations into a piecewise-constant estimate, with fewer change-points as λ grows.
Why the l1 penalty gives exact zerosElliptical contours of the squared loss around the least-squares estimate, with the l1 ball (a diamond) and the l2 ball (a dashed circle) of the same radius. The smallest contour that meets the l1 ball touches it at the corner where the first coordinate is zero; the l2 ball is met at a point where neither coordinate is zero. In the animation, a contour grows from the estimate until it meets the l2 ball, then the l1 ball.ββ1 = 0ℓ1ℓ2β1β2(a) penalty balls and loss contours
Lasso regularization pathCoefficients of the Lasso as the penalty lambda decreases, on simulated data with 60 observations and 20 variables, of which 3 are in the model. All coefficients are exactly zero for a large penalty; the three variables of the model enter first and grow largest. In the animation, the path is drawn as lambda decreases.a variable outside the model (true coefficient 0)a variable outside the model (true coefficient 0)a variable outside the model (true coefficient 0)a variable outside the model (true coefficient 0)a variable outside the model (true coefficient 0)a variable outside the model (true coefficient 0)a variable outside the model (true coefficient 0)a variable outside the model (true coefficient 0)a variable outside the model (true coefficient 0)a variable outside the model (true coefficient 0)a variable outside the model (true coefficient 0)a variable outside the model (true coefficient 0)a variable outside the model (true coefficient 0)a variable outside the model (true coefficient 0)a variable outside the model (true coefficient 0)a variable outside the model (true coefficient 0)a variable outside the model (true coefficient 0)a variable in the model (true coefficient 2)a variable in the model (true coefficient -1.5)a variable in the model (true coefficient 1.2)coefficientsλ decreasing →3 variables in the model17 others(b) Lasso path, simulated data
The geometry of sparsity. (a) Growing around the least-squares estimate (white dot), the contours of a quadratic loss first touch the ℓ2 ball (dashed circle) at a point where no coordinate vanishes, then the ℓ1 ball (diamond) at a corner, where β1 = 0 exactly. (b) The Lasso path on simulated data (n = 60, p = 20, 3 variables in the model): every coefficient is zero for a large penalty λ, and as λ decreases, the three variables of the model enter first.

Deep Learning for Time Series

Deep learning models for time series: anomaly detection with a patch-based transformer that scores each patch by its reconstruction error (PatchTrAD), and the normalization of large causal time-series models trained on heterogeneous collections of signals.

Anomaly detection with an autoencoder

An autoencoder learns the pattern of a time series from the stretch left of the dashed line, then flags as an anomaly what it fails to reconstruct (its error is the curve at the bottom).

Papers

Interdisciplinary Project: Artificial Intelligence for Mechanics (AI4Meca)

  • 2024–2025

    Unsupervised Deep Clustering of combined data from multi-Structural Health Monitoring Techniques obtained on Smart Polymer-Matrix Composites embedded with Piezoelectric Transducers

    Data fusion and clustering for structural health monitoringSchematic: a smart polymer-matrix composite specimen with two PZT transducers and a PVDF film is tested in load-unload tension. Its three signals are merged and passed to a convolutional autoencoder, whose latent codes are clustered. Digital image correlation (DIC) and acoustic emission (AE) serve as external validation. In the animation, the signals run into the fusion and through the autoencoder, and the latent codes gather into their clusters.smart compositePZTPVDFPZTload–unload teststhree signalsdatafusionzconvolutionalautoencoderclusters in thelatent spaceDICAEexternalvalidation
    Health monitoring of a smart composite. The signals of two PZT transducers and a PVDF film, recorded in load–unload tensile tests, are fused and clustered in the latent space of a convolutional autoencoder; digital image correlation (DIC) and acoustic emission (AE) serve as external validation. Schematic of the project, after Dolbachian et al. (2026).

    This project is a first collaboration with Matériaux et Surfaces team of Roberval Laboratory in UTC. It concerns data fusion and clustering methods utilizing deep neural networks (DNN) to classify heterogeneous data from different acquisition methods. A machine learning algorithm, specifically a convolutional autoencoder, was evaluated by clustering datasets obtained from load-unload tensile tests of smart specimens embedding PZTs and PVDF transducers. These piezoelectric transducers were employed to collect multi-source data for SHM purposes. Additionally, external equipment such as DIC and AE were used for both validation and the initial testing of the DNN configuration. The project highlights the feasibility of using DNN architecture to classify multi-acquired and merged data for SHM.

    Paper Composites Multi-source Sensor Data Fusion Framework for Structural Health Monitoring of Polymer-Matrix Composites (PMC) Based on Latent-Space Clustering Using a Convolutional AutoencoderMechanics of Advanced Materials and Structures, 2026

  • 2025–2028

    Generative Deep Learning for Atomistically Engineered Materials: Synergistic Integration of Molecular Dynamics Simulations, Experiments and Data Augmentation

    Generative models to augment materials dataSchematic: data from molecular dynamics simulations and from experiments are few. Generative models (GANs, VAEs and hybrids) produce new samples with the same structure; the augmented data serve to predict material behavior across scales, and the physical relevance of the generated samples is checked against experiments and simulations. In the animation, the data run through the loop and the generated samples appear one by one.moleculardynamicsexperimentslimited datagenerative modelsGAN · VAEand hybridsaugmented datameasuredor simulatedgeneratedpredictionsacross scalesvalidation ofphysical relevance
    Augmenting scarce materials data. Molecular dynamics simulations and experiments give few data: generative models (GANs, VAEs and hybrids) learn their structure and produce new samples, and the augmented data support predictions of material behavior across scales. The physical relevance of the generated samples is to be checked against experiments and atomistic simulations. Schematic of the project.

    This project is a second collaboration with Matériaux et Surfaces team of Roberval Laboratory in UTC. The primary objective of the project is to advance the development of nanostructured materials with tailored properties through novel approaches. It proposes an integrated, data-driven approach to expedite the development of advanced nanostructured materials. Using machine learning-driven data augmentation—specifically GANs, VAEs, and hybrid architectures—we address the constraints of limited datasets in materials science. This strategy complements existing experimental and atomistic modeling efforts, allowing robust predictions of material behavior across scales. It reduces time and cost associated with iterative experimentation and simulation. Moving forward, deeper validation of the synthetic data’s physical relevance—via experiments and atomistic simulations—will be crucial.

Interdisciplinary Project: Artificial Intelligence for Chemistry (AI4Chem)

  • 2025–2026

    Machine Learning Prediction Modelling for Chemistry with emphasis on High-Through Experiment

    Machine learning guiding high-throughput experimentsIllustration on simulated values: on a 96-well plate, 33 wells have been measured. A model fitted to them (here a Gaussian process) predicts the response of every well and proposes three unmeasured wells for the next runs: one where the predicted response is highest, two in the corner where the model is least certain. Running them and refitting closes the loop of exploration and optimisation. The animation runs four rounds of this loop.measured wellsML modelfitted to thempredicted responseround 1 of 4round 2 of 4round 3 of 4round 4 of 4high responsenot yet runproposed next runsrun, refit:explore, optimizemost uncertain
    Choosing the next experiments. A Gaussian process fitted to the measured wells predicts the response of every well and proposes three for the next runs: in the first round, the well with the highest predicted response and two wells where the model is least certain. Running them and refitting closes the loop between exploration and optimization, here over four rounds. Illustration on simulated values.

    This project is a collaboration with the team Activités Microbiennes et Bioprocédés (MAB) of TIMR Laboratory U High-throughput experimentation in chemistry enables rapid and automated exploration of chemical space, facilitating the discovery of new drugs. Integrating machine learning techniques with these high-throughput methods can further accelerate and enhance the exploration and optimization of chemical space.

Interdisciplinary Project: Neural Network for Unfolding Spectra

  • 2023–2026

    Neutron Spectrum Unfolding with Convolutional Neural Networks

    Neural networks for unfolding neutron spectraSchematic: seven reaction rates measured by an activation spectrometer are given to a convolutional neural network (transposed convolutions, or a U-net) trained on simulated spectra, which predicts the neutron spectrum over 1024 energy bins, from thermal to fast neutrons. The spectrum drawn is an illustrative shape. In the animation, the counts build up, the layers of the network light up in turn and the spectrum is drawn from thermal to fast energies.7 reaction ratesdetector countsCNNtransposed convolutionsor U-nettraining datasimulatedspectraneutron spectrum1024 energy binsthermalenergy, log scalefast
    Unfolding neutron spectra. A convolutional neural network trained on simulated spectra predicts the neutron spectrum, from thermal to fast energies, directly from the reaction rates measured by an activation spectrometer. Schematic of the project, after Bouhadida et al. (2023) and Hmede et al. (2026); the spectrum drawn is only an illustration.

    Unfolding a neutron spectrum means recovering the energy distribution of a neutron field from a few energy-integrated detector measurements, such as the reaction rates of an activation spectrometer; it is needed in radiation protection, nuclear reactor physics and criticality safety. Bayesian unfolding methods start from an initial estimate of the solution, which can bias the result. Convolutional neural networks trained on large sets of simulated spectra predict the spectrum directly from the measurements. Two architectures, one built from residual transposed convolution blocks and one a modified U-net, recover spectra from thermal to fast energies with high accuracy; a newer architecture was validated on Serpent simulations of californium-252 spectra and on MCNP simulations of the Silene reactor.