Captures data influence changes during training.
problem Traditional influence functions fail for modern training methods.
method Formalized trajectory-specific LOO influence, using data value embedding.
result Data influence varies by training stage, early and late stages have greater impact.
Model inference, such as model comparison, model checking, and model selection, is an important part of model development. Leave-one-out cross-validation (LOO) is a general approach for assessing the generalizability of a model, but unfortunately, LOO does not scale well to large datasets. We propose a combination of u…
A new method improves robustness and efficiency of Bayesian LOO-CV.
problem Computational expense and unreliability of classical LOO-CV in high-dimensional Bayesian models.
method Proposes a mixture estimator to compute Bayesian LOO-CV criteria with finite asymptotic variance.
result Improved robustness and efficiency in high-dimensional problems.
Improved LOO cross-validation for function approximation.
problem Estimating the Integrated Squared Error (ISE) for function approximation.
method Weighted Leave-One-Out cross-validation based on Gaussian Process.
result Significantly more precise ISE estimation compared to unweighted LOO.
The future predictive performance of a Bayesian model can be estimated using Bayesian cross-validation. In this article, we consider Gaussian latent variable models where the integration over the latent values is approximated using the Laplace method or expectation propagation (EP). We study the properties of several B…
In this paper we consider the problem of Gaussian process classifier (GPC) model selection with different Leave-One-Out (LOO) Cross Validation (CV) based optimization criteria and provide a practical algorithm using LOO predictive distributions with such criteria to select hyperparameters. Apart from the standard avera…
New method improves transductive learning predictions with multiplicative oracle inequalities.
problem Improving transductive learning predictions with known covariates.
method Median of Level-Set Aggregation (MLSA) for transductive LOO prediction.
result Proved multiplicative oracle inequality for LOO error.
LOO prediction method improves generalization guarantees for arbitrary datasets.
problem Understanding LOO error guarantees in fully transductive settings for arbitrary datasets.
method Median of Level-Set Aggregation (MLSA) for empirical-risk level sets.
result Multiplicative oracle inequality for LOO error with complexity scaling.
MoNODEs improve neural ODEs by separating dynamic states from static factors.
problem Learning non-linear dynamics with variations across trajectories.
method Introduces time-invariant modulator variables to separate dynamic states from static factors.
result Consistently improves model generalization and far-horizon forecasting.
A method to improve surrogate model accuracy using multiple fidelity models.
problem Efficiently combining models of varying accuracy and computational cost.
method Multifidelity Gaussian process models and leave-one-out cross-validation.
result Reduced LOO-CV error at the highest fidelity through adaptive learning.
Study shows minimizing the norm of the ERM solution stabilizes kernel ridge-less regression.
problem Stability of kernel ridge-less regression.
method Minimizing the norm of the ERM solution to minimize CV stability.
result Interpolating solution with minimum norm minimizes CV stability.
A reflexion space is generalization of a symmetric space introduced by O. Loos. We generalize locally symmetric spaces to local reflexion spaces in the similar way. We investigate, when local reflexion spaces are equivalently given by a locally flat Cartan connection of certain type.
New method optimizes hyperparameters for non-smooth problems efficiently.
problem Efficiently tuning hyperparameters for non-smooth cost functions.
method Combines hyperparameter search with proximal gradient updates.
result Method converges to local optimum of LOO validation error.
LOO-StabCP speeds up CP for multiple predictions.
problem Balancing computational efficiency and prediction accuracy in CP.
method Leave-One-Out Stable Conformal Prediction (LOO-StabCP) using algorithmic stability.
result LOO-StabCP is faster and more accurate than RO-StabCP.
New algorithms optimize actions under time-varying constraints without projecting.
problem Optimizing actions under time-varying constraints without projecting.
method Projection-free algorithms using linear optimization oracle.
result Guaranteed ildeO(T3/4) regret and O(T7/8) constraints violation. A new KDE model prevents singular solutions and accelerates optimization for probabilistic modeling.
problem Adapting to varying densities in data regions for probabilistic modeling.
method Adaptive KDE model with individual bandwidths, LOO-MLL criterion, and modified EM algorithm.
result The proposed models prevent singular solutions and have promising performance.
We classify homotopes of classical symmetric spaces (studied in Part I of this work). Our classification uses the fibered structure of homotopes: they are fibered as symmetric spaces, with flat fibers, over a non-degenerate base; the base spaces correspond to inner ideals in Jordan pairs. Using that inner ideals in cla…
A method for non-parametric conditional distribution estimation using CRPS-optimal binning.
problem Non-parametric conditional distribution estimation.
method Partitioning covariate-sorted observations into bins to minimize LOO-CRPS, selecting K by K-fold cross-validation of test CRPS.
result Produces narrower prediction intervals with near-nominal coverage compared to split-conformal competitors.
New proof for symmetric spaces with rectangular lattices.
problem Characterizing symmetric spaces with rectangular unit lattices.
method Explicit construction of isometric embeddings and analysis of root systems.
result Symmetric spaces with rectangular unit lattices are symmetric R-spaces.
Statistical machine learning models should be evaluated and validated before putting to work. Conventional k-fold Monte Carlo Cross-Validation (MCCV) procedure uses a pseudo-random sequence to partition instances into k subsets, which usually causes subsampling bias, inflates generalization errors and jeopardizes the r…
Recursive partitioning approaches producing tree-like models are a long standing staple of predictive modeling, in the last decade mostly as ``sub-learners'' within state of the art ensemble methods like Boosting and Random Forest. However, a fundamental flaw in the partitioning (or splitting) rule of commonly used tre…
Algorithm learns mixtures of Markov chains and MDPs from short trajectories.
problem Learning mixtures of Markov chains and MDPs from short unlabeled trajectories.
method Subspace estimation, spectral clustering, EM algorithm, model estimation, classification.
result 96.6% average accuracy on a mixture of two MDPs in gridworld, outperforming EM algorithm with random initialization.
The paper proves LOO CV is reliable under estimator stability.
problem Ensuring the reliability of leave-one-out cross validation.
method Using concentration inequalities based on logarithmic Sobolev inequality.
result LOO CV is a valid procedure under estimator stability.
Hua domain, named after Chinese mathematician Loo-Keng Hua, is defined as a domain in Cn fibered over an irreducible bounded symmetric domain Ω⊂Cd(d<n) with the fiber over z∈Ω being a (n−d)-dimensional generalized complex ellipsoid Σ(z). In general, a Hua domain is a nonhom…
A method merges two pretrained diffusion experts to improve image quality and likelihood.
problem Trade-off between image quality and data likelihood in diffusion models.
method Combining two pretrained diffusion experts by switching between them along the denoising trajectory.
result The merged model consistently matches or outperforms its base components, improving or preserving both likelihood and sample quality.
Combines causal learning with dynamical systems for practical model identification.
problem Lack of practical, identifiable models for causal inference in dynamical systems.
method Draws connection between causal representation learning and dynamical systems, applying identifiable methods to scalable differentiable solvers.
result Learned explicitly controllable models for trajectory-specific parameters.
This is the pdf -version of the author's Ph.D. thesis (1995, ULB, Belgium). The notion of symeplectic symmertic space is introduced and studied via Lie theoretical and symplectic geoemetrical methods. The first chapter concerns basic poperties, however, an explicit formula for the Loos connection in the symplectic fram…
Deep neural networks obtain state-of-the-art performance on a series of tasks. However, they are easily fooled by adding a small adversarial perturbation to input. The perturbation is often human imperceptible on image data. We observe a significant difference in feature attributions of adversarially crafted examples f…
A Banach symmetric space in the sense of O. Loos is a smooth Banach manifold M endowed with a multiplication map μ:M×M→M such that each left multiplication map μx:=μ(x,⋅) (with x∈M) is an involutive automorphism of (M,μ) with the isolated fixed point x. We show that morphisms of …
StrTransformer recovers sources without labels by optimizing latent matrices and enforcing structural constraints.
problem Unsupervised blind source recovery in signal processing.
method Source-wise structured Transformer framework with latent source matrix optimization, structural regularization, and branch-specific weights.
result StrTransformer learns distinct temporal-scale structures and recovers source-aligned latent trajectories.
Dynamic skewness models improve financial time series analysis.
problem Modeling financial time series with skewness and heavy tails.
method Dynamic skewness stochastic volatility models with penalized priors and HMC estimation.
result Penalized priors outperform classical choices in model performance.
Improves visualization of high-dimensional data by correcting misleading artifacts in neighbor embedding methods.
problem Misleading visual artifacts in t-SNE and UMAP due to lack of data-independent manifold learning interpretations.
method LOO-map framework that extends embedding maps to the entire input space, identifying and correcting map discontinuities.
result Developed point-wise diagnostic scores to detect unreliable embedding points and improve hyperparameter selection.
Backward Conformal Prediction offers flexible control over prediction set sizes while ensuring coverage guarantees.
problem Providing reliable prediction sets with controlled sizes in applications like medical diagnosis.
method Defines a rule that constrains prediction set sizes based on observed data, adapting coverage levels.
result Maintains computable coverage guarantees while ensuring interpretable, well-controlled prediction set sizes.