Paper analyzes mathematical theory behind out-of-sample DR extensions.
problem Developing a solid mathematical foundation for out-of-sample DR extensions.
method Utilizes RKHS theory to treat DR extension as an extension of the identity on RKHS defined on X.
result Shows Nyström-type DR extension as an orthogonal projection and provides conditions for exact DR extension.
Proposes a parametric t-SNE without perplexity tuning.
problem Non-parametric t-SNE's perplexity parameter limits DR quality.
method Multi-scale parametric t-SNE with deep neural network.
result Produces reliable embeddings with competitive neighborhood preservation.
New models improve classification model performance, especially robust to small training sets.
problem Improving classification model performance, especially robust to small training sets.
method Distributionally robust AUC maximization models using Kantorovich metric and hinge loss function.
result The proposed DR-AUC models outperform standard models in general and worst-case out-of-sample performance.
This paper proposes an out-of-sample extension framework for a global manifold learning algorithm (Isomap) that uses temporal information in out-of-sample points in order to make the embedding more robust to noise and artifacts. Given a set of noise-free training data and its embedding, the proposed framework extends t…
The paper proves limit theorems for graph embeddings out-of-sample.
problem Proving limit theorems for graph embeddings out-of-sample.
method Least-squares and maximum-likelihood objectives for adjacency and Laplacian spectral embeddings.
result Out-of-sample extensions based on these objectives obey central limit theorems and concentration inequalities.
Two approaches extend graph embedding to unseen vertices.
problem Extending graph embedding to new data points.
method Least-squares and maximum-likelihood formulations.
result Both methods estimate latent positions with the same error rate under latent position models.
New model-free DR-RL algorithm with finite sample complexity.
problem Limited model-free DR-RL methods with convergence guarantees or sample complexities.
method Integrates Multi-level Monte Carlo (MLMC) technique with threshold mechanism.
result First model-free DR-RL approach with finite sample complexity for total variation and Chi-square divergence.
Dimensionality reduction methods are very common in the field of high dimensional data analysis. Typically, algorithms for dimensionality reduction are computationally expensive. Therefore, their applications for the analysis of massive amounts of data are impractical. For example, repeated computations due to accumula…
Regression aims at estimating the conditional mean of output given input. However, regression is not informative enough if the conditional density is multimodal, heteroscedastic, and asymmetric. In such a case, estimating the conditional density itself is preferable, but conditional density estimation (CDE) is challeng…
ProbDR framework interprets DR algorithms as probabilistic inference.
problem Efficiently compressing high-dimensional data into lower dimensions.
method ProbDR variational framework that treats DR as probabilistic inference.
result ProbDR unifies various DR algorithms and enables probabilistic reasoning.
The paper studies continuous submodular functions and their optimization.
problem Maximizing continuous submodular functions in poly. time.
method Characterization of continuous submodularity, operations preserving it, and algorithms for constrained maximization.
result Continuous submodularity is equivalent to a weak DR property, leading to continuous DR-submodular functions with the full DR property.
Let $(M, \dr M)$ be a 3-manifold with incompressible boundary that admits a convex co-compact hyperbolic metric. We consider the hyperbolic metrics on M such that $\dr M$ looks locally like a hyperideal polyhedron, and we characterize the possible dihedral angles. We find as special cases the results of Bao and Bonah…
This paper explores autoencoders for estimating intrinsic dimensionality.
problem Estimating the intrinsic dimensionality of random vectors.
method Use of autoencoders for dimension estimation, focusing on architectural choices and regularization techniques.
result Autoencoders can be adapted for intrinsic dimension estimation, addressing questions beyond classic DR/DE techniques.
Landmark Diffusion Maps reduce manifold learning complexity for high-volume data streams.
problem Complexity of out-of-sample extensions in manifold learning techniques.
method Landmark Diffusion Maps (L-dMaps) using pruned spanning trees or k-medoids to select landmark points.
result Up to 50-fold speedups in out-of-sample extension with less than 4% errors in manifold reconstruction.
Let (M,∂M) be a compact 3-manifold with boundary which admits a complete, convex co-compact hyperbolic metric. For each hyperbolic metric g on M such that $\dr M$ is smooth and strictly convex, the induced metric on $\dr M$ has curvature K>−1, and each such metric on $\dr M$ is obtained for a unique ch…
This paper establishes non-asymptotic learning bounds for the DR covariate shift adaptation.
problem Distribution shift between training and test domains in machine learning.
method Doubly-robust (DR) estimator combining density ratio estimation and pilot regression model.
result First non-asymptotic learning bounds for DR covariate shift adaptation.
We consider the problem of vertex classification for graphs constructed from the latent position model. It was shown previously that the approach of embedding the graphs into some Euclidean space followed by classification in that space can yields a universally consistent vertex classifier. However, a major technical d…
This paper presents a new method for dimensionality reduction and out-of-sample extension.
problem Dimensionality reduction and out-of-sample extension in high-dimensional data.
method Adaptive non-linear embedding using positive semi-definite kernel eigenvectors.
result The embedding method is more robust to outliers compared to spectral embedding.
A new DR method for HSI classification improves accuracy with limited samples.
problem Challenges in DR for HSI classification with limited training samples.
method Graph-based spatial and spectral regularized local scaling cut (SSRLSC).
result Improved classification accuracy compared to spectral-only methods.
Dr.VAE improves drug response prediction accuracy.
problem Improving accuracy of drug response prediction.
method Two deep generative models based on Variational Autoencoders.
result Dr.VAE outperforms benchmarks by 3-11% AUROC and 2-30% AUPR.
Enhances supervised visualization for unseen data using autoencoders and random forest.
problem Lack of generalization to unseen test sets in supervised dimensionality reduction.
method Combines autoencoder and random forest proximities for out-of-sample extension.
result 40% reduction in training time with 10% of training data, achieving consistent quality.
New method for probabilistic modeling of integer submodular functions.
problem Lack of probabilistic modeling for integer submodular functions.
method Proposed Generalized Multilinear Extension and block-coordinate ascent algorithm.
result Demonstrated effectiveness and viability on real-world datasets.
Several popular graph embedding techniques for representation learning and dimensionality reduction rely on performing computationally expensive eigendecompositions to derive a nonlinear transformation of the input data space. The resulting eigenvectors encode the embedding coordinates for the training samples only, an…
A new probabilistic model for CCA reduces data complexity without vectorization.
problem Reducing data complexity for two-dimensional canonical correlation analysis.
method A latent variable model for matrix-variate data with two variational inference approaches.
result The proposed methods outperform existing probabilistic and non-probabilistic CCA approaches.
FEALM learns features for better nonlinear DR of hidden patterns.
problem DR misses important patterns on distorted manifolds.
method FEALM generates optimized projections using an optimization algorithm and neighbor-shape dissimilarity.
result FEALM captures important patterns on hidden manifolds.
DR-MCTS improves decision quality and sample efficiency in complex environments.
problem Improving decision quality and sample efficiency in complex environments.
method Integrates Doubly Robust off-policy estimation into Monte Carlo Tree Search (MCTS).
result DR-MCTS achieves superior performance in Tic-Tac-Toe and VirtualHome tasks.
Non-linear manifold learning enables high-dimensional data analysis, but requires out-of-sample-extension methods to process new data points. In this paper, we propose a manifold learning algorithm based on deep learning to create an encoder, which maps a high-dimensional dataset and its low-dimensional embedding, and …
Wasserstein DR optimizes decisions under uncertain distributions.
problem Learning decisions from uncertain data with limited samples.
method Wasserstein distributionally robust optimization (DR) approach.
result Optimal decisions can be computed efficiently and have strong guarantees.
Constructs diffeological moduli stacks for Higgs and flat bundles on Kähler manifolds
problem Establishing an equivalence between diffeological substacks of Higgs and flat bundles
method Using diffeological moduli stacks
result Shows equivalence of categories between semistable Higgs bundles and flat bundles
New DR-IC estimator reduces bias and variance in OPE.
problem Estimating value of a target policy using logged data from a different policy.
method DR-IC estimator that combines parametric reward model and context-based switching rule.
result DR-IC estimator outperforms state-of-the-art OPE algorithms.
StableDR stabilizes doubly robust learning for biased recommendation data.
problem Data missing not at random in recommender systems.
method StableDR, a stabilized doubly robust learning approach.
result StableDR achieves bounded bias, variance, and generalization error.
Paper defines and evaluates DR complex for persistent homology.
problem Computing persistent homology of Euclidean point cloud data.
method Delaunay-Rips complex construction for speed and stability.
result DR produces stable persistence diagrams under point cloud perturbations.
New algorithms solve DR-submodular maximization with faster convergence.
problem Maximizing monotone DR-submodular functions under convex constraints.
method Introduced strongly DR-submodular functions and proposed SDRFW and PGA algorithms.
result SDRFW achieves optimal approximation ratio after fewer iterations.
DR technique helps deep models learn faster and better.
problem Vanishing gradients and local minima in deep model training.
method DR technique applies penalties on hidden units to improve learning.
result DR improves convergence and generalization in deep neural networks.
Paper tackles non-monotone DR-submodular maximization with approximation and regret guarantees.
problem Maximizing non-monotone DR-submodular functions over specific sets.
method Frank-Wolfe algorithm for general convex sets, Stochastic Gradient Ascent for down-closed convex sets.
result First approximation guarantees for both offline and online settings.
Online and stochastic learning has emerged as powerful tool in large scale optimization. In this work, we generalize the Douglas-Rachford splitting (DRs) method for minimizing composite functions to online and stochastic settings (to our best knowledge this is the first time DRs been generalized to sequential version).…
Dimensionality reduction (DR) is often used as a preprocessing step in classification, but usually one first fixes the DR mapping, possibly using label information, and then learns a classifier (a filter approach). Best performance would be obtained by optimizing the classification error jointly over DR mapping and cla…
Neural networks struggle with identity relations; DR units improve generalization.
problem Neural networks fail to generalize identity relations.
method Exploring various factors in neural network architecture and learning process, including number of hidden layers, activation function, and data representation.
result DR units improve generalization, leading to almost perfect test accuracy in mid fusion setting.
Paper tackles online DR-submodular maximization with various convex sets.
problem Maximizing DR-submodular functions online over different convex sets.
method Develops online algorithms with approximation guarantees for various convex sets.
result Achieves 1/e-approximation ratio with O(T2/3) regret for down-closed sets. New DR method improves robustness in high-dimensional treatment effects.
problem Estimating dynamic treatment effects with high-dimensional confounders.
method Proposes a novel DR representation for intermediate conditional outcome models.
result Achieves superior robustness guarantees with high-dimensional confounders.
Two algorithms maximize DR-submodular functions under convex constraints.
problem Maximizing non-monotone DR-submodular functions under convex constraints.
method Developed two algorithms with provable guarantees: a two-phase algorithm with 1/4 approximation and a Frank-Wolfe variant with 1/e approximation.
result Proved strong relation between stationary points and global optimum for DR-submodular functions.
DR-NMF uses unfolded ISTA for speech separation, offering interpretability and speed.
problem Speech separation in noisy environments.
method DR-NMF is a recurrent neural network that unfolds ISTA iterations for NMF of spectrograms.
result DR-NMF outperforms NMF and LSTM networks in speech separation.
Interactive DR framework for comparing datasets.
problem Limited flexibility in existing DR methods for comparative analysis.
method Unified linear comparative analysis (ULCA) with interactive optimization and visualization.
result ULCA and optimization algorithm improve comparative analysis efficiency and flexibility.
Dr.S recommends cancer drugs based on genomic data.
problem Personalizing cancer treatments using genomic information.
method Machine learning to identify optimal drug-gene associations.
result Developed a Drug Recommendation System (Dr.S) for cancer cell lines.
A new DR formulation improves metric learning for faster and more stable performance.
problem Learning embeddings for class separation in metric learning.
method Distance-ratio (DR) formulation for metric learning.
result DR formulation achieves improved or comparable generalization performances.
Paper evaluates competence measures for DRS systems.
problem Choosing the best measure to quantify competence in DRS systems is challenging.
method Reviewed and adapted eight competence measures for regression problems, compared them on 15 datasets, and evaluated three DRS systems.
result DRS systems outperform individual regressors and static systems, but competence measure choice depends on the problem.
Paper tackles online learning for DR management with incentives.
problem Estimating baseline consumption in DR programs with consumer incentives.
method Online learning scheme using least-squares with perturbed reward prices.
result Achieves low regret of $\mathcal{O}\left((\log{T})^2
ight)$ compared to optimal.
Paper analyzes an algorithm for maximizing non-concave functions with budget constraints.
problem Maximizing non-concave functions with budget constraints under DR-submodularity.
method Generalized Sequential algorithm for online monotone DR-submodular function maximization.
result First competitive ratio bound matches known tight bound for linear objective functions.