Study SGD dynamics in sequence models, revealing training phases and influence of sequence length.
problem Understanding SGD in sequence models like attention networks.
method Derived closed-form population loss and analyzed SGD dynamics for SSI models.
result Two distinct training phases: escape from uninformative initialization and alignment with target subspace.
Paper introduces Prob-SSI for robust OMA in noisy data.
problem Challenges in estimating modal parameters from noisy data.
method Probabilistic formulation of SSI, robust Prob-SSI algorithm.
result Robust Prob-SSI outperforms conventional SSI in corrupted data.
Paper tackles continual learning with single-index models, proving regret bounds.
problem Continual learning with single-index models across multiple tasks.
method Proposes a randomized strategy to learn a common single-index and task-specific link functions.
result Proves regret bounds for the proposed strategy under various loss function assumptions.
Bayesian SSI improves modal parameter uncertainty in operational systems.
problem Uncertainty in modal parameters due to stochastic operational systems and lack of forcing information.
method Proposes a Bayesian stochastic subspace identification (SSI) algorithm with a hierarchical probabilistic model and two inference schemes (Markov Chain Monte Carlo and variational Bayes).
result Posterior distributions over modal properties are obtained, showing lower variance for mean values coinciding with natural frequencies.
Reducing the incidence of surgical site infections (SSIs) is one of the objectives of the French nosocomial infection control program. Manual monitoring of SSIs is carried out each year by the hospital hygiene team and surgeons at the University Hospital of Bordeaux. Our goal was to develop an automatic detection algor…
This study compares Markowitz and Single-Index models for Malaysian stocks.
problem Optimizing portfolio selection for Malaysian stocks using different models.
method Applied Markowitz and Single-Index models to 10-year historical data of 10 stocks and a risk-free asset.
result Comparison of minimum variance and maximum Sharpe portfolios for both models under various constraints.
SGD shows distinct phases in learning single-index models, achieving optimal sample complexity and regret.
problem Learning single-index models with SGD in adaptive data settings.
method Stochastic gradient descent (SGD) with an optimal learning rate schedule.
result SGD achieves near-optimal sample complexity and regret guarantees across both burn-in and learning phases.
New neural networks learn single-index models efficiently.
problem Learning low-dimensional structure in high-dimensional data.
method Shallow neural networks with frozen biases, studied via gradient flow.
result Generalization guarantees match near-optimal sample complexity.
Gradient descent dynamics studied for DEQs in linear and single-index models.
problem Understanding gradient descent dynamics for DEQs.
method Rigorously studied gradient descent dynamics for DEQs in linear and single-index models.
result Gradient descent converges to a global minimizer for linear DEQs and single-index models.
New method approximates M-estimator and predictions without solving fixed-point equations.
problem Characterize behavior of M-estimator and predictions in single index models.
method Develops data-driven observable adjustments to proximal operators.
result Empirical distributions of M-estimator and predictions are approximated without solving fixed-point equations.
Mamba efficiently learns low-dimensional targets in-context via feature extraction.
problem Learning low-dimensional targets in context for computational efficiency.
method Test-time feature learning of a single-index model using Mamba's pretrained linear-time sequence model.
result Mamba achieves efficient in-context learning of low-dimensional targets via feature extraction.
New method uses spherical harmonics to simplify learning single-index models.
problem Learning single-index models with unknown one-dimensional projections.
method Proposes using spherical harmonics instead of Hermite polynomials to capture rotational symmetry.
result Characterizes the complexity of learning single-index models under arbitrary spherically symmetric input distributions.
New algorithm reduces regret for single-index bandits to nearly optimal.
problem Optimizing rewards from unknown projections of high-dimensional contexts.
method Two-phase algorithm: estimate projection direction, reduce to 1D bandit, use UCB.
result Achieved optimal regret of ildeO(T2/3). Improved regret for single-index bandits with optimal algorithm.
problem Optimizing rewards from unknown one-dimensional projections of high-dimensional contexts.
method Two-phase algorithm: first estimate projection direction, then discretize and use UCB.
result Achieved optimal regret of ildeO(T2/3). Efficiently learns Single-Index Models with constant factor approximation.
problem Learning Single-Index Models under L22 loss with unknown link functions. method An efficient algorithm using alignment sharpness for optimization.
result Achieves constant factor approximation to optimal loss for various distributions and link functions.
We investigate the variety of a portfolio of stocks in normal and extreme days of market activity. We show that the variety carries information about the market activity which is not present in the single-index model and we observe that the variety time evolution is not time reversal around the crash days. We obtain th…
Kernelized bandit algorithm tackles adaptive contextual bandits with single-index models.
problem Adaptive contextual bandits with single-index models and unknown link functions.
method Kernelized ε-greedy algorithm combining Stein-based index estimation and kernel ridge regression for reward functions.
result Unified framework for simultaneous learning and inference in single-index contextual bandits.
New method learns SIMs with arbitrary monotone activations without strong distributional assumptions.
problem Learning Single-Index Models with arbitrary monotone activations.
method Based on omniprediction with calibrated multiaccuracy and Bregman divergences.
result First agnostic learning result for SIMs with arbitrary monotone activations.
Improved SGD learning for single index models reduces sample complexity.
problem Learning a single index model with optimal sample complexity.
method Using smoothed loss in online SGD to reduce sample complexity.
result Online SGD with smoothed loss achieves optimal sample complexity of dk⋆/2. Study shows computational and statistical gaps in Gaussian Single-Index Models.
problem Statistical and computational trade-offs in high-dimensional regression problems.
method Analysis of SQ and LDP frameworks, partial-trace algorithm.
result Computational algorithms require significantly more samples than information-theoretic limits.
Neural networks can achieve optimal sample complexity for learning single-index models.
problem Achieving optimal computational-statistical tradeoff in learning Gaussian single-index models.
method Unified gradient-based algorithm for training a two-layer neural network, adaptable to various loss and activation functions.
result Sample complexity of ds⋆/2∨d matches the SQ lower bound up to a polylogarithmic factor. Proposes a new semi-parametric framework for batched bandits with covariates.
problem Sequential decision-making with batched feedback and contextual information.
method Batched single-Index Dynamic binning and Successive arm elimination (BIDS) using single-index regression.
result Achieves minimax-optimal rates for nonparametric batched bandits.
Paper uses deep Ritz method for solving stationary Schrödinger equation, proving convergence and feature emergence.
problem Solving stationary Schrödinger equation with high-dimensional features.
method Deep Ritz method, gradient descent, single-index model, two-neuron model.
result Gradient descent converges to near-optimal solution, feature emergence observed in two-neuron model.
Develops a robust model for skewed and heavy-tailed data in periodontal studies.
problem Skewed and heavy-tailed data in periodontal pocket depth measurements.
method Flexible two-piece scale Student-t error distribution and deep neural network with monotonicity constraints.
result Robust mode-based estimation resistant to outliers with clinical interpretability.
The study analyzes how neural reward models learn features for policy optimization in a Gaussian single-index model.
problem Reward modeling in policy optimization and its impact on downstream value.
method Two-stage neural reward model: first learns hidden direction, then fits readout layer.
result For any feature-learning temperature above a dimension-free threshold, a constant fraction of neurons recover the hidden direction.
The paper studies implicit regularization in over-parameterized models for high-dimensional data.
problem Understanding implicit regularization in over-parameterized models for high-dimensional data.
method The paper designs regularization-free algorithms for the high-dimensional single index model and provides theoretical guarantees for the induced implicit regularization phenomenon.
result The proposed methods achieve minimax optimal statistical rates of convergence and outperform classical methods with explicit regularization.
Deep learning maps tongue movements to speech sounds for voiceless individuals.
problem Developing silent speech interfaces for individuals without a larynx.
method Hybrid spatio-temporal 3D convolutions and feature shuffling for formant estimation and tracking from ultrasound tongue images.
result Best model achieves R-squared of 99.96% for vowel formant regression.
Deep single-index Fréchet regression for metric space-valued outputs
problem Predicting outputs in non-Euclidean spaces
method DeSI (Deep Single-Index Fréchet Regression)
result Interpretable index direction for inputs
Randomly biased data makes complex models as easy to learn as simple ones.
problem Learning complex models like multi-index and sparse Boolean functions.
method Introducing a small random shift in the first moment of the data distribution.
result Randomly biased data makes Gaussian single index models and sparse Boolean functions as easy to learn as linear functions.
Full-batch GD outperforms one-pass SGD in learning a single-index model with quadratic activation.
problem Learning a single-index model with quadratic activation using gradient descent.
method Full-batch gradient descent compared to one-pass stochastic gradient descent (SGD) on a correlation loss.
result Full-batch GD requires only n≃d samples for strong recovery, while one-pass SGD requires n≳dlogd samples. Proposes a method for valid inference in GPLSIMs with longitudinal data.
problem Challenges in longitudinal data inference due to within-subject correlation and unstable variance estimation.
method Profile estimating-equation approach using spline approximation and block empirical likelihood.
result Block empirical likelihood ratio statistic with Wilks-type chi-square limit for joint inference.
This paper proposes a Lasso-type estimator for a high-dimensional sparse parameter identified by a single index conditional moment restriction (CMR). In addition to this parameter, the moment function can also depend on a nuisance function, such as the propensity score or the conditional choice probability, which we es…
New method uses simple sensor intentions to learn complex tasks.
problem Defining reward schemes for exploration in robotic systems.
method Introduce simple sensor intentions (SSIs) to define auxiliary tasks.
result Learning system can solve complex robotic tasks using only raw sensor streams.
New method optimizes policies without assuming known link functions between preferences and rewards.
problem Policy alignment with unknown and unrestricted link functions.
method Formulates an f-divergence-constrained reward maximization problem, learning policies directly. result Induces a semiparametric single-index binary choice model for policy alignment.
This work uses diffusion models for accurate signal recovery from semi-parametric models.
problem Recovering signals from semi-parametric single index models with discontinuous link functions.
method Proposes an efficient reconstruction method using diffusion models that requires one round of sampling and inversion.
result Demonstrates more accurate reconstructions with fewer evaluations compared to competing methods.
Vanilla SGD learns SIM from anisotropic data without explicit covariance estimation.
problem Learning SIM from anisotropic Gaussian inputs.
method Vanilla Stochastic Gradient Descent (SGD) trained on SIM with anisotropic input.
result Vanilla SGD adapts to anisotropic data's covariance structure.
This work analyzes a two-stage algorithm for single index models, showing precise asymptotics of gradient descent.
problem Learning single index models with non-convex optimization.
method Spectral initialization followed by gradient descent, with detailed analysis of dynamics and asymptotics.
result Gradient descent converges to long-time fixed points in the large system limit, representing mean field behavior.
STNN-DDI predicts drug interactions using substructure-aware neural networks.
problem Predicting drug-drug interactions (DDIs) to avoid side effects in poly-drug treatments.
method Designing a novel Substructure-ware Tensor Neural Network (STNN-DDI) that learns a 3-D tensor of substructure-substructure interactions.
result Significant improvement in AUC, AUPR, Accuracy, and Precision compared to state-of-the-art models.
Single Index Models (SIMs) are simple yet flexible semi-parametric models for classification and regression. Response variables are modeled as a nonlinear, monotonic function of a linear combination of features. Estimation in this context requires learning both the feature weights, and the nonlinear function. While met…
Proposes a transfer learning framework for sparse SIMs without raw source data.
problem Lack of direct access to raw source data and known link functions in transfer learning.
method Source-data-free framework based on SIM, using summary statistics and a multilayer perceptron.
result Consistent improvements over existing approaches in synthetic and real-world data.
New algorithm finds best subset in high-dimensional data models.
problem Finding the best subset of predictors in high-dimensional data models.
method Proposes a scalable algorithm using a generalized information criterion.
result Directly proves consistency and oracle property for the best-subset selection.
Empirical Bayes method improves Gaussian sequence model inference.
problem Estimating parameters in correlated Gaussian sequence models.
method Maximum Composite Marginal Likelihood (CML) estimator, leveraging geometric Brascamp-Lieb inequality.
result CML estimator converges at rate \( n_*^{-1/2} \) in weighted Hellinger distance.
Develops a new model for network estimation from multi-variate data.
problem Network estimation from multi-variate point process or time series data.
method Semi-parametric approach based on the monotone single-index multi-variate autoregressive model (SIMAM).
result Achieves optimal rates of convergence and superior performance in prediction and network estimation.
Generalized Linear Models (GLMs) and Single Index Models (SIMs) provide powerful generalizations of linear regression, where the target variable is assumed to be a (possibly unknown) 1-dimensional function of a linear predictor. In general, these problems entail non-convex estimation procedures, and, in practice, itera…
Enhances SDR via Hellinger correlation for better data dependency understanding.
problem Improving sufficient dimension reduction in single-index models.
method Developed a new method using Hellinger correlation for detecting the dimension reduction subspace.
result Significantly enhances and outperforms existing SDR methods through deeper data dependency understanding.
Single Index Models (SIMs) are simple yet flexible semi-parametric models for machine learning, where the response variable is modeled as a monotonic function of a linear combination of features. Estimation in this context requires learning both the feature weights and the nonlinear function that relates features to ob…
Single index model is a powerful yet simple model, widely used in statistics, machine learning, and other scientific fields. It models the regression function as g(<a,x>), where a is an unknown index vector and x are the features. This paper deals with a nonlinear generalization of this framework to allow for a regre…
We propose denoising dictionary learning (DDL), a simple yet effective technique as a protection measure against adversarial perturbations. We examined denoising dictionary learning on MNIST and CIFAR10 perturbed under two different perturbation techniques, fast gradient sign (FGSM) and jacobian saliency maps (JSMA). W…