A new framework improves LSTM performance without adding more parameters.
problem Improving LSTM performance without increasing model complexity.
method A unifying framework of bilinear LSTMs that balances hidden state vector size and weight matrix approximation quality.
result Bilinear LSTMs achieve superior performance compared to linear LSTMs without additional parameters.
First, classes of Markov processes that scale exactly with a Hurst exponent H are derived in closed form. A special case of one class is the Tsallis density, advertised elsewhere as nonlinear diffusion or diffusion with nonlinear feedback. But the Tsallis model is only one of a very large class of linear diffusion with…
Reformulates RBM for unified linear and nonlinear dimensionality reduction.
problem Traditional RBM limitations in handling both linear and nonlinear data.
method Reformulates RBM using MAP and EM, introduces deterministic CD algorithm.
result Reformulated RBM can outperform PCA in nonlinear dimensionality reduction.
Method discovers nonlinear relations from time series data.
problem Identifying directional relations from nonlinear interactions in time series.
method Minimum predictive information regularization method for deep learning.
result Substantially outperforms other methods for learning nonlinear relations.
This study compares various superlearner and deep learning architectures (machine-learning-based and neural-network-based) for classification problems across several simulated and industrial datasets to assess performance and computational efficiency, as both methods have nice theoretical convergence properties. Superl…
Study analyzes learning dynamics in nonlinear perceptrons using stochastic-process approach.
problem Understanding the roles of nonlinearity and input-data distribution in neural network learning.
method Stochastic-process approach to derive flow equations for learning dynamics in nonlinear perceptrons.
result Input-data noise affects learning speed differently under supervised and reinforcement learning.
Study compares bundles with connections to prehomogeneous geometries.
problem Comparing bundle structures to prehomogeneous geometries.
method Analyzes fibered manifolds, jet bundles, and nonlinear PDEs.
result Identifies similarities and differences between bundle structures and prehomogeneous geometries.
The paper revisits discriminative vs. generative classifiers, showing naive Bayes requires fewer samples.
problem Comparing discriminative and generative classifiers in multiclass settings.
method Theoretical analysis and simulations of naive Bayes vs. logistic regression.
result Multiclass naive Bayes requires fewer samples to approach asymptotic error compared to logistic regression.
The paper introduces a new pairs trading model using nonlinear and non-Gaussian state-space models.
problem Developing a robust trading strategy for pairs of assets with non-Gaussian and heteroskedastic innovations.
method A nonlinear and non-Gaussian state-space model for the spread between two assets, with mean reversion modeled as a mean-reverting process.
result The new trading strategy yields significantly higher returns and Sharpe ratios compared to existing methods.
Donor-aware scRNA-seq benchmarks improve classification accuracy in inflammatory bowel disease.
problem Influenza disease classification from scRNA-seq data is prone to donor-level confounding.
method Developed and evaluated three feature representations across two IBD cohorts.
result Compartment-stratified CLR composition and GatedStructuralCFN embeddings outperform linear models in classification accuracy.
Boosting strategies for merging vs. ensembling studies analyzed.
problem Deciding between merging and ensembling studies for boosting.
method Analytical transition point and bias-variance decomposition for boosting with linear learners.
result Theoretical guidelines for merging vs. ensembling studies.
New method selects direct causal parents from large sets of variables.
problem Inferring direct causal parents from many variables, especially nonlinear and cyclic.
method One-vs.-the-rest feature selection approach with theoretical guarantees.
result Significant improvements over existing methods.
A common method of generalizing binary to multi-class classification is the error correcting code (ECC). ECCs may be optimized in a number of ways, for instance by making them orthogonal. Here we test two types of orthogonal ECCs on seven different datasets using three types of binary classifier and compare them with t…
The support vector machine (SVM) is a powerful and widely used classification algorithm. This paper uses the Karush-Kuhn-Tucker conditions to provide rigorous mathematical proof for new insights into the behavior of SVM. These insights provide perhaps unexpected relationships between SVM and two other linear classifier…
Unified framework for linear attribution methods in deep learning.
problem Separate theoretical foundations of XAI attribution methods.
method GRALIS (Gradient-Riesz Averaged Locally-Integrated Shapley) framework.
result Unified representation theory for linear attribution methods.
JULIA combines multi-linear and nonlinear models for tensor completion.
problem Complex patterns in real-world tensors require a unified model.
method JULIA unifies multi-linear and nonlinear models with flexible component assignment and efficient alternating optimization.
result JULIA outperforms existing methods in large-scale tensor completion.
Explains linearizing a nonlinear connection on a pullback bundle.
problem Clarifying the geometric meaning of linearized connections.
method Fiberwise linear approximation of a vector bundle connection.
result Clarifies the geometric meaning of linearized connections.
Study solves inverse problems for equations with fractional nonlinearities.
problem Solving inverse problems for semilinear elliptic equations with fractional power nonlinearities.
method Higher order linearization method adapted for fractional order.
result Results of previous studies remain valid for general power nonlinearities.
This paper introduces TDA and TSI for better business analytics.
problem Nonlinear, multi-scale business datasets under-represented by traditional tools.
method Topological Data Analysis (TDA) and Topological Stability Index (TSI).
result TSI reveals structural variability in business data.
We learn linear models from nonlinear systems using multiple trajectories and regularization.
problem Identifying linear models from data when the underlying dynamics are nonlinear.
method Multiple trajectories data acquisition followed by regularized least squares.
result Learn linearized dynamics with arbitrarily small error given enough samples.
The paper introduces a framework to assess nonlinear causality in financial markets.
problem Identifying and quantifying co-dependence between financial instruments.
method Transfer entropy and convergent cross-mapping methods to assess linear and nonlinear causality.
result Stock indices exhibit significant nonlinear causality, and correlation underestimates causality.
We identify linear models from nonlinear systems with initialization constraints.
problem Identifying linear models from nonlinear systems with initialization constraints.
method Multiple trajectories-based deterministic data acquisition algorithm followed by regularized least squares.
result We provide a finite sample error bound on the learned linearized dynamics.
EML-CD discovers causal mechanisms from neural networks in a structured way.
problem Extracting causal mechanisms from neural network weights is ill-posed.
method Integrates EML operator into causal structure learning, representing each edge mechanism as a gated EML binary tree.
result Achieves SHD=11.2 +/- 0.4 on real data, matching or outperforming existing methods.
The paper provides a non-asymptotic error bound for linear system identification under nonlinear policies.
problem System identification for linear systems with nonlinear and/or time-varying policies under i.i.d. random excitation noises.
method Least square estimation with non-asymptotic error bound for bounded state and action trajectories.
result The error bound is consistent with linear policies and generalizes existing guarantees.
Proposes a partially linear structure to capture nonlinear relationships in mixture of experts models.
problem Suboptimal estimates due to linearity assumption in mixture of experts models.
method Introduces a partially linear structure that incorporates unspecified functions to capture nonlinear relationships.
result Establishes the identifiability of the proposed model under mild conditions and introduces a practical estimation algorithm.
This paper compares classical parametric methods with recently developed Bayesian methods for system identification. A Full Bayes solution is considered together with one of the standard approximations based on the Empirical Bayes paradigm. Results regarding point estimators for the impulse response as well as for conf…
In recent years, the number of papers on Alzheimer's disease classification has increased dramatically, generating interesting methodological ideas on the use machine learning and feature extraction methods. However, practical impact is much more limited and, eventually, one could not tell which of these approaches are…
RFMs transition from linear to nonlinear under specific input-label correlation.
problem Understanding the transition from linear to nonlinear behavior in RFMs.
method Analyzing RFMs under spiked covariance designs, characterizing the interaction between anisotropy and input-label correlation.
result The RFM generalization error is governed by the strength of input-label correlation, leading to a clear nonlinear advantage above a specific boundary.
VR game data for P300 BCI with raccoon vs demon stimuli.
problem Developing confidence metrics for P300 BCI.
method Multiclass labeled P300 dataset in VR game context.
result Estimation of model's confidence in stimulus predictions.
Proposes σ-PCA to learn identifiable linear transformations without whitening.
problem Cannot identify axes with equal variances in PCA.
method Unified model for linear and nonlinear PCA, introducing a missing piece to eliminate rotational indeterminacy.
result Eliminates subspace rotational indeterminacy in PCA.
Policy gradient converges to globally optimal policy in nearly linear-quadratic systems.
problem Finding optimal policies in nonlinear control systems with partial information.
method Policy gradient algorithm designed for nearly linear-quadratic regulators with small Lipschitz nonlinear components.
result Policy gradient algorithm converges to globally optimal policy with linear rate.
This paper analyzes the sample complexity of two timescale reinforcement learning algorithms.
problem Analyzing the sample complexity of two timescale reinforcement learning algorithms.
method Non-asymptotic analysis of linear and nonlinear TDC and Greedy-GQ algorithms under Markovian sampling with constant stepsize.
result The paper provides non-asymptotic convergence results for two timescale linear and nonlinear TDC and Greedy-GQ algorithms.
We present a simple, general technique for reducing the sample complexity of matrix and tensor decomposition algorithms applied to distributions. We use the technique to give a polynomial-time algorithm for standard ICA with sample complexity nearly linear in the dimension, thereby improving substantially on previous b…
Regularizes RNNs to handle long-range dependencies and multiple time scales.
problem Identifying nonlinear dynamical systems with varying time scales and long-range dependencies.
method A simple regularization scheme for vanilla RNNs with ReLU activation.
result Regularized RNNs can solve long-range dependency problems and express slow time scales.
This paper analyzes a simplified strategy for nonlinear control using local linear models and iLQR updates.
problem Nonlinear policy optimization in control systems.
method Iterative estimation of local linear models and iLQR-like policy updates.
result Demonstrates polynomial sample complexity and overcomes exponential problem horizon dependence.
Develops a new approach to study nonlinear PDEs and their singularities.
problem Understanding the propagation domains of solutions to nonlinear PDEs.
method Derived geometric machinery and sheaf theory to study nonlinear PDEs and their singular supports.
result Estimates the domains of propagation for solutions of non-linear systems.
For many years, a combination of principal component analysis (PCA) and independent component analysis (ICA) has been used for blind source separation (BSS). However, it remains unclear why these linear methods work well with real-world data that involve nonlinear source mixtures. This work theoretically validates that…
CDSSL improves representation quality by integrating linear and nonlinear dependencies.
problem Scarcity of labeled data and neglect of nonlinear dependencies in SSL.
method CDSSL combines linear correlations and nonlinear dependencies using HSIC in RKHS.
result CDSSL enhances representation quality on diverse benchmarks.
In many compressive sensing problems today, the relationship between the measurements and the unknowns could be nonlinear. Traditional treatment of such nonlinear relationships have been to approximate the nonlinearity via a linear model and the subsequent un-modeled dynamics as noise. The ability to more accurately ch…
The paper reviews identifiability in linear and nonlinear models, from Gaussian to non-Gaussian.
problem Identifiability issues in latent-variable and structural-equation models, especially in nonlinear cases.
method Review of identifiability theory for linear and nonlinear models, including factor analysis and structural equation models.
result Even nonparametric nonlinear models can be estimated with additional assumptions.
New linear algorithms improve wSVMs for multiclass probability estimation.
problem Estimating conditional probabilities for multiclass problems.
method Proposed baseline learning and OVA learning schemes to improve wSVMs.
result Linear algorithms achieve optimal computational efficiency and good estimation accuracy.
Nonlinear RNNs' memory capacity varies widely, making it impractical.
problem The usefulness of memory capacity as a metric for linear RNNs is questioned.
method Analysis of random nonlinear RNNs with varying input scales.
result Memory capacity of nonlinear RNNs is arbitrary and impractical.
This work examines the stability of GD and SGD near minima, revealing nonlinear dynamics that differ from linear analysis.
problem The stability of optimization algorithms like GD and SGD near minima is not well understood.
method The authors derive an exact criterion for stable oscillations of GD near minima in the multivariate setting, considering high-order derivatives.
result Nonlinear dynamics can diverge in expectation even if a single batch is unstable, challenging linear analysis.
Extends linear MDP to handle nonlinear rewards.
problem Restrictive linear MDP assumption limits real-world applicability.
method Proposes Generalized Linear MDP (GLMDP) with GLMs for rewards.
result Develops offline RL algorithms achieving suboptimality guarantees.
Kernel method approximates Koopman operator eigenfunctions.
problem Complexity of computing Koopman operator spectra.
method Kernel-based approach to construct principal eigenfunctions.
result Principal eigenfunctions match linearization eigenvalues.
We simplify complex regression coefficients using linearization and feature comparison.
problem Interpreting high-dimensional regression coefficients from nonlinear responses.
method Developed a linearization method to derive feature coefficients and compare them with regression coefficients.
result Shows how regression coefficients relate to linearized feature coefficients and how they change under regularization.
Measurement noise limits the advantage of nonlinear models over linear models in biomedical prediction
problem Nonlinear models vs. linear models in biomedical prediction
method Measurement reliability
result Measurement noise blurs the population-optimal predictor
Extends importance sampling to nonlinear models using adjoint operators.
problem Lack of tools for identifying important data points in nonlinear models.
method Introduces adjoint operator for nonlinear maps, generalizes norm and leverage scores.
result Generalized scores provide approximation guarantees for nonlinear mappings.