New classifier combines locally linear kernels for fast and accurate non-linear classification.
problem Developing a fast and accurate non-linear classifier.
method Combines locally linear classifiers using a ℓ1 Multiple Kernel Learning (MKL) problem with scalable MKL training for streaming kernels. result The resulting classifier achieves high accuracy with fast inference time.
Study examines various non-linear activation functions for improving neural network performance on MNIST classification.
problem Improving neural network performance on MNIST classification tasks.
method Empirical analysis of non-linear activation functions, including depth and weight initialization effects.
result Optimal neural network architecture with best activation function and weight initialization yields impressive results.
LCC algorithm maps instances to a central space for better classification.
problem Improving classification accuracy for various datasets.
method Formulated as a quadratic program, simplified to a linear program, uses kernel functions for non-linear cases.
result LCC outperforms other methods in accuracy on standard datasets.
GraphAIR improves graph representation learning by capturing non-linear interactions.
problem Challenges in capturing non-linear interactions in graph data.
method Integrates neighborhood aggregation and interaction modeling.
result Demonstrates improved performance on node classification and link prediction tasks.
Improves sparse recovery with non-linear Fourier features.
problem Sparse recovery challenges with non-linear Fourier features.
method Characterizes sufficient data points for perfect recovery.
result Sufficient data points depend on kernel matrix.
This paper explains a mechanism called phase collapse that improves image classification accuracy.
problem Understanding the role of non-linearities and convolutional filters in image classification.
method Demonstrates phase collapse as a mechanism that eliminates spatial variability and linearly separates classes.
result Phase collapse improves classification accuracy, while thresholding operators degrade performance.
Simpler GNNs perform well on graph classification tasks.
problem Understanding what Graph Neural Networks (GNNs) learn and their complexity.
method Dissected GNNs into graph filtering and set function, linearizing them separately.
result Linear graph filtering with non-linear set function is efficient and powerful.
New method distinguishes feature relevance in non-linear contexts.
problem Finding relevant features with preserved redundancies.
method Random forest models and statistical methods.
result Distinguishes strong from weak feature relevance in non-linear problems.
Detects adversarial examples with non-linear dimensionality reduction.
problem Vulnerability of deep neural networks to adversarial examples.
method Combining non-linear dimensionality reduction and density estimation.
result Effective detection of adversarial examples by non-adaptive attackers.
Extends OC-KSR for multi-task one-class classification.
problem Improving one-class classification performance with shared information.
method Linear and non-linear structure learning mechanisms for multi-task one-class classification.
result Improved performance on multiple one-class problems.
This paper evaluates different non-linear activation functions for deep neural networks on MNIST.
problem Improving performance of deep neural networks on MNIST classification task.
method Introduction and evaluation of various non-linear activation functions, analysis of deeper networks, and investigation of weight initialisation methods.
result Different activation functions have distinct characteristics and can improve performance on MNIST.
Improved classification model for high-cardinality categorical predictors.
problem Handling high-cardinality categorical predictors and non-linear data.
method Data-driven binning of spline functions and shrinkage estimators.
result Improved classification precision with interpretable predictors.
Novel approach for creating interpretable classifiers using bilevel optimization of split-rules in NLDTs.
problem Creating highly accurate and easily interpretable classifiers for practical applications.
method Representing classifiers as assemblies of simple mathematical rules using NLDTs with evolutionary bilevel optimization.
result The approach ensures interpretability while achieving high accuracy on various classification problems.
Haar scattering networks improve pattern recognition across various tasks.
problem Improving pattern recognition in diverse tasks like regression and classification.
method Stacking convolutional filters based on Haar wavelets followed by non-linear operators.
result Outperformed best algorithms in 4 out of 18 data classification problems.
LQF linearizes deep models for better interpretability.
problem Lack of interpretability in deep neural networks.
method Simple modifications to architecture, loss function, and optimization.
result Comparable performance to non-linear fine-tuning, with interpretability.
Langevin algorithms improve training of very deep neural networks, especially for image classification.
problem Training very deep neural networks is challenging due to increased non-linearity and the risk of getting stuck in local minima.
method Comparison of Langevin and non-Langevin algorithms for training deep neural networks, introduction of Layer Langevin algorithm.
result Langevin algorithms, especially Layer Langevin, lead to significant improvements in training deep neural networks, particularly for image classification tasks.
In the two parts of this paper we solve a problem of De Rham, proving that Reidemeister torsion invariants determine topological equivalence of linear G-representations, for G a finite cyclic group. Methods in controlled K-theory and surgery theory are developed to establish, and effectively calculate, a necessary and …
A new game-theoretic approach optimizes complex rate metrics.
problem Optimizing non-decomposable performance metrics and rate constraints.
method Extending two-player game approaches to a three-player game, seeking equilibrium.
result Generalizes and improves upon existing algorithms for constrained optimization.
We propose a novel kernel based post selection inference (PSI) algorithm, which can not only handle non-linearity in data but also structured output such as multi-dimensional and multi-label outputs. Specifically, we develop a PSI algorithm for independence measures, and propose the Hilbert-Schmidt Independence Criteri…
GP-KAN uses Gaussian Processes in KANs for robust, parameter-efficient non-linear modeling.
problem Non-linear modeling with limited parameters and uncertainty estimates.
method Integrates Gaussian Processes into Kolmogorov Arnold Networks (KANs) for robust non-linear modeling.
result GP-KAN achieves 98.5% accuracy on MNIST with 80k parameters compared to 1.5M for state-of-the-art models.
This study evaluates handwriting features to diagnose Parkinson's disease.
problem Diagnosing Parkinson's disease through handwriting analysis.
method Kinematic, geometrical, and non-linear features were evaluated using K-nearest neighbors, support vector machines, and random forest classifiers.
result Up to 93.1% accuracy in classifying Parkinson's disease and healthy subjects.
Develops a Riemannian archetypal analysis for interpretable non-linear data.
problem Limited performance of classical archetypal analysis on non-linear data.
method Riemannian geometry for data-driven pullback, geodesic convex combinations, convex relaxation followed by non-convex refinement.
result Combines interpretability of classical archetypal analysis with expressive power of modern non-linear models.
A scattering transform defines a signal representation which is invariant to translations and Lipschitz continuous relatively to deformations. It is implemented with a non-linear convolution network that iterates over wavelet and modulus operators. Lipschitz continuity locally linearizes deformations. Complex classes o…
Rocket algorithm classifies time-series data efficiently using random projections and natural sparsity.
problem Time-series classification challenges in diverse fields.
method Random convolutional kernels, non-linear transformation, compressed sensing framework.
result Rocket algorithm preserves discriminative patterns in time-series data and expresses inherent sparsity.
New approach predicts stock price synchronization using RNNs and LSTMs.
problem Forecasting synchronization of stock prices in the Indian market.
method Utilizing recurrence plots and CRQA for non-linear analysis, RNNs and LSTMs for prediction.
result Accuracy of 0.98 and F1 score of 0.83 in predicting stock price synchronization.
The least-squares support vector machine is a frequently used kernel method for non-linear regression and classification tasks. Here we discuss several approximation algorithms for the least-squares support vector machine classifier. The proposed methods are based on randomized block kernel matrices, and we show that t…
Selecting important features in non-linear or kernel spaces is a difficult challenge in both classification and regression problems. When many of the features are irrelevant, kernel methods such as the support vector machine and kernel ridge regression can sometimes perform poorly. We propose weighting the features wit…
New findings show DNC is not optimal for deep models, revealing a low-rank bias.
problem Theoretical limitations of DNC in non-linear models and multi-class classification.
method Analysis of non-linear models of arbitrary depth in multi-class classification.
result DNC stops being optimal for DUFM when going beyond two layers or two classes, due to a low-rank bias.
The fundamental tool in the classification of orthogonal coordinate systems in which the Hamilton-Jacobi and other prominent equations can be solved by a separation of variables are second order Killing tensors which satisfy the Nijenhuis integrability conditions. The latter are a system of three non-linear partial dif…
This paper explores how MoE layers improve deep learning performance.
problem Understanding the Mixture-of-Experts (MoE) layer in deep learning.
method Formal study of MoE layer's effectiveness and mechanism.
result MoE layer improves performance by leveraging cluster structure and non-linearity.
Deep convolutional networks provide state of the art classifications and regressions results over many high-dimensional problems. We review their architecture, which scatters data with a cascade of linear filter weights and non-linearities. A mathematical framework is introduced to analyze their properties. Computation…
Layer-wise relevance propagation (LRP) is a recently proposed technique for explaining predictions of complex non-linear classifiers in terms of input variables. In this paper, we apply LRP for the first time to natural language processing (NLP). More precisely, we use it to explain the predictions of a convolutional n…
NON model improves tabular data classification accuracy.
problem Tabular data classification in real-world applications.
method Field-wise network, across field network, operation fusion network.
result NON significantly outperforms state-of-the-art models.
Paper presents a new method for multiclass classification using hyperplane arrangements.
problem Developing efficient multiclass classifiers.
method Mixed integer programming formulations with hyperplane arrangements, kernel trick adaptation, and dimensionality reductions.
result Our proposal outperforms other methods in multiclass classification tasks.
Analyzes generalization error in generalized linear models, explaining double descent phenomenon.
problem Understanding generalization of machine learning models in high dimensions.
method Develops a framework to characterize asymptotic generalization error for generalized linear models.
result Rigorously explains the double descent phenomenon in generalized linear models.
Survey of text classification algorithms for complex documents.
problem Understanding and classifying complex texts using machine learning.
method Discusses various text feature extractions, dimensionality reduction methods, and classification algorithms.
result Overview of text classification techniques and their limitations.
Hadwiger's Theorem states that Euclidean-invariant convex-continuous valuations of definable sets are linear combinations of intrinsic volumes. We lift this result from sets to data distributions over sets, specifically, to definable real-valued functions on n-dimensional Euclidean space. This generalizes intrinsic vol…
Randomized ICA and LDA methods improve HSI classification efficiency.
problem Overcome curse of dimensionality in hyperspectral images.
method Proposes RFFICA and RFFLDA using Random Fourier features to handle non-linearities and scalability issues.
result Demonstrates improved overall and per-class accuracies compared to conventional kernel ICA and LDA.
HCBM improves deep learning explainability by non-linear concept aggregation.
problem Lack of explainable and accurate predictions in deep learning for high-stake decisions.
method Introduce Hoeffding Concept Bottleneck Models (HCBM) using Hoeffding functional decomposition of gradient-boosted trees for non-linear and sparse concept aggregation.
result HCBM outperforms standard linear CBM and is robust to interconcept leakage.
Two new algorithms reduce feature space while preserving non-linear relationships.
problem High-dimensional data and overfitting issues.
method Bias-variance analysis for non-linear transformations and generalized linear models.
result Competitive performance on regression and classification tasks.
A new method for one-class classification using ellipsoidal encapsulation.
problem One-class classification for data optimization.
method Iterative transformation into an optimized subspace with regularization terms.
result Better results in one-class classification compared to existing methods.
Two Fisher information matrix estimators are analyzed for neural networks, focusing on their variances and trade-offs.
problem Estimating the Fisher information matrix in neural networks due to its high computational cost.
method Examined two popular diagonal Fisher information matrix estimators and their variances in neural networks for regression and classification.
result The variances of the estimators depend on the non-linearity with respect to different parameter groups and should not be neglected.
Python library automates feature engineering and selection for linear models.
problem Difficulties in training and explaining complex machine learning models.
method Automated feature engineering and selection for linear models.
result Improves prediction accuracy of linear models while retaining interpretability.
Water saturation is an important property in reservoir engineering domain. Thus, satisfactory classification of water saturation from seismic attributes is beneficial for reservoir characterization. However, diverse and non-linear nature of subsurface attributes makes the classification task difficult. In this context,…
New method builds robust trees from noisy data.
problem Building accurate classification trees from noisy labeled data.
method Combines SVM-like splitting rules and label noise detection.
result Effective in detecting and mitigating label noise.
Study uses EEG features HFD and SampEn to detect depression with high accuracy.
problem Diagnosing depression reliably and accurately.
method Applied Higuchi Fractal Dimension and Sample Entropy on EEG signals using seven machine learning algorithms.
result Good classification possible even with small EEG data, achieving high accuracy.
Develops a risk-averse classification method based on coherent risk measures.
problem Designing a classifier that considers risk in classification problems.
method Uses coherent measures of risk and risk sharing ideas to design a risk-averse classifier.
result The risk-sharing classification problem is equivalent to an optimization problem with unequal weights.
We present an approach to model time series data from resting state fMRI for autism spectrum disorder (ASD) severity classification. We propose to adopt kernel machines and employ graph kernels that define a kernel dot product between two graphs. This enables us to take advantage of spatio-temporal information to captu…