We consider the online multiclass linear classification under the bandit feedback setting. Beygelzimer, Pál, Szörényi, Thiruvenkatachari, Wei, and Zhang [ICML'19] considered two notions of linear separability, weak and strong linear separability. When examples are strongly linearly separable with margin , they prese…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper investigates how data augmentation improves linear separation of manifold data.
New algorithms improve blind source separation for linear-quadratic mixtures.
In this paper, we presented a novel semi-supervised one-class classification algorithm which assumes that class is linearly separable from other elements. We proved theoretically that class is linearly separable if and only if it is maximal by probability within the sets with the same mean. Furthermore, we presented an…
Adam optimizes linear classifiers with separable data.
The paper extends optimal transport for linear separability of sheared distributions in supervised learning.
Neural networks use their hidden layers to transform input data into linearly separable data clusters, with a linear or a perceptron type output layer making the final projection on the line perpendicular to the discriminating hyperplane. For complex data with multimodal distributions this transformation is difficult t…
We study the problem of efficient online multiclass linear classification with bandit feedback, where all examples belong to one of classes and lie in the -dimensional Euclidean space. Previous works have left open the challenge of designing efficient algorithms with finite mistake bounds when the data is linear…
Guiding the design of neural networks is of great importance to save enormous resources consumed on empirical decisions of architectural parameters. This paper constructs shallow sigmoid-type neural networks that achieve 100% accuracy in classification for datasets following a linear separability condition. The separab…
Spectral analysis shows neural networks separate from linear methods in approximating functions.
A new algorithm finds a separating hyperplane with fewer updates.
New PCstar algorithm discovers causal structure of max-linear Bayesian networks.
Shallow nonlinear networks can separate classes linearly with polynomially scaling width.
We introduce a new approach for designing computationally efficient learning algorithms that are tolerant to noise, and demonstrate its effectiveness by designing algorithms with improved noise tolerance guarantees for learning linear separators. We consider both the malicious noise model and the adversarial label nois…
StrADiff separates sources from mixtures without labels, using structured priors.
Study high-dimensional Bayesian linear regression using variational inference.
Develops large-sample theory for non-stationary source separation.
Adversarial noises are linearly separable for random neural networks.
Transformers can learn noisy linear systems with depth and IID data.
A new method separates data points using entropy minimization over a hypercube.
We show linear XOR classification is possible and propose equality separation for anomaly detection.
Graph convolution improves linear separability and generalizes to out-of-distribution data.
We develop randomized (block) coordinate descent (CD) methods for linearly constrained convex optimization. Unlike most CD methods, we do not assume the constraints to be separable, but let them be coupled linearly. To our knowledge, ours is the first CD method that allows linear coupling constraints, without making th…
BELIEF framework interprets GLMs using binary linear models.
New findings on robust learning with well-separated data.
We address the problem of causal discovery from data, making use of the recently proposed causal modeling framework of modular structural causal models (mSCM) to handle cycles, latent confounders and non-linearities. We introduce σ-connection graphs (σ-CG), a new class of mixed graphs (containing undirected, bidirected…
An important issue in neural network research is how to choose the number of nodes and layers such as to solve a classification problem. We provide new intuitions based on earlier results by An et al. (2015) by deriving an upper bound on the number of nodes in networks with two hidden layers such that linear separabili…
We present an approach to learn the dynamics of multiple objects from image sequences in an unsupervised way. We introduce a probabilistic model that first generate noisy positions for each object through a separate linear state-space model, and then renders the positions of all objects in the same image through a high…
KPCA improves OoD detection by separating InD and OoD data.
slimTrain simplifies DNN training by separating features and adapting hyperparameters.
We provide new results concerning label efficient, polynomial time, passive and active learning of linear separators. We prove that active learning provides an exponential improvement over PAC (passive) learning of homogeneous linear separators under nearly log-concave distributions. Building on this, we provide a comp…
This work examines a semi-blind single-channel source separation problem. Our specific aim is to separate one source whose local structure is approximately known, from another a priori unspecified background source, given only a single linear combination of the two sources. We propose a separation technique based on lo…
Paper develops robust methods for panel data with latent groups, improving inference under group separation violations.
New algorithms reduce slate bandit regret for large slates, outperforming existing methods.
Reservoir computing's success depends on mapping different input time series to separable states.
This paper proposes a new evaluation metric and boosting method for weight separability in neural network design. In contrast to general visual recognition methods designed to encourage both intra-class compactness and inter-class separability of latent features, we focus on estimating linear independence of column vec…
Online learning of linear operators between infinite-dimensional spaces is possible but with limitations.
Paper optimizes hyperspherical prototypes for better class separation.
Singing voice separation attempts to separate the vocal and instrumental parts of a music recording, which is a fundamental problem in music information retrieval. Recent work on singing voice separation has shown that the low-rank representation and informed separation approaches are both able to improve separation qu…
Grokking occurs in simple binary logistic classification near linear separability and noise.
For many years, a combination of principal component analysis (PCA) and independent component analysis (ICA) has been used for blind source separation (BSS). However, it remains unclear why these linear methods work well with real-world data that involve nonlinear source mixtures. This work theoretically validates that…
Algorithm finds frequencies, amplitudes, and phases of sinusoids in noisy data.
This paper presents a new ensemble learning method for classification problems called projection pursuit random forest (PPF). PPF uses the PPtree algorithm introduced in Lee et al. (2013). In PPF, trees are constructed by splitting on linear combinations of randomly chosen variables. Projection pursuit is used to choos…
The Trek Separation Theorem (Sullivant et al. 2010) states necessary and sufficient conditions for a linear directed acyclic graphical model to entail for all possible values of its linear coefficients that the rank of various sub-matrices of the covariance matrix is less than or equal to n, for any given n. In this pa…
We present a novel blind source separation (BSS) method, called information geometric blind source separation (IGBSS). Our formulation is based on the log-linear model equipped with a hierarchically structured sample space, which has theoretical guarantees to uniquely recover a set of source signals by minimizing the K…
This paper explains a mechanism called phase collapse that improves image classification accuracy.
Study of eigenvalues in nonlinear kernels for classification of separable data.
The paper improves prediction and testing for signals from a linear combination of translated features with Gaussian noise.