Study on size and depth of neural networks for approximating benign functions, showing barriers and explicit results.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This work introduces sequential neural beamforming, which alternates between neural network based spectral separation and beamforming based spatial separation. Our neural networks for separation use an advanced convolutional architecture trained with a novel stabilized signal-to-noise ratio loss function. For beamformi…
Estimates statistical power for cluster analysis in biomedical research.
The paper connects decision tree interpretability and robustness through separation.
The main result of this article is a refinement of the well-known subgroup separability results of Hall and Scott for free and surface groups. We show that for any finitely generated subgroup, there is a finite dimensional representation of the free or surface group that separates the subgroup in the induced Zariski to…
The Douglas Rachford algorithm is an algorithm that converges to a minimizer of a sum of two convex functions. The algorithm consists in fixed point iterations involving computations of the proximity operators of the two functions separately. The paper investigates a stochastic version of the algorithm where both funct…
Proves a representation that separates a subgroup in RAAGs.
Existing depth separation results for constant-depth networks essentially show that certain radial functions in , which can be easily approximated with depth networks, cannot be approximated by depth networks, even up to constant accuracy, unless their size is exponential in . However, the func…
In this paper, we presented a novel semi-supervised one-class classification algorithm which assumes that class is linearly separable from other elements. We proved theoretically that class is linearly separable if and only if it is maximal by probability within the sets with the same mean. Furthermore, we presented an…
The issue addressed in this paper is that of testing for common breaks across or within equations of a multivariate system. Our framework is very general and allows integrated regressors and trends as well as stationary regressors. The null hypothesis is that breaks in different parameters occur at common locations and…
Diffusion models can memorize training data, limiting their creativity and privacy.
We provide a detailed study on the implicit bias of gradient descent when optimizing loss functions with strictly monotone tails, such as the logistic loss, over separable datasets. We look at two basic questions: (a) what are the conditions on the tail of the loss function under which gradient descent converges in the…
Study on hyperbolic groups, focusing on separability and splittings.
This paper compares Transformers and RNNs in various tasks, showing size differences.
This article establishes the performance of stochastic blockmodels in addressing the co-clustering problem of partitioning a binary array into subsets, assuming only that the data are generated by a nonparametric process satisfying the condition of separate exchangeability. We provide oracle inequalities with rate of c…
Proves depth 2 neural networks can't approximate certain functions as well as depth 3 networks.
Heuristic algorithm for portfolio optimization reduces solve times to milliseconds.
This paper proposes a novel proximal-gradient algorithm for a decentralized optimization problem with a composite objective containing smooth and non-smooth terms. Specifically, the smooth and nonsmooth terms are dealt with by gradient and proximal updates, respectively. The proposed algorithm is closely related to a p…
We quantify the separation between the numbers of labeled examples required to learn in two settings: Settings with and without the knowledge of the distribution of the unlabeled data. More specifically, we prove a separation by multiplicative factor for the class of projections over the Boolean hypercube o…
We find the wealth distribution for an economic agent in the financial market, in analogy with standard derivation of generaliz Boltzman (Tsallis) factor in statistical mechanics. In this respect, Tsallis entropic index separates two different regimes, the large and small size market. The Pareto like wealth distributio…
Stochastic Gradient Descent phases explained for deep networks.
Numerous algorithms are used for nonnegative matrix factorization under the assumption that the matrix is nearly separable. In this paper, we show how to make these algorithms efficient for data matrices that have many more rows than columns, so-called "tall-and-skinny matrices". One key component to these improved met…
In this paper we address the question of the size distribution of firms. To this aim, we use the Bloomberg database comprising multinational firms within the years 1995-2003, and analyze the data of the sales and the total assets of the separate financial statement of the Japanese and the US companies, and make a compa…
This paper investigates how data augmentation improves linear separation of manifold data.
A fast method estimates Gaussian mixture components without iterative fitting.
Recent deep learning approaches have achieved impressive performance on speech enhancement and separation tasks. However, these approaches have not been investigated for separating mixtures of arbitrary sounds of different types, a task we refer to as universal sound separation, and it is unknown how performance on spe…
Blind source separation algorithms such as independent component analysis (ICA) are widely used in the analysis of neuroimaging data. In order to leverage larger sample sizes, different data holders/sites may wish to collaboratively learn feature representations. However, such datasets are often privacy-sensitive, prec…
New algorithm recovers mixture means even with many outliers.
The paper proves barriers to approximating functions with small weights and depth in neural networks.
Stochastic Gradient Descent (SGD) is a central tool in machine learning. We prove that SGD converges to zero loss, even with a fixed (non-vanishing) learning rate - in the special case of homogeneous linear classifiers with smooth monotone loss functions, optimized on linearly separable data. Previous works assumed eit…
Capitalizing on the need for addressing the existing challenges associated with gesture recognition via sparse multichannel surface Electromyography (sEMG) signals, the paper proposes a novel deep learning model, referred to as the XceptionTime architecture. The proposed innovative XceptionTime is designed by integrati…
Let be a function of the form for . We give a simple proof that shows that poly-size depth two neural networks with (exponentially) bounded weights cannot approximate $f…
Deep neural networks and decision trees operate on largely separate paradigms; typically, the former performs representation learning with pre-specified architectures, while the latter is characterised by learning hierarchies over pre-specified features with data-driven architectures. We unite the two via adaptive neur…
In this paper we study decomposition methods based on separable approximations for minimizing the augmented Lagrangian. In particular, we study and compare the Diagonal Quadratic Approximation Method (DQAM) of Mulvey and Ruszczyński and the Parallel Coordinate Descent Method (PCDM) of Richtárik and Takáč. We show that …
In this article, we study connections between representation theory and efficient solutions to the conjugacy problem on finitely generated groups. The main focus is on the conjugacy problem in conjugacy separable groups, where we measure efficiency in terms of the size of the quotients required to distinguish a distinc…
Paper explains why robust generalization is hard in deep learning models.
Gradient methods avoid overfitting on separable data.
We consider the problem of learning causal networks with interventions, when each intervention is limited in size under Pearl's Structural Equation Model with independent errors (SEM-IE). The objective is to minimize the number of experiments to discover the causal directions of all the edges in a causal graph. Previou…
Speech separation refers to extracting each individual speech source in a given mixed signal. Recent advancements in speech separation and ongoing research in this area, have made these approaches as promising techniques for pre-processing of naturalistic audio streams. After incorporating deep learning techniques into…
We introduce a new and improved characterization of the label complexity of disagreement-based active learning, in which the leading quantity is the version space compression set size. This quantity is defined as the size of the smallest subset of the training data that induces the same version space. We show various a…
We determine the information-theoretic cutoff value on separation of cluster centers for exact recovery of cluster labels in a -component Gaussian mixture model with equal cluster sizes. Moreover, we show that a semidefinite programming (SDP) relaxation of the -means clustering method achieves such sharp threshol…
We solve minimal separator problems in AMP chain graphs and improve structure learning algorithms.
A new text classification method using separable convolution reduces memory consumption.
Amino acid sequence portrays most intrinsic form of a protein and expresses primary structure of protein. The order of amino acids in a sequence enables a protein to acquire a particular stable conformation that is responsible for the functions of the protein. This relationship between a sequence and its function motiv…
Paper calculates distances between strata in Teichmüller space, proving a constant separation.
We introduce a Markovian single point process model, with random intensity regulated through a buffer mechanism and a self-exciting effect controlling the arrival stream to the buffer. The model applies the principle of the Hawkes process in which point process jumps generate a shot-noise intensity field. Unlike the Ha…
Study shortest non-separating curves on non-orientable surfaces, proving NP-hardness and tractability.
A new RNN architecture reduces model size and improves performance.