Bayesian units improve speech recognition with minimal parameters.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This article provides the role of big idea statisticians in future of Big Data Science. We describe the `United Statistical Algorithms' framework for comprehensive unification of traditional and novel statistical methods for modeling Small Data and Big Data, especially mixed data (discrete, continuous).
Most of the parameters in large vocabulary models are used in embedding layer to map categorical features to vectors and in softmax layer for classification weights. This is a bottle-neck in memory constraint on-device training applications like federated learning and on-device inference applications like automatic spe…
New algorithm minimizes regret in multi-armed bandits with network interference.
Randomly chosen primary hidden units and derived secondary units reduce neural network complexity.
Algorithm removes units and layers of neural networks without losing accuracy.
Algorithm learns two-layer residual units using ReLU activations from samples.
We present a novel neural network algorithm, the Tensor Switching (TS) network, which generalizes the Rectified Linear Unit (ReLU) nonlinearity to tensor-valued hidden units. The TS network copies its entire input vector to different locations in an expanded representation, with the location determined by its hidden un…
Investment tool predicts higher returns for Madrid real estate units.
A fast ML method solves complex combinatorial auction problems.
A Bayesian factor graph reduced to normal form consists in the interconnection of diverter units (or equal constraint units) and Single-Input/Single-Output (SISO) blocks. In this framework localized adaptation rules are explicitly derived from a constrained maximum likelihood (ML) formulation and from a minimum KL-dive…
A new method for optimizing regression problems with ReLU units converges.
The paper develops a multi-unit soft sensing model for virtual flow meters that improves few-shot learning.
This paper analyzes and improves the effectiveness of BN techniques in controlling Internal Covariate Shift.
A new algorithm optimizes softmax units in large language models.
New framework for estimating treatment effects in experiments with network interference.
A novel quantum model improves RBM performance and is efficiently trainable.
New framework segments 3D scenes using neural algorithms and sub-Riemannian geometry.
The spectral -support norm enjoys good estimation properties in low rank matrix learning problems, empirically outperforming the trace norm. Its unit ball is the convex hull of rank matrices with unit Frobenius norm. In this paper we generalize the norm to the spectral -support norm, whose additional para…
Deep neural network learning can be formulated as a non-convex optimization problem. Existing optimization algorithms, e.g., Adam, can learn the models fast, but may get stuck in local optima easily. In this paper, we introduce a novel optimization algorithm, namely GADAM (Genetic-Evolutionary Adam). GADAM learns deep …
We aim to create the highest possible quality of treatment-control matches for categorical data in the potential outcomes framework. Matching methods are heavily used in the social sciences due to their interpretability, but most matching methods do not pass basic sanity checks: they fail when irrelevant variables are …
New algorithm speeds up fitting GLLVMs to large datasets.
We investigate the problem of factorizing a matrix into several sparse matrices and propose an algorithm for this under randomness and sparsity assumptions. This problem can be viewed as a simplification of the deep learning problem where finding a factorization corresponds to finding edges in different layers and valu…
We give the first dimension-efficient algorithms for learning Rectified Linear Units (ReLUs), which are functions of the form with . Our algorithm works in the challenging Reliable Agnostic learning model of Kalai, Kanade, and Ma…
Many natural systems, such as neurons firing in the brain or basketball teams traversing a court, give rise to time series data with complex, nonlinear dynamics. We can gain insight into these systems by decomposing the data into segments that are each explained by simpler dynamic units. Building on switching linear dy…
New techniques improve 16-bit training accuracy without 32-bit units.
Improves causal graph learning on dependent binary data.
A new neural model evolves to learn at the synaptic level.
Researchers create a teapot model for Mandelbrot set, proving connectedness.
Optimizes portfolios with discrete units using simulated annealing.
In this work, we propose an infinite restricted Boltzmann machine~(RBM), whose maximum likelihood estimation~(MLE) corresponds to a constrained convex optimization. We consider the Frank-Wolfe algorithm to solve the program, which provides a sparse solution that can be interpreted as inserting a hidden unit at each ite…
There is growing interest in applying machine learning methods to Electronic Medical Records (EMR). Across different institutions, however, EMR quality can vary widely. This work investigated the impact of this disparity on the performance of three advanced machine learning algorithms: logistic regression, multilayer p…
Permutation of any two hidden units yields invariant properties in typical deep generative neural networks. This permutation symmetry plays an important role in understanding the computation performance of a broad class of neural networks with two or more hidden units. However, a theoretical study of the permutation sy…
METASET selects diverse unit cells for efficient data-driven metamaterial design.
Facial expression analysis based on machine learning requires large number of well-annotated data to reflect different changes in facial motion. Publicly available datasets truly help to accelerate research in this area by providing a benchmark resource, but all of these datasets, to the best of our knowledge, are limi…
Restricted Boltzmann Machine (RBM) is a bipartite graphical model that is used as the building block in energy-based deep generative models. Due to numerical stability and quantifiability of the likelihood, RBM is commonly used with Bernoulli units. Here, we consider an alternative member of exponential family RBM with…
The bienergy of smooth maps between Riemannian manifolds, when restricted to unit vector fields, yields two different variational problems depending on whether one takes the full functional or just the vertical contribution. Their critical points, called biharmonic unit vector fields and biharmonic unit sections, form …
Characterizes magnetic unit vector fields on Lie groups.
New algorithm reveals piecewise affine structure of neural networks.
With the abundance of data in recent years, interesting challenges are posed in the area of recommender systems. Producing high quality recommendations with scalability and performance is the need of the hour. Singular Value Decomposition(SVD) based recommendation algorithms have been leveraged to produce better result…
New term ADS describes how machine learning can change user behavior.
We give a polynomial-time algorithm for learning neural networks with one layer of sigmoids feeding into any Lipschitz, monotone activation function (e.g., sigmoid or ReLU). We make no assumptions on the structure of the network, and the algorithm succeeds with respect to {\em any} distribution on the unit ball in …
The paper examines A/B tests in recommendation systems to detect biased algorithm comparisons due to shared data.
We propose a method to decrease the number of hidden units of the restricted Boltzmann machine while avoiding decrease of the performance measured by the Kullback-Leibler divergence. Then, we demonstrate our algorithm by using numerical simulations.
This paper tackles federated learning for automatic latent variable selection in multi-output Gaussian processes.
In this paper we propose and investigate a novel nonlinear unit, called unit, for deep neural networks. The proposed unit receives signals from several projections of a subset of units in the layer below and computes a normalized norm. We notice two interesting interpretations of the unit. First…
A new CNN-based algorithm improves Fourier ptychography for faster, more robust image reconstruction.
New MBL hidden Born machine learns various tasks.