We establish upper bounds for the minimal number of hidden units for which a binary stochastic feedforward network with sigmoid activation probabilities and a single hidden layer is a universal approximator of Markov kernels. We show that each possible probabilistic assignment of the states of n output units, given t…
We present the data on wealth and income distributions in the United Kingdom, as well as on the income distributions in the individual states of the USA. In all of these data, we find that the great majority of population is described by an exponential distribution, whereas the high-end tail follows a power law. The di…
In this article we present an alternative model for the distribution of household incomes in the United States. We provide arguments from two differing perspectives which both yield the proposed income distribution curve, and then fit this curve to empirical data on household income distribution obtained from the Unite…
In this paper we propose and investigate a novel nonlinear unit, called Lp unit, for deep neural networks. The proposed Lp unit receives signals from several projections of a subset of units in the layer below and computes a normalized Lp norm. We notice two interesting interpretations of the Lp unit. First…
We study unit horizontal bundles associated with Riemannian submersions. First we investigate metric properties of an arbitrary unit horizontal bundle equipped with a Riemannian metric of the Cheeger-Gromoll type. Next we examine it from the Gromov-Hausdorff convergence theory point of view, and we state a collapse the…
We have presented a novel technique of detecting intermittencies in a financial time series of the foreign exchange rate data of U.S.- Euro dollar(US/EUR) using a combination of both statistical and spectral techniques. This has been possible due to Continuous Wavelet Transform (CWT) analysis which has been popularly a…
Conditional restricted Boltzmann machines are undirected stochastic neural networks with a layer of input and output units connected bipartitely to a layer of hidden units. These networks define models of conditional probability distributions on the states of the output units given the states of the input units, parame…
Bayesian units improve speech recognition with minimal parameters.
problem Improving speech recognition models with fewer parameters.
method Derived Bayesian recurrent units integrated into deep learning frameworks.
result Adding Bayesian units improves speech recognition performance.
A new method for online prediction uncertainty quantification in non-exchangeable panel data.
problem Challenges in quantifying predictive uncertainty for non-exchangeable panel data.
method Online conformal prediction framework for non-exchangeable panel data, using similarity weights and adaptive miscoverage levels.
result Improves coverage on worst-covered target units through adaptive interval-width allocation.
New LTC RNNs can approximate any continuous system with fewer units.
problem Approximating continuous dynamical systems with neural networks.
method Introducing LTC RNNs with variable time-constant synaptic transmission.
result LTC RNNs can approximate any n-dimensional continuous dynamical system. We generalize recent theoretical work on the minimal number of layers of narrow deep belief networks that can approximate any probability distribution on the states of their visible units arbitrarily well. We relax the setting of binary units (Sutskever and Hinton, 2008; Le Roux and Bengio, 2008, 2010; Montúfar and Ay,…
We present a probabilistic variant of the recently introduced maxout unit. The success of deep neural networks utilizing maxout can partly be attributed to favorable performance under dropout, when compared to rectified linear units. It however also depends on the fact that each maxout unit performs a pooling operation…
ReLU activations lead to smoother learning curves compared to sigmoidal activations in neural networks.
problem Comparing the performance of ReLU and sigmoidal activations in neural networks.
method Analytical computation of learning curves in shallow networks with different activation functions.
result ReLU networks exhibit continuous transitions in performance, while sigmoidal networks show discontinuous transitions.
Researchers develop a method to measure treatment effects in settings with shared states.
problem Measuring treatment effects in settings with shared states like prices, recommendations, or social signals.
method Double machine learning (DML) theorem with conditions for efficient inference under shared-state interference.
result Efficient estimation of average direct effect (ADE) and global average treatment effect (GATE) in various models.
Bayesian SHMM discovers acoustic units from unlabeled speech.
problem Discovering language-specific acoustic units from unlabeled speech.
method Bayesian Subspace Hidden Markov Model (SHMM) trained on labeled data to find new acoustic units on target language.
result Significantly outperforms previous HMM-based systems and compares favorably with Variational Auto Encoder-HMM.
Improves TTS accuracy by correcting context-dependent units.
problem Improves text-to-speech accuracy through speaker adaptation.
method Statistical model predicting context-dependent phonetic unit classes and their mean error values.
result Corrected boundaries of units improve TTS accuracy compared to HMM segmentation.
Paper introduces a simpler gated RNN structure to better capture long-term dependencies.
problem Difficulty in learning long-term dependencies in RNNs.
method Proposes a grouped distributor unit (GDU) with partitioned hidden states and adaptive update rates.
result GDU outperforms LSTM and GRU on various tasks, including pathological and natural data.
A new LSTM variant retains semantic information across time.
problem Maintaining semantic information in LSTM hidden states over time.
method Persistent Recurrent Unit (PRU) with a feedforward layer.
result PRU outperforms conventional LSTM in three tasks.
New ODE-Block handles stateful layers with continuous-in-depth functions using basis functions.
problem Handling stateful layers in ODE-Nets.
method Formulate ODE-Block using continuous-in-depth functions with basis function expansions.
result Enables state-of-the-art performance and reduces memory footprint.
We prove that the refined approach -- our extension of the Yakovenko et al. formalism -- is universal in the sense that it describes well both household incomes in the European Union and the individual incomes in the United States for social classes of any income. This formalism allowed the study of the impact of the r…
This work proposes efficient unit pruning for neural networks to reduce model size.
problem Redundant parameters in neural networks can be pruned without performance loss.
method Proposes unit-wise pruning over parameter pruning, introduces saliency scores, and defines dead units.
result 5x model size reduction on MNIST using unit-wise pruning.
New proof for symmetric spaces with rectangular lattices.
problem Characterizing symmetric spaces with rectangular unit lattices.
method Explicit construction of isometric embeddings and analysis of root systems.
result Symmetric spaces with rectangular unit lattices are symmetric R-spaces.
Study shows structured reservoirs improve deep ESN performance.
problem Improving performance of deep reservoir computing networks.
method Investigated structured reservoir topologies in deep ESNs.
result Structured reservoirs significantly enhance predictive performance.
Dropout learning is analyzed as ensemble learning to prevent overfitting.
problem Overfitting in deep learning models.
method Dropout learning ignores some inputs and hidden units with a probability, p, and combines them with the learned network.
result Combining neglected hidden units with the learned network can be seen as ensemble learning.
Lipschitz RNNs improve stability and performance in various tasks.
problem Improving stability and performance of RNNs.
method Introduced a Lipschitz recurrent unit with a linear and Lipschitz nonlinear component for stability analysis.
result Lipschitz RNNs outperform existing units on benchmark tasks.
New MBL hidden Born machine learns various tasks.
problem Learning from quantum many-body systems.
method MBL dynamics and hidden units for training.
result Enhanced trainability and stability in learning.
CRUs model irregular time series with continuous hidden states.
problem Handling irregular time intervals in sequential data.
method Continuous Recurrent Units (CRUs) that integrate hidden states via a linear stochastic differential equation.
result CRUs outperform methods based on neural ordinary differential equations in irregular time series interpolation.
Linear RC shows hierarchical temporal patterns in state signals.
problem Understanding hierarchical temporal representations in deep RNNs.
method Used linear recurrent units and frequency analysis on state signals.
result Linear RC reveals intrinsic hierarchical temporal structure.
Non-parametric method predicts multi-stream longitudinal data evolution.
problem Predicting the evolution of multi-stream longitudinal data for an in-service unit.
method Decomposes each stream into eigenfunctions and FPC scores, uses Gaussian process prior and empirical Bayesian updating.
result Framework outperforms state-of-the-art approaches and achieves high predictive accuracy.
New deep learning model robust to adversarial attacks using stochastic LWTA units.
problem Adversarial robustness in deep learning networks.
method Introduces deep networks with stochastic LWTA activations, combining them with Bayesian non-parametric tools.
result Achieves high robustness to adversarial perturbations, outperforming state-of-the-art methods.
Improved GRU model with weighted time-delay feedback for long-term dependencies.
problem Modeling long-term dependencies in sequential data.
method Introducing a gated recurrent unit (GRU) with a weighted time-delay feedback mechanism.
result τ-GRU outperforms state-of-the-art models on various tasks.
New model explains complex, nonlinear systems with simpler units that switch based on observations.
problem Complex, nonlinear systems with switching dynamics.
method Recurrent switching linear dynamical systems (recSLDS) model.
result Models the switching behavior of simpler units based on observations or latent states.
Convolutional LSTM networks outperform GRU in EEG seizure detection.
problem Seizure detection in EEG signals.
method Comparison of LSTM and GRU units, hybrid CNN-RNN architecture, various initialization and regularization methods.
result Convolutional LSTM networks achieve 30% sensitivity at 6 false alarms per 24 hours.
Examines US equity risk premiums amid COVID-19.
problem Analyzing equity risk premiums during the pandemic.
method Not specified in the abstract.
result Not specified in the abstract.
Since learning is typically very slow in Boltzmann machines, there is a need to restrict connections within hidden layers. However, the resulting states of hidden units exhibit statistical dependencies. Based on this observation, we propose using l1/l2 regularization upon the activation possibilities of hidden unit…
This study analyzes economic policy uncertainty indices using visibility graphs.
problem Understanding the role of economic policy uncertainty in global economies.
method Visibility graph algorithm applied to economic policy uncertainty indices.
result The economic policy uncertainty indices exhibit persistent behavior and scale-free networks.
New Convolutional Unit improves Batch Whitening performance.
problem Improving the efficiency and effectiveness of Batch Whitening.
method Proposes a new Convolutional Unit that aligns with Batch Whitening theory and empirically analyzes the original Convolutional Unit.
result Significantly improved performance on multiple image classification datasets.
Paper proposes dual recurrent attention units for VQA models.
problem Comprehending visual and textual data for accurate question answering.
method Introduces and evaluates recurrent attention mechanisms in VQA models.
result Dual Recurrent Attention Units (RAUs) improve VQA performance.
IC-Network improves CNNs by integrating elastic collision units.
problem Designing more effective basic units in neural networks.
method Developed IC layer and IC block units combining the IC structure with convolution operations.
result Significant performance improvements in existing CNNs, reducing top-1 error from 22.85% to 21.49% on imagenet.
RUM improves RNN's long-term memory by using unitary matrices.
problem Limited capacity of RNN to manipulate long-term memory.
method Proposes Rotational Unit of Memory (RUM) with unitary matrices.
result RUM learns long-term dependencies and improves state-of-the-art results.
Improved video prediction with bijective Gated Recurrent Units.
problem Ill-posed future video prediction with high variability and error propagation.
method Introduces bijective Gated Recurrent Units for state sharing in auto-encoders.
result Significant reduction in computational cost and memory usage compared to state-of-the-art approaches.
The disbalance of Supply and Demand is typically considered as the driving force of the markets. However, the measurement or estimation of Supply and Demand at price different from the execution price is not possible even after the transaction. An approach in which Supply and Demand are always matched, but the rate $I=…
Using an analog of the boundary element method in engineering and science, we analyze and model unemployment rate in Austria, Italy, the Netherlands, Sweden, Switzerland, and the United States as a function of inflation and the change in labor force. Originally, the model linking unemployment to inflation and labor for…
The paper introduces a US crime index to assess financial losses from property and cyber crimes.
problem Lack of indices evaluating crime's financial impact on investments.
method Developed an index-based insurance portfolio using FBI financial losses data.
result Real estate, ransomware, and government impersonation are major risk contributors.
Better neural arithmetic logic units improve cell counting model generalization.
problem Neural networks struggle with high cell counts outside training data range.
method Introduced Neural Arithmetic Logic Units (NALU) for arithmetic operations in existing architectures.
result Improved cell counting accuracy for higher numeric ranges with better generalization.
ENRNN uses eigenvalue normalization for short-term memory in RNNs.
problem Vanishing/exploding gradient problem and long-term dependency modeling.
method Eigenvalue normalization of recurrent matrix to simulate short-term memory.
result ENRNN outperforms existing RNN variants in experiments.
In this work, we propose a novel recurrent neural network (RNN) architecture. The proposed RNN, gated-feedback RNN (GF-RNN), extends the existing approach of stacking multiple recurrent layers by allowing and controlling signals flowing from upper recurrent layers to lower layers using a global gating unit for each pai…
Machine learning predicts homicide clearance rates with SHAP explaining key features.
problem Predicting and explaining homicide clearance rates in the US.
method Nine algorithmic approaches compared; XGBoost selected. SHAP used for feature importance.
result XGBoost best predicts national homicide clearance rates; SHAP reveals key features.