Polynomial neural networks explore thresholds for maximum expressiveness.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New method trains neural networks with threshold activation functions efficiently.
Overwhelming theoretical and empirical evidence shows that mildly overparametrized neural networks -- those with more connections than the size of the training data -- are often able to memorize the training data with accuracy. This was rigorously proved for networks with sigmoid activation functions and, very …
Study active learning of PTFs with derivative access.
This work interprets GELU and related activations via a first-order loss function.
"Sparse" neural networks, in which relatively few neurons or connections are active, are common in both machine learning and neuroscience. Whereas in machine learning, "sparsity" is related to a penalty term that leads to some connecting weights becoming small or zero, in biological brains, sparsity is often created wh…
Study proposes an active subsampling method for estimating individualized thresholds in high-dimensional data.
In active learning, the user sequentially chooses values for feature and an oracle returns the corresponding label . In this paper, we consider the effect of feature noise in active learning, which could arise either because itself is being measured, or it is corrupted in transmission to the oracle, or the o…
Percolation on complex networks has been used to study computer viruses, epidemics, and other casual processes. Here, we present conditions for the existence of a network specific, observation dependent, phase transition in the updated posterior of node states resulting from actively monitoring the network. Since tradi…
Novel active learning framework using sparse approximation for efficient model training.
Improved learning bounds for corrupted data using thresholded gradient descent.
Paper proposes a new activation function to reduce overfitting and large weight update issues.
In this work, we derive a generic overcomplete frame thresholding scheme based on risk minimization. Overcomplete frames being favored for analysis tasks such as classification, regression or anomaly detection, we provide a way to leverage those optimal representations in real-world applications through the use of thre…
Giving provable guarantees for learning neural networks is a core challenge of machine learning theory. Most prior work gives parameter recovery guarantees for one hidden layer networks, however, the networks used in practice have multiple non-linear layers. In this work, we show how we can strengthen such results to d…
We consider the problem of online active learning to collect data for regression modeling. Specifically, we consider a decision maker with a limited experimentation budget who must efficiently learn an underlying linear population model. Our main contribution is a novel threshold-based algorithm for selection of most i…
This article studies the financial integration between the six main Latin American markets and the US market in a nonlinear framework. Using the threshold cointegration techniques of Hansen and Seo (2002), we show significant threshold stock market linkages between Mexico, Chile and the US. Thus, the dynamics of these …
Bayesian nonparametric model predicts user activity and intervention success.
A cost-effective approach to label acquisition using active learning markets.
Activation functions are essential for deep learning methods to learn and perform complex tasks such as image classification. Rectified Linear Unit (ReLU) has been widely used and become the default activation function across the deep learning community since 2012. Although ReLU has been popular, however, the hard zero…
New oracle uses uncertainty for active classification with noisy feedback.
We present a new algorithm based on an gradient ascent for a general Active Exploration bandit problem in the fixed confidence setting. This problem encompasses several well studied problems such that the Best Arm Identification or Thresholding Bandits. It consists of a new sampling rule based on an online lazy mirror …
In this paper we consider two semimartingales driven by diffusions and jumps. We allow both for finite activity and for infinite activity jump components. Given discrete observations we disentangle the {\it integrated covariation} (the covariation between the two diffusion parts, indicated by IC) from the co-jumps. Thi…
Study proposes active learning method for estimating robust regions in uncertain function evaluations.
We consider a univariate semimartingale model for (the logarithm of) an asset price, containing jumps having possibly infinite activity (IA). The nonparametric threshold estimator of the integrated variance IV proposed in Mancini 2009 is constructed using observations on a discrete time grid, and precisely it sums up t…
Unified formula for training dynamics of linear networks combining lazy and balanced regimes.
New model improves community detection in networks with strong assortativity.
Paper proposes efficient AL algorithms for optimizing product performance under environmental variability.
Hidden Markov Models detect hand gestures from wearable sEMG signals.
This paper describes the black hole threshold in a moduli space of spherically symmetric spacetimes.
DIET-SNN optimizes SNNs for faster, lower-energy image classification.
Active learning method improves local model validity estimation.
Deep Convolutional Sparse Coding (D-CSC) is a framework reminiscent of deep convolutional neural networks (DCNNs), but by omitting the learning of the dictionaries one can more transparently analyse the role of the activation function and its ability to recover activation paths through the layers. Papyan, Romano, and E…
In this paper, we propose exact passive-aggressive (PA) online algorithms for learning to rank. The proposed algorithms can be used even when we have interval labels instead of actual labels for examples. The proposed algorithms solve a convex optimization problem at every trial. We find exact solution to those optimiz…
Power spectrum densities for the number of tick quotes per minute (market activity) on three currency markets (USD/JPY, EUR/USD, and JPY/EUR) for periods from January 1999 to December 2000 are analyzed. We find some peaks on the power spectrum densities at a few minutes. We develop the double-threshold agent model and …
We present a novel active learning algorithm for community detection on networks. Our proposed algorithm uses a Maximal Expected Model Change (MEMC) criterion for querying network nodes label assignments. MEMC detects nodes that maximally change the community assignment likelihood model following a query. Our method is…
The mushroom body is the key network for the representation of learned olfactory stimuli in Drosophila and insects. The sparse activity of Kenyon cells, the principal neurons in the mushroom body, plays a key role in the learned classification of different odours. In the specific case of the fruit fly, the sparseness o…
We describe a framework for designing efficient active learning algorithms that are tolerant to random classification noise and are differentially-private. The framework is based on active learning algorithms that are statistical in the sense that they rely on estimates of expectations of functions of filtered random e…
For using neural networks in safety critical domains, it is important to know if a decision made by a neural network is supported by prior similarities in training. We propose runtime neuron activation pattern monitoring - after the standard training process, one creates a monitor by feeding the training data to the ne…
A fall is an abnormal activity that occurs rarely, so it is hard to collect real data for falls. It is, therefore, difficult to use supervised learning methods to automatically detect falls. Another challenge in using machine learning methods to automatically detect falls is the choice of engineered features. In this p…
Adaptive batch sizes improve active learning efficiency and flexibility.
Study on critical points in random neural networks, revealing three regimes based on activation function.
The relaxation dynamics of aftershocks after large volatility shocks are investigated based on two high-frequency data sets of the Shanghai Stock Exchange Composite (SSEC) index. Compared with previous relevant work, we have defined main financial shocks based on large volatilities rather than large crashes. We find th…
ZDP detects drift in large language models without labels, proving key theorems and metrics.
The paper analyzes methods for sparse Bayesian regression in nonlinear system identification.
Quantile gradient boosted trees outperform other models in predicting NO2 concentration distributions.
In this paper, we consider the problem of recovering a sparse signal based on penalized least squares formulations. We develop a novel algorithm of primal-dual active set type for a class of nonconvex sparsity-promoting penalties, including , bridge, smoothly clipped absolute deviation, capped and mini…
Market activity scales near a constant of 0.632 in intrinsic time.
In this letter, an age of information (AoI)-aware transmission power and resource block (RB) allocation technique for vehicular communication networks is proposed. Due to the highly dynamic nature of vehicular networks, gaining a prior knowledge about the network dynamics, i.e., wireless channels and interference, in o…