This research secures deployed sentiment analysis models by identifying and defending against attack vectors.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In semi-supervised learning, virtual adversarial training (VAT) approach is one of the most attractive method due to its intuitional simplicity and powerful performances. VAT finds a classifier which is robust to data perturbation toward the adversarial direction. In this study, we provide a fundamental explanation why…
In this paper, we prove that depth with nonlinearity creates no bad local minima in a type of arbitrarily deep ResNets with arbitrary nonlinear activation functions, in the sense that the values of all local minima are no worse than the global minimum value of corresponding classical machine-learning models, and are gu…
The critical locus of the loss function of a neural network is determined by the geometry of the functional space and by the parameterization of this space by the network's weights. We introduce a natural distinction between pure critical points, which only depend on the functional space, and spurious critical points, …
Proposes a new factor to improve BAB strategies by recognizing bad-beta assets.
We propose to improve trust region policy search with normalizing flows policy. We illustrate that when the trust region is constructed by KL divergence constraints, normalizing flows policy generates samples far from the 'center' of the previous policy iterate, which potentially enables better exploration and helps av…
The abstract constructs a set of bad 3-orbifolds and shows how any bad 3-orbifold can be transformed into a good one.
Recent work has noted that all bad local minima can be removed from neural network loss landscapes, by adding a single unit with a particular parameterization. We show that the core technique from these papers can be used to remove all bad local minima from any loss landscape, so long as the global minimum has a loss o…
New attacks reduce bad queries in black-box classifiers, improving effectiveness.
Study finds stocks with common firm fears earn lower returns.
This paper introduces a gradient analysis framework to improve language model performance by rewarding good examples and penalizing bad ones.
New proof shows knot Floer thickness limits bad domains in diagrams.
We provide two fundamental results on the population (infinite-sample) likelihood function of Gaussian mixture models with components. Our first main result shows that the population likelihood function has bad local maxima even in the special case of equally-weighted mixtures of well-separated and spherical…
In this paper, we consider a dynamic asset pricing model in an approximate fractional economy to address empirical regularities related to both investor protection and past information. Our newly developed model features not only in terms with a controlling shareholder who diverts a fraction of the output, but also goo…
Due to the insufficient measurements in the distribution system state estimation (DSSE), full observability and redundant measurements are difficult to achieve without using the pseudo measurements. The matrix completion state estimation (MCSE) combines the matrix completion and power system model to estimate voltage b…
Neural Architecture Search (NAS) has emerged as a promising technique for automatic neural network design. However, existing MCTS based NAS approaches often utilize manually designed action space, which is not directly related to the performance metric to be optimized (e.g., accuracy), leading to sample-inefficient exp…
NICE learns a representation to avoid bad controls in causal inference.
In deep learning, \textit{depth}, as well as \textit{nonlinearity}, create non-convex loss surfaces. Then, does depth alone create bad local minima? In this paper, we prove that without nonlinearity, depth alone does not create bad local minima, although it induces non-convex loss surface. Using this insight, we greatl…
We show that every bad orbifold vector bundle can be realized as the restriction of a good orbifold vector bundle to a suborbifold of the base space. We give an explicit construction of this result in which the Chen-Ruan orbifold cohomology of the two base spaces are isomorphic (as additive groups). This construction i…
This study improves knowledge distillation for RNN-T models with noisy labels.
MCN improves deep neural networks by bettering local minima and generalizing well.
Computational models in fields such as computational neuroscience are often evaluated via stochastic simulation or numerical approximation. Fitting these models implies a difficult optimization problem over complex, possibly noisy parameter landscapes. Bayesian optimization (BO) has been successfully applied to solving…
In this work, we investigate semi-supervised learning (SSL) for image classification using adversarial training. Previous results have illustrated that generative adversarial networks (GANs) can be used for multiple purposes. Triple-GAN, which aims to jointly optimize model components by incorporating three players, ge…
Graph attacks can be successful with just a few bad nodes.
Asymmetries in volatility spillovers are highly relevant to risk valuation and portfolio diversification strategies in financial markets. Yet, the large literature studying information transmission mechanisms ignores the fact that bad and good volatility may spill over at different magnitudes. This paper fills this gap…
Training generative adversarial networks requires balancing of delicate adversarial dynamics. Even with careful tuning, training may diverge or end up in a bad equilibrium with dropped modes. In this work, we improve CS-GAN with natural gradient-based latent optimisation and show that it improves adversarial dynamics b…
Graph Neural Networks (GNNs) are efficient approaches to process graph-structured data. Modelling long-distance node relations is essential for GNN training and applications. However, conventional GNNs suffer from bad performance in modelling long-distance node relations due to limited-layer information propagation. Ex…
Using back-propagation and its variants to train deep networks is often problematic for new users. Issues such as exploding gradients, vanishing gradients, and high sensitivity to weight initialization strategies often make networks difficult to train, especially when users are experimenting with new architectures. Her…
One of the main difficulties in analyzing neural networks is the non-convexity of the loss function which may have many bad local minima. In this paper, we study the landscape of neural networks for binary classification tasks. Under mild assumptions, we prove that after adding one special neuron with a skip connection…
Deep neural networks (DNNs) are known for their vulnerability to adversarial examples. These are examples that have undergone small, carefully crafted perturbations, and which can easily fool a DNN into making misclassifications at test time. Thus far, the field of adversarial research has mainly focused on image model…
We identify a class of over-parameterized deep neural networks with standard activation functions and cross-entropy loss which provably have no bad local valley, in the sense that from any point in parameter space there exists a continuous path on which the cross-entropy loss is non-increasing and gets arbitrarily clos…
Analyzes minima of deep linear networks with weight decay.
We detect and quantify asymmetries in volatility spillovers using the realized semivariances of petroleum commodities: crude oil, gasoline, and heating oil. During the 1987--2014 period we document increasing spillovers from volatility among petroleum commodities that substantially change after the 2008 financial crisi…
Highly connected orbifolds are rare but exist.
The counting grid is a grid of microtopics, sparse word/feature distributions. The generative model associated with the grid does not use these microtopics individually. Rather, it groups them in overlapping rectangular windows and uses these grouped microtopics as either mixture or admixture components. This paper bui…
PyBADS optimizes complex functions quickly and reliably.
Reviews recent findings on neural network landscapes.
Generative Adversarial Networks (GANs) have become a popular method to learn a probability model from data. In this paper, we aim to provide an understanding of some of the basic issues surrounding GANs including their formulation, generalization and stability on a simple benchmark where the data has a high-dimensional…
We investigate the difficulties of training sparse neural networks and make new observations about optimization dynamics and the energy landscape within the sparse regime. Recent work of \citep{Gale2019, Liu2018} has shown that sparse ResNet-50 architectures trained on ImageNet-2012 dataset converge to solutions that a…
Bad models can teach well by replicating noise.
Proposes a method to simulate data for testing credit risk scorecard stability.
In this paper, we study a simple and generic framework to tackle the problem of learning model parameters when a fraction of the training samples are corrupted. We first make a simple observation: in a variety of such settings, the evolution of training accuracy (as a function of training epochs) is different for clean…
A simple modification improves GAN performance by discarding bad samples.
The study analyzes local minima in ReLU networks and finds low probability of bad local minima.
The lattice cohomology of a plumbed 3--manifold associated with a connected negative definite plumbing graph is an important tool in the study of topological properties of , and in the comparison of the topological properties with analytic ones when is realized as complex analytic singularity link. By defini…
Despite their great success in recent years, deep neural networks (DNN) are mainly black boxes where the results obtained by running through the network are difficult to understand and interpret. Compared to e.g. decision trees or bayesian classifiers, DNN suffer from bad interpretability where we understand by interpr…
For the problem of high-dimensional sparse linear regression, it is known that an -based estimator can achieve a "fast" rate on the prediction error without any conditions on the design matrix, whereas in absence of restrictive conditions on the design matrix, popular polynomial-time methods only guarante…
New rule reduces exploration regret to logarithmic, improving bad episode handling.