Overfitting frequently occurs in deep learning. In this paper, we propose a novel regularization method called Drop-Activation to reduce overfitting and improve generalization. The key idea is to drop nonlinear activation functions by setting them to be identity functions randomly during training time. During testing, …
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Deep neural networks have dramatically achieved great success on a variety of challenging tasks. However, most successful DNNs have an extremely complex structure, leading to extensive research on model compression.As a significant area of progress in model compression, traditional gradual pruning approaches involve an…
Deep neural networks (DNNs) have been proven to have many redundancies. Hence, many efforts have been made to compress DNNs. However, the existing model compression methods treat all the input samples equally while ignoring the fact that the difficulties of various input samples being correctly classified are different…
Out-of-distribution (OOD) detection approaches usually present special requirements (e.g., hyperparameter validation, collection of outlier data) and produce side effects (e.g., classification accuracy drop, slower energy-inefficient inferences). We argue that these issues are a consequence of the SoftMax loss anisotro…
New metrics reveal oversmoothing in GNNs more accurately than traditional methods.
Chunking is a significant CL problem, accounting for half of performance drop, and current methods don't address it.
Round balls minimize liquid drop model volumes ≤ 1.
Machine learning is currently dominated by largely experimental work focused on improvements in a few key tasks. However, the impressive accuracy numbers of the best performing models are questionable because the same test sets have been used to select these models for multiple years now. To understand the danger of ov…
Empirical study shows GANs overfit and drop modes when training is deterministic.
The study examines mass drop and multiplicity in mean curvature flow.
Deep neural networks are vulnerable against adversarial examples. In this paper, we propose to train and test the networks with randomly subsampled images with high drop rates. We show that this approach significantly improves robustness against adversarial examples in all cases of bounded L0, L2 and L_inf perturbation…
This paper analyzes the configurations of shapes that shows a spacelike liquid drop in Minkowski space deposited over a spacelike plane . We assume the presence of a uniform gravity field directed toward and that the volume of the drop is prescribed. Our interest are the liquid drops that are critical points of …
System predicts HIV patients at risk of dropping out of care.
This study assesses the impact of non-IID data in federated learning, revealing significant performance drops.
Dropping a tiny fraction of preferences can significantly alter the rankings of top LLMs.
When training clinical prediction models from electronic health records (EHRs), a key concern should be a model's ability to sustain performance over time when deployed, even as care practices, database systems, and population demographics evolve. Due to de-identification requirements, however, current experimental pra…
Study addresses RTB model performance drops due to distribution shifts.
Dropout has proven to be an effective technique for regularization and preventing the co-adaptation of neurons in deep neural networks (DNN). It randomly drops units with a probability during the training stage of DNN. Dropout also provides a way of approximately combining exponentially many different neural networ…
Paper resolves Huisken's conjecture without strict genus drop theorem.
Study on liquidation games with market drop-out, proving unique equilibria.
Drop-Muon updates only some layers, speeding up training.
Deep learning models struggle with new data in stock price trend prediction.
RDIS fills missing values in time series data explicitly.
The Frank-Wolfe (FW) algorithm has been widely used in solving nuclear norm constrained problems, since it does not require projections. However, FW often yields high rank intermediate iterates, which can be very expensive in time and space costs for large problems. To address this issue, we propose a rank-drop method …
Under-parameterization hinders deep RL's efficiency.
The past few years have witnessed the fast development of different regularization methods for deep learning models such as fully-connected deep neural networks (DNNs) and Convolutional Neural Networks (CNNs). Most of previous methods mainly consider to drop features from input data and hidden layers, such as Dropout, …
Study connects hyperbolic geometry to membrane shapes.
In this paper we compare market price fluctuations with the response to fundamental price drops within the Lux-Marchesi model which is able to reproduce the most important stylized facts of real market data. Major differences can be observed between the decay of spontaneous fluctuations and of changes due to external p…
This paper shows the susceptibility of spectrogram-based audio classifiers to adversarial attacks and the transferability of such attacks to audio waveforms. Some commonly used adversarial attacks to images have been applied to Mel-frequency and short-time Fourier transform spectrograms, and such perturbed spectrograms…
Select-DC reduces GFLOPS for uncertainty estimation in neural networks.
Method diagnoses model performance under distribution shifts.
This work introduces a method to attribute model performance drops to distribution shifts.
Paper proposes a method to maintain ASR performance on new tasks without forgetting old ones.
Google uses continuous streams of data from industry partners in order to deliver accurate results to users. Unexpected drops in traffic can be an indication of an underlying issue and may be an early warning that remedial action may be necessary. Detecting such drops is non-trivial because streams are variable and noi…
Lower Ricci curvature bound prevents first Betti number from dropping more than dimension in collapsing manifolds.
Improves performance of deep GCNs by controlling node feature variance.
Enhances time-series modeling by dropping patches, improving efficiency and adaptability.
CODA uses a new dropout technique inspired by constructivism learning to improve deep learning performance.
Study reveals statistical bias in dataset replication, reducing accuracy drop from 11-14% to 3.6%.
A new loss function improves neural networks' out-of-distribution detection without side effects.
The study of heavy-tailed distributions in economic and financial systems has been widely addressed since financial time series has become a research subject.After the eighties, several "highly improbable" market drops were observed (e.g. the 1987 stock market drop known as "Black Monday" and on even more recent ones, …
Deep neural networks achieve state-of-the-art results on several tasks while increasing in complexity. It has been shown that neural networks can be pruned during training by imposing sparsity inducing regularizers. In this paper, we investigate two techniques for group-wise pruning during training in order to improve …
Random walks on Fuchsian Schottky groups have harmonic measures with lower dimension.
Study on stability of 3D sessile drops, identifying degenerate kernel.
Study evaluates how well question-answering models generalize to new data types.
Many tasks in natural language understanding require learning relationships between two sequences for various tasks such as natural language inference, paraphrasing and entailment. These aforementioned tasks are similar in nature, yet they are often modeled individually. Knowledge transfer can be effective for closely …
We study the long time existence theory for a non local flow associated to a free boundary problem for a trapped non liquid drop. The drop has free boundary components on two horizontal plates and its free energy is anisotropic and axially symmetric. For axially symmetric initial surfaces with sufficiently large volume…
Dropout improves MIL performance on noisy WSI classification.