This work identifies redundant tests in conditional-independence-based discovery that can improve graphical model accuracy.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new cross-validation method reduces redundancy and improves model performance.
Redundancy in AI perception systems doesn't guarantee independent error occurrences.
Optimized GPRNN reduces model complexity and overfitting, improving performance.
The study analyzes optimization trajectories in neural networks to reveal redundancy and redundancy-reducing strategies.
TS-RSR improves batch Bayesian Optimization by minimizing redundancy and focusing on high uncertainty points.
New method reconstructs stable sentiment signals from sparse news data.
Scientific discovery is limited by hypothesis redundancy, and hybrid methods can exploit non-local exploration.
With the advent of Big Data era, data reduction methods are highly demanded given its ability to simplify huge data, and ease complex learning processes. Concretely, algorithms that are able to filter relevant dimensions from a set of millions are of huge importance. Although effective, these techniques suffer from the…
Laboratory testing and medication prescription are two of the most important routines in daily clinical practice. Developing an artificial intelligence system that can automatically make lab test imputations and medication recommendations can save costs on potentially redundant lab tests and inform physicians of a more…
Williams and Beer (2010) proposed a nonnegative mutual information decomposition, based on the construction of redundancy lattices, which allows separating the information that a set of variables contains about a target variable into nonnegative components interpretable as the unique information of some variables not p…
New method quantifies redundant information using information bottleneck.
Redundancy improves learning stability and generalization in structured systems.
Transformers reduce redundancy by focusing on invariant relational quantities.
Convolutional neural networks (CNN) are generally designed with a heuristic initialization of network architecture and trained for a certain task. This often leads to overparametrization after learning and induces redundancy in the information flow paths within the network. This robustness and reliability is at the inc…
We present an efficient algorithm for simultaneously training sparse generalized linear models across many related problems, which may arise from bootstrapping, cross-validation and nonparametric permutation testing. Our approach leverages the redundancies across problems to obtain significant computational improvement…
Data acquisition, storage and management have been improved, while the key factors of many phenomena are not well known. Consequently, irrelevant and redundant features artificially increase the size of datasets, which complicates learning tasks, such as regression. To address this problem, feature selection methods ha…
Paper proposes redundancy-free features for zero-shot object recognition.
This work explains scaling laws as redundancy laws in deep learning.
Proposes a novel network for CTR prediction by learning modality-specific and modality-invariant representations.
Study on neural networks to identify redundancy issues in safe machine learning.
This works extends the Random Embedding Bayesian Optimization approach by integrating a warping of the high dimensional subspace within the covariance kernel. The proposed warping, that relies on elementary geometric considerations, allows mitigating the drawbacks of the high extrinsic dimensionality while avoiding the…
Robust subgroup discovery finds non-redundant, statistically significant subgroups.
We simplify SSL by approximating redundant structural components with low-rank factorization.
Redundancy in deep neural network (DNN) models has always been one of their most intriguing and important properties. DNNs have been shown to overparameterize, or extract a lot of redundant features. In this work, we explore the impact of size (both width and depth), activation function, and weight initialization on th…
This paper introduces a new measure to identify model redundancy in compressed CNNs.
Deep neural network is a state-of-art method in modern science and technology. Much statistical literature have been devoted to understanding its performance in nonparametric estimation, whereas the results are suboptimal due to a redundant logarithmic sacrifice. In this paper, we show that such log-factors are not nec…
Feature selection is among the most important components because it not only helps enhance the classification accuracy, but also or even more important provides potential biomarker discovery. However, traditional multivariate methods is likely to obtain unstable and unreliable results in case of an extremely high dimen…
Current clinical practice to monitor patients' health follows either regular or heuristic-based lab test (e.g. blood test) scheduling. Such practice not only gives rise to redundant measurements accruing cost, but may even lead to unnecessary patient discomfort. From the computational perspective, heuristic-based test …
Contrastive learning works well with redundant data views.
This work introduces RISE to explain LLMs more reliably by distinguishing essential context.
SAE-FiRE extracts key financial info from long documents, improving earnings surprise predictions.
Locality-sensitive hashing speeds up web app security testing.
In dynamic selection (DS) techniques, only the most competent classifiers, for the classification of a specific test sample are selected to predict the sample's class labels. The more important step in DES techniques is estimating the competence of the base classifiers for the classification of each specific test sampl…
The study reveals flaws in pruning criteria and proposes a new assumption for better filter selection.
Information theory provides ideas for conceptualising information and measuring relationships between objects. It has found wide application in the sciences, but economics and finance have made surprisingly little use of it. We show that time series data can usefully be studied as information -- by noting the relations…
New method distinguishes feature relevance in non-linear contexts.
New framework improves multivariate time series forecasting by minimizing redundant information.
Paper compresses deep neural networks by eliminating redundant neurons.
Large datasets have been crucial to the success of deep learning models in the recent years, which keep performing better as they are trained with more labelled data. While there have been sustained efforts to make these models more data-efficient, the potential benefit of understanding the data itself, is largely unta…
FactorMiner discovers financial alpha factors with low redundancy.
Wide neural networks' last hidden layers split into groups of redundant neurons.
FlipOut prunes neural networks by flipping weights' signs, achieving high sparsity.
Spectral dimensionality reduction algorithms are widely used in numerous domains, including for recognition, segmentation, tracking and visualization. However, despite their popularity, these algorithms suffer from a major limitation known as the "repeated Eigen-directions" phenomenon. That is, many of the embedding co…
This paper enhances ML algorithms by improving data locality and reducing redundancy.
Machine learning (ML) is probably the first and foremost used technique to deal with the size and complexity of the new generation of data. In this paper, we analyze one of the means to increase the performances of ML algorithms which is exploiting data locality. Data locality and access patterns are often at the heart…
In 1980 J. Powell proposed that five specific elements sufficed to generate the Goeritz group of any Heegaard splitting of . This conjecture remains unresolved for genus . Here a short argument shows that one of his proposed generators is redundant, in fact a consequence of three of the other four.
Improved change point detection using matched filters for non-parametric tests.