New research shows testing IIA in discrete choice is nearly impossible with current sample sizes.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
When applied to high-dimensional datasets, feature selection algorithms might still leave dozens of irrelevant variables in the dataset. Therefore, even after feature selection has been applied, classifiers must be prepared to the presence of irrelevant variables. This paper investigates a new training method called Co…
Protocol minimizes disclosure in classification tasks.
Regularization is a popular technique in machine learning for model estimation and avoiding overfitting. Prior studies have found that modern ordered regularization can be more effective in handling highly correlated, high-dimensional data than traditional regularization. The reason stems from the fact that the ordered…
We focus on credal nets, which are graphical models that generalise Bayesian nets to imprecise probability. We replace the notion of strong independence commonly used in credal nets with the weaker notion of epistemic irrelevance, which is arguably more suited for a behavioural theory of probability. Focusing on direct…
New method better identifies irrelevant variables for more accurate treatment effect estimation.
Estimating the difficulty level of math word problems is an important task for many educational applications. Identification of relevant and irrelevant sentences in math word problems is an important step for calculating the difficulty levels of such problems. This paper addresses a novel application of text categoriza…
The paper investigates how irrelevant features affect clustering performance.
Bayesian methods detect significant IIA violations in similarity choice data.
Testing independence is of significant interest in many important areas of large-scale inference. Using extreme-value form statistics to test against sparse alternatives and using quadratic form statistics to test against dense alternatives are two important testing procedures for high-dimensional independence. However…
The problem of finding a reduced dimensionality representation of categorical variables while preserving their most relevant characteristics is fundamental for the analysis of complex data. Specifically, given a co-occurrence matrix of two variables, one often seeks a compact representation of one variable which preser…
Paper extends sparse alternatives to softmax for continuous domains, enabling efficient attention mechanisms.
It is inconceivable how chaotic the world would look to humans, faced with innumerable decisions a day to be made under uncertainty, had they been lacking the capacity to distinguish the relevant from the irrelevant---a capacity which computationally amounts to handling probabilistic independence relations. The highly …
The paper tackles statistical and computational challenges in learning correlated reward models.
Modern machine learning methods often require more data for training than a single expert can provide. Therefore, it has become a standard procedure to collect data from external sources, e.g. via crowdsourcing. Unfortunately, the quality of these sources is not always guaranteed. As additional complications, the data …
We present a general framework for hypothesis testing on distributions of sets of individual examples. Sets may represent many common data sources such as groups of observations in time series, collections of words in text or a batch of images of a given phenomenon. This observation pattern, however, differs from the c…
Improved RL for TBGs by pruning irrelevant tokens and bootstrapping.
LLMs are vulnerable to task-irrelevant data changes, limiting their use for data fitting.
In this article, we define an independence system for a classical knot diagram and prove that the independence system is a knot invariant for alternating knots. We also discuss the exchange property for minimal unknotting sets. Finally, we show that there are knot diagrams where the independence system is a matroid and…
Metaheuristic algorithms (MAs) have seen unprecedented growth thanks to their successful applications in fields including engineering and health sciences. In this work, we investigate the use of a deep learning (DL) model as an alternative tool to do so. The proposed method, called MaNet, is motivated by the fact that …
This paper simplifies OPE in large state spaces using state abstractions.
New theory connects geometry without relying on connections.
In this note, we complete the classification of quasi-alternating Montesinos links. We show that the quasi-alternating Montesinos links are precisely those identified independently by Qazaqzeh-Chbili-Qublan and Champanerkar-Ording. A consequence of our proof is that a Montesinos link is quasi-alternating if and onl…
Method learns representations invariant to task-irrelevant details in reinforcement learning tasks.
Study shows optimal rates for independence testing via U-statistic permutation tests.
Implicit feedback is the simplest form of user feedback that can be used for item recommendation. It is easy to collect and domain independent. However, there is a lack of negative examples. Existing works circumvent this problem by making various assumptions regarding the unconsumed items, which fail to hold when the …
OSIRIS reduces variance in off-policy evaluation by omitting irrelevant states.
Ito-Takimura recently defined a splice-unknotting number for knot diagrams. They proved that this number provides an upper bound for the crosscap number of any prime knot, asking whether equality holds in the alternating case. We answer their question in the affirmative. (Ito has independently proven the same …
High-dimensional data acquired from biological experiments such as next generation sequencing are subject to a number of confounding effects. These effects include both technical effects, such as variation across batches from instrument noise or sample processing, or institution-specific differences in sample acquisiti…
The method approximates stationary distributions of Markov models by truncating irrelevant states.
Anisotropic neural network selects relevant features from datasets.
New unsupervised learning technique learns independent kernels for better machine learning tasks.
Selecting important features in non-linear or kernel spaces is a difficult challenge in both classification and regression problems. When many of the features are irrelevant, kernel methods such as the support vector machine and kernel ridge regression can sometimes perform poorly. We propose weighting the features wit…
In this paper, we build an organization of high-dimensional datasets that cannot be cleanly embedded into a low-dimensional representation due to missing entries and a subset of the features being irrelevant to modeling functions of interest. Our algorithm begins by defining coarse neighborhoods of the points and defin…
Independent component analysis (ICA) is a method for recovering statistically independent signals from observations of unknown linear combinations of the sources. Some of the most accurate ICA decomposition methods require searching for the inverse transformation which minimizes different approximations of the Mutual I…
Entropy regularized OT test assesses independence between samples.
Alternative to likelihood-based LSNM model selection, residual independence testing is more robust to noise misspecification.
The paper sets thresholds for testing correlation in hypergraphs, distinguishing between independent and correlated states.
Simple proof of knot genus theorem using Alexander polynomial.
Sparse PCA selects variables with FDR control for improved performance.
A new method tests conditional independence by transforming it into an unconditional problem using transport maps.
XGBoost fails to accurately identify relevant features, while interpretable methods do.
Proposes Infomax and Domain-Independent Representations for robust causal inference.
Many successful methods have been proposed for learning low dimensional representations on large-scale networks, while almost all existing methods are designed in inseparable processes, learning embeddings for entire networks even when only a small proportion of nodes are of interest. This leads to great inconvenience,…
A key element in transfer learning is representation learning; if representations can be developed that expose the relevant factors underlying the data, then new tasks and domains can be learned readily based on mappings of these salient factors. We propose that an important aim for these representations are to be unbi…
A network removes irrelevant structures from chest radiographs for better analysis.
New bounds for knot complexity based on Jones polynomial coefficients.
We present a novel recurrent neural network (RNN) based model that combines the remembering ability of unitary RNNs with the ability of gated RNNs to effectively forget redundant/irrelevant information in its memory. We achieve this by extending unitary RNNs with a gating mechanism. Our model is able to outperform LSTM…