Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

98195293390 · Jun 202019922001200920182026
48 results for Uninformative Features

Reconstruction-based learning produces uninformative features for perception tasks.

problem Misalignment between reconstruction-based learning and perception tasks.
method Investigated the impact of input space reconstruction on feature learning for perception tasks.
result Reconstruction-based learning allocates model capacity to a subspace with uninformative features for perception tasks.

Adding uninformative labels improves tumor segmentation in low-data mammography.

problem Improving tumor segmentation in mammography with limited data.
method Used seemingly uninformative labels from non-expert annotators to turn a multi-label task into a multi-class problem.
result Performance gains in tumor segmentation are achieved in low-data settings with additional uninformative labels.

Neural networks compress uninformative input directions, improving test error.

problem Data lie in a high-dimensional space but labels vary along a lower-dimensional manifold.
method One-hidden layer network trained with gradient descent, analyzing weight evolution and compression.
result Compression factor λ ∼ √p improves test error, with β Feature > β Lazy.

Sparse Convex Clustering improves clustering performance in high-dimensional data.

problem Distortion in convex clustering performance with uninformative features.
method Introduces Sparse Convex Clustering with an adaptive group-lasso penalty and a tuning criterion based on clustering stability.
result Demonstrates improved clustering performance through feature selection.

The paper sets bounds on how much regret is unavoidable in adaptive LQR with unknown B-matrix.

problem Understanding the limits of adaptive LQR with unknown B-matrix.
method Local asymptotic minimax regret lower bounds using van Trees' inequality and Bellman error representation.
result Logarithmic regret is impossible if the parametrization induces an uninformative optimal policy.

New optimization criteria improve variational autoencoders for clearer images and latent features.

problem Improving clarity and informativeness of variational autoencoders' latent features and samples.
method Proposed new optimization criteria and a sequential VAE model.
result New criteria help generate clearer images and more informative latent features.

Modeling market dynamics with informed and uninformed traders and fads.

problem Optimizing market making in a market with fads, informed, and uninformed traders.
method Characterizing the optimal liquidity provision problem in a market with fads, informed, and uninformed traders, considering both complete and partial information.
result The price of liquidity is a function of the proportion of informed traders, and strategies ignoring fads underperform.

Study shows mutual funds add little value for uninformed investors.

problem Understanding the performance of actively managed equity mutual funds for uninformed investors.
method Constructed a reference portfolio using prices and supply information, analyzed various subsets of funds, and compared to market index.
result Mutual funds provide insignificant alpha for uninformed investors, with negative and significant alpha when compared to the market index.

Informed traders strategically reveal noisier signals, making prices less responsive to public information.

problem How informed traders strategically reveal signals impacts market prices and utility.
method Modeling a market with an informed trader, an uninformed trader, and liquidity providers, proving equilibrium existence.
result In equilibrium, the insider strategically reveals a noisier signal, making prices less responsive to public information.

Study dynamic equilibrium with insider and general uninformed agent preferences.

problem Analyzing asymmetric information and general utility functions in a continuous-time economy.
method Introducing a new method to prove existence of a partial communication equilibrium (PCE) for agents with general utility functions.
result Identify the equilibrium price in the small and large risk aversion limits for agents with power utility.

Study automates feature selection and clustering for HFT stock price forecasting.

problem Manual feature selection and clustering for high-frequency trading (HFT) stock price forecasting.
method Dual competitive feature importance mechanism and clustering via shallow neural network topology.
result Enhanced forecasting ability of the RBFNN regressor through automated feature selection and clustering.

UAFS selects features to improve imputation and prediction accuracy in datasets with missing data.

problem Missing data challenges imputation and feature selection in high-dimensional datasets.
method Uncertainty-aware feature selection (UAFS) that incorporates missingness uncertainty.
result UAFS improves imputation accuracy and prediction accuracy across various datasets and levels of missingness.

New method speeds up Bayesian inference for complex simulators.

problem Challenges in Bayesian inference for complex stochastic simulators with intractable likelihood functions.
method Optimization Monte Carlo framework reformulated as deterministic optimization problems with gradient-based methods.
result Accurate posterior inference with reduced runtimes compared to existing methods.

New geometric interpretation explains over-parameterized models and adversarial perturbations.

problem Geometric understanding of over-parameterized regression and adversarial perturbations.
method Alternative geometric interpretation of regression in feature space.
result Adversarial perturbations are a natural feature of biased models due to underlying geometry.

Study shows overparametrization can shift and bend loss landscapes, affecting signal recovery.

problem Understanding how overparametrization affects loss landscapes in neural networks.
method Field theory analysis of Hessian spectrum at initialization.
result Overparametrization can shift the BBP transition point, potentially reaching weak-recovery threshold.

A simple strategy optimizes broker-client trading, reducing price discounts for informed traders.

problem Optimizing broker-client trading to balance client flow and informed trader losses.
method Modelled as a stochastic control problem, derived optimal strategy in closed form, introduced algorithm.
result Optimal strategy reduces price discounts for informed traders, balancing client flow and informed trader losses.

Paper identifies problematic baselines in Shapley value explanations and proposes a reweighting mechanism.

problem Identifying and addressing the suboptimality of baselines in Shapley value feature importance analysis.
method Analyzed suboptimality of baselines, identified problematic baseline, generalized uninformativeness, and designed a reweighting mechanism.
result Proposed uncertainty-based reweighting mechanism effectively accelerates computation and improves explanation quality.

The paper optimizes trading strategies for assets modeled by a randomized Brownian bridge.

problem Optimizing trading strategies for assets with uninformative noise and unknown terminal prices.
method Modeling asset price evolution with an exponential randomized Brownian bridge and solving for optimal trading strategies numerically.
result Disconnected continuation/exercise regions appear under certain prior distributions.

DWTS uses observational data to improve clinical trial efficiency.

problem Lack of definitive conclusions from randomized clinical trials due to insufficient patient cohorts and confounding biases.
method DWTS combines observational data with randomized clinical trials using Doubly Debiased LASSO (DDL) to identify reliable covariates.
result DWTS reduces cumulative regret in clinical trials compared to standard methods.

Letter analyzes training dynamics of a nonlinear contrastive learning model in high dimensions.

problem Understanding training dynamics of nonlinear contrastive learning models in high-dimensional settings.
method High-dimensional analysis using McKean-Vlasov PDEs and low-dimensional ODEs.
result The model's performance evolves according to specific ODEs, revealing features like feature learnability and noise effects.

Study explores GAN dynamics for high-dimensional subspace learning.

problem Subspace learning in high-dimensional datasets.
method Single-layer GAN model with multi-feature discriminators.
result GANs outperform conventional methods in capturing informative subspace.

We introduce a new representation learning algorithm suited to the context of domain adaptation, in which data at training and test time come from similar but different distributions. Our algorithm is directly inspired by theory on domain adaptation suggesting that, for effective domain transfer to be achieved, predict…

2014-12-15abs ↗pdf ↗

This paper studies how AMMs can minimize losses from arbitrage while retaining uninformed trading activity.

problem Minimizing losses from arbitrage in AMMs while retaining uninformed trading activity.
method Modeling arbitrage dynamics and sensitivity to fee choices, mapping to a random walk with a reward scheme.
result AMMs can maximize value retention by optimizing fee structures.

We extend the theory of asymmetric information in mispricing models for stocks following geometric Brownian motion to constant relative risk averse investors. Mispricing follows a continuous mean--reverting Ornstein--Uhlenbeck process. Optimal portfolios and maximum expected log--linear utilities from terminal wealth f…

2011-01-06abs ↗pdf ↗

Graph Neural Networks outperform the Weisfeiler-Lehman algorithm in representation power.

problem Limited representation power of Graph Neural Networks compared to the Weisfeiler-Lehman algorithm.
method Algebraic analysis using eigenvalue decomposition of graph operators.
result Graph Neural Networks produce more discriminative representations than the Weisfeiler-Lehman algorithm.

The paper discusses the impact of prior densities on Bayesian model selection.

problem The sensitivity of marginal likelihood to prior choice in Bayesian model selection.
method Analyzes the role of prior densities in model selection, discusses improper priors, and proposes solutions.
result Marginal likelihood can be sensitive to prior choice, but improper priors can still be used with caution.

PARIS reduces imbalanced regression datasets by pruning uninformative samples.

problem Imbalanced regression where models focus on high-frequency regions, ignoring rare but impactful events.
method PARIS uses the representer theorem to compute a closed-form representer deletion residual for iterative pruning of the training set.
result PARIS reduces training set by up to 75% while preserving or improving overall performance, outperforming other methods.

New algorithm reduces online learning regret in uninformed Markov games.

problem Achieving no external regret in uninformed Markov games is impossible.
method Empirical Nash-value regret, parameter-free algorithm, adaptive restart.
result Achieves O(min{K+(CK)1/3,LK})O(\min \{\sqrt{K} + (CK)^{1/3},\sqrt{LK}\}) regret bound.

New method estimates position bias without manual interventions for better search engine rankings.

problem Presentation bias confounds relevance signals in search engines.
method Proposes a method for consistent propensity estimation without manual relevance judgments.
result Initial studies confirm scalability, accuracy, and robustness of the approach.

SaliencyMix augments images with salient patches to improve model generalization.

problem Improving deep learning model generalization through better data augmentation.
method Carefully selects salient patches from images and mixes them with the target image.
result Achieves state-of-the-art top-1 error rates and robustness against adversarial attacks.

A technique called 'prior laundering' uses legacy reconstructions to create uncertainty in Bayesian inverse problems.

problem Uncertainty in Bayesian inverse problems when data is uninformative.
method Using an archive of legacy reconstructions to create uncertainty in the posterior distribution, averaging the legacy posterior over measurements.
result The uncertainty reported in the posterior is inherited from the legacy reconstructions, not from the data itself.

Predicts the number of polynomial additions in Buchberger's algorithm using machine learning.

problem Predict the number of polynomial additions in Buchberger's algorithm.
method Multiple linear regression and recursive neural network models trained on ideal generator statistics.
result Machine learning can predict the number of polynomial additions in Buchberger's algorithm.

A new framework CyGen models joint distributions using cyclic conditionals.

problem Modeling a joint distribution using only two conditional models without relying on an uninformative prior.
method Developed a general theory for operable equivalence criteria for compatibility and sufficient conditions for determinacy. Proposed CyGen framework and methods to achieve compatibility and determinacy.
result CyGen better fits data and captures more representative features compared to models using an uninformative prior.

This paper investigates the equilibrium interactions between trading targets and private information in a multi-period Kyle (1985) market. There are two investors who each follow dynamic trading strategies: A strategic portfolio rebalancer who engages in order splitting to reach a cumulative trading target and an uncon…

2015-02-07abs ↗pdf ↗

Tree-based models outperform deep learning on tabular data, especially for medium-sized datasets.

problem Understanding why tree-based models outperform deep learning on tabular data.
method Extensive benchmarks of tree-based and deep learning models on 45 datasets, accounting for hyperparameters.
result Tree-based models remain state-of-the-art on medium-sized tabular data, even without hyperparameter optimization.