Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

285785113 · Jun 202019922001200920172026
48 results for hypothesis removal

A statistical framework for removing unwanted data domains in machine learning.

problem Removing unwanted data domains in machine learning while preserving desired performance.
method Modeling domains as probability distributions and using hypothesis testing to select samples to remove.
result Characterization of allowable edited data distributions and removal-preservation Pareto frontiers for various distribution families.

The recently proposed Lottery Ticket Hypothesis of Frankle and Carbin (2019) suggests that the performance of over-parameterized deep networks is due to the random initialization seeding the network with a small fraction of favorable weights. These weights retain their dominant status throughout training -- in a very r…

2019-05-19abs ↗pdf ↗

Study on surfaces in flag threefold with constraints on twistor fibers.

problem Understanding the arrangement and existence of twistor fibers in surfaces of specific bidegree.
method Analyzing surfaces of bidegree (1,d) in the flag threefold, proving existence and non-existence of twistor fibers.
result Existence and non-existence of surfaces containing specific numbers of twistor fibers, with improved results for d=2 and d=3.

We generalize the Weinstein-Moser theorem on the existence of nonlinear normal modes (i.e., periodic orbits) near an equilibrium in a Hamiltonian system to a theorem on the existence of relative periodic orbits near a relative equilibrium in a Hamiltonian system with continuous symmetries. More specifically we signific…

1999-06-01abs ↗pdf ↗

Recent work establishes dataset difficulty and removes annotation artifacts via partial-input baselines (e.g., hypothesis-only models for SNLI or question-only models for VQA). When a partial-input baseline gets high accuracy, a dataset is cheatable. However, the converse is not necessarily true: the failure of a parti…

2019-05-14abs ↗pdf ↗

A new method for joint noise removal and trend estimation from sparse signals.

problem Jointly removing noise and estimating trends from sparse signals.
method PENDANTSS combines SOOT/SPOQ penalties with BEADS algorithm in a Trust-Region block alternating variable metric forward-backward approach.
result Outperforms comparable methods in deconvolving analytical chemistry signals.

New theorem removes uniform finite upper bound for shrinkability of null decompositions.

problem Shrinkability of null decompositions with non-singleton elements.
method Defining squeezable and squashable subsets, proving their equivalence, and applying these definitions to null decompositions.
result Any null decomposition of a compact metric space whose non-singleton elements are recursively squeezable is shrinkable.

Optimal rates for vector-valued regression on various norms.

problem Optimal rates for vector-valued ridge regression on continuous norms.
method Combining standard capacity assumptions with tensor product constructions of vector-valued interpolation spaces.
result Optimal rates for vector-valued ridge regression, independent of output space dimension.

Ahpatron improves online kernel learning with tighter mistake bounds.

problem Improving mistake bounds in online kernel learning with budget constraints.
method Introducing Ahpatron, a new model that uses an aggressive updating rule and a budget maintenance mechanism to approximate AVP.
result Ahpatron achieves tighter mistake bounds compared to previous models.

Recent advances in statistical theory, together with advances in the computational power of computers, provide alternative methods to do mass-univariate hypothesis testing in which a large number of univariate tests, can be properly used to compare MEEG data at a large number of time-frequency points and scalp location…

2014-06-25abs ↗pdf ↗

Symmetric observations don't necessarily imply symmetric causal explanations.

problem Inferring causal models from observed correlations is challenging and computationally intensive.
method An explicit example using a tripartite probability distribution over binary events.
result Symmetries in observations cannot be used to reduce the hypothesis space of causal models.

The paper characterizes hyperbolic manifolds and graphs verifying a specific isoperimetric inequality.

problem Understanding the relationship between hyperbolicity and isoperimetric inequalities in manifolds and graphs.
method Characterization of hyperbolic manifolds and graphs with isoperimetric inequality, using Gromov boundary.
result Having a pole is a necessary condition for verifying the isoperimetric inequality, which can be removed.

We identify a condition for regularity of optimal transport maps that requires only three derivatives of the cost function, for measures given by densities that are only bounded above and below. This new condition is equivalent to the weak Ma-Trudinger-Wang condition when the cost is C4C^4. Moreover, we only require (n…

2012-12-19abs ↗pdf ↗

The search for efficient, sparse deep neural network models is most prominently performed by pruning: training a dense, overparameterized network and removing parameters, usually via following a manually-crafted heuristic. Additionally, the recent Lottery Ticket Hypothesis conjectures that, for a typically-sized neural…

2019-12-10abs ↗pdf ↗

A new method removes biases in data integration by using surrogate control outcomes.

problem Data integration methods can be biased due to data-dependent processes.
method Post-integrated inference method using surrogate control outcomes to account for latent heterogeneity.
result The method provides consistent and efficient estimators under minimal assumptions and potential misspecifications.

Study finds TVL doesn't predict cryptocurrency returns.

problem Assumption of TVL predicting returns in crypto markets.
method Examined TVL-sorted portfolios against crypto market returns, using various TVL measures.
result TVL-sorted portfolios' returns are linear functions of crypto market returns, replicable with standard tools.

Researchers compute Wodzicki residue for pseudo-differential operators on compact Lie groups.

problem Computing the Wodzicki residue for pseudo-differential operators on compact Lie groups.
method Analytic continuation of traces and matrix-valued symbols.
result Main theorem complementary to [2], removing ellipticity hypothesis.

We clarify the status of log-periodicity associated with speculative bubbles preceding financial crashes. In particular, we address Feigenbaum's [2001] criticism and show how it can be rebuked. Feigenbaum's main result is as follows: ``the hypothesis that the log-periodic component is present in the data cannot be reje…

2001-06-26abs ↗pdf ↗

New approach removes data influence in high dimensions with single step.

problem Efficiently removing data influence in high-dimensional settings with strong convexity and smoothness assumptions.
method Introduces ε-Gaussian certifiability and analyzes Newton method performance.
result Single Newton step followed by Gaussian noise achieves privacy and accuracy.

Let X be a hyperbolic surface and H the fundamental group of a hyperbolic 3-manifold that fibers over the circle with fiber X. Using the Birman exact sequence, H embeds in the mapping class group Mod(Y) of the surface Y obtained by removing a point from X. We prove that a subgroup G in H is convex cocompact in Mod(Y) i…

2012-08-13abs ↗pdf ↗

Pruning is a well-established technique for removing unnecessary structure from neural networks after training to improve the performance of inference. Several recent results have explored the possibility of pruning at initialization time to provide similar benefits during training. In particular, the "lottery ticket h…

2019-03-05abs ↗pdf ↗

The problem of non-stationarity in financial markets is discussed and related to the dynamic nature of price volatility. A new measure is proposed for estimation of the current asset volatility. A simple and illustrative explanation is suggested of the emergence of significant serial autocorrelations in volatility and …

2009-11-26abs ↗pdf ↗

Study on the spectrum of drift Laplacian on Ricci expanders.

problem Analyzing the spectrum of the drift Laplacian on Ricci expanders.
method Investigation of discrete spectrum under proper potential function, asymptotic behavior of potential function, and computation of eigenvalues.
result Discrete spectrum of the drift Laplacian on Ricci expanders with bounded Ricci curvature.

Study explores DNNs' reliance on existing vs. new features in physiological signals.

problem Understanding how deep neural networks discover new features in physiological signals.
method Proposes a method to remove hand-engineered features and force DNNs to learn new representations.
result DNNs often rediscover known features, but can also learn new ones.

Good data stewardship requires removal of data at the request of the data's owner. This raises the question if and how a trained machine-learning model, which implicitly stores information about its training data, should be affected by such a removal request. Is it possible to "remove" data from a machine-learning mode…

2019-11-08abs ↗pdf ↗

AutoAnchor uses cross-attention to improve text-to-image model unlearning.

problem Mitigating harmful or copyrighted content in text-to-image models.
method Two-stage framework that automatically synthesizes manifold-proximal anchors using cross-attention consistency loss.
result Effective robust and unbiased unlearning across various baselines.

Selective removal of data subsets can efficiently unlearn unwanted distributions.

problem Efficiently removing unwanted data subsets without losing important information.
method Formalized as distributional unlearning, using Kullback-Leibler divergence constraints to select a small subset of data.
result Proposed method achieves corresponding log-loss bounds and is quadratically more sample-efficient than random removal.

A framework for hypothesis testing on attributed graphs using sampling.

problem Statistical testing on graph data, especially large attributed graphs.
method Sampling-based framework with PHASE and PHASEopt for accurate and efficient hypothesis testing.
result PHASE and PHASEopt improve accuracy and efficiency of hypothesis testing in attributed graphs.

Adversarial attacks and defenses are currently active areas of research for the deep learning community. A recent review paper divided the defense approaches into three categories; gradient masking, robust optimization, and adversarial example detection. We divide gradient masking and robust optimization differently: (…

2019-10-23abs ↗pdf ↗

The paper proves removable singularity for nonlocal minimal graphs.

problem Proving removable singularities for nonlocal minimal graphs.
method Analyzing (s,1)(s, 1)-capacity zero compact sets to ensure graphs are minimal in the entire domain.
result Nonlocal minimal graphs are removable in the entire domain if they are minimal in a set of (s,1)(s, 1)-capacity zero.

Removing spurious features can hurt model accuracy and disproportionately affect different groups.

problem Interference from spurious features in robust model performance across different groups.
method Characterization and analysis of spurious feature removal in noiseless overparameterized linear regression.
result Removal of spurious features can decrease accuracy and disproportionately affect different groups, even in balanced datasets.

We present a study of generalization for data-dependent hypothesis sets. We give a general learning guarantee for data-dependent hypothesis sets based on a notion of transductive Rademacher complexity. Our main result is a generalization bound for data-dependent hypothesis sets expressed in terms of a notion of hypothe…

2019-04-09abs ↗pdf ↗