Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

3468101135 · Jun 202019922001200920182026
48 results for paired hypotheses discrepancy

Proposes PHD to measure domain discrepancy for complex models.

problem Insufficient domain discrepancy measures for complex models.
method Introduces PHD, a novel discrepancy measure for complex models.
result PHD is computationally efficient and applicable to multi-class classification.

Do two data samples come from different distributions? Recent studies of this fundamental problem focused on embedding probability distributions into sufficiently rich characteristic Reproducing Kernel Hilbert Spaces (RKHSs), to compare distributions by the distance between their embeddings. We show that Regularized Ma…

2013-05-02abs ↗pdf ↗

This paper introduces localized discrepancy theories for unsupervised domain adaptation.

problem Improving generalization bounds for unsupervised domain adaptation.
method Localized discrepancies defined on the hypothesis space after localization, leading to smaller and asymmetric values.
result Improved generalization bounds and sample complexity reduction.

Unified method for MMD variance estimation improves accuracy and computational efficiency.

problem Variance estimation for MMD in nonparametric testing.
method Unified finite-sample characterization of MMD variance through U-statistic and Hoeffding decomposition; exact acceleration method for univariate case.
result Unified estimators improve accuracy and computational efficiency for MMD variance.

Boosting combines weak hypotheses to create accurate predictions under bounded VC dimension.

problem How to combine weak hypotheses to achieve accurate predictions efficiently.
method Designing a novel boosting algorithm with complex aggregation rules for bounded VC dimension classes.
result The new boosting algorithm requires fewer weak hypotheses than classical lower bounds, provided they belong to a bounded VC class.

A new measure scales MMD to assess distribution closeness.

problem Testing statistical significance of distribution closeness.
method Norm-adaptive MMD (NAMMD) for distributional discrepancy.
result NAMMD-based DCT has higher test power than MMD-based DCT.

We present a new analysis of the problem of learning with drifting distributions in the batch setting using the notion of discrepancy. We prove learning bounds based on the Rademacher complexity of the hypothesis set and the discrepancy of distributions both for a drifting PAC scenario and a tracking scenario. Our boun…

2012-05-19abs ↗pdf ↗

New MMD estimators detect differences in missing paired data.

problem Handling missing data in matched pairs with complex distributions.
method Maximum mean discrepancy (MMD) estimators for complex data with missing values.
result Valid and consistent estimators detect differences in data distributions.

More powerful feature selection tests using selective inference.

problem Selection bias in feature selection leading to specious analysis.
method Conditioning on minimal selection event using Maximum Mean Discrepancy and Hilbert Schmidt Independence Criterion with multiscale bootstrap.
result Proposed test is more powerful in most scenarios.

This work tackles domain generalization by minimizing discrepancy between domains.

problem Domain generalization: learning to handle unseen domains with i.i.d. data assumptions violated.
method The approach involves minimizing discrepancy between domains using a lemma and deriving a generalization bound.
result Low risk over unseen domains can be achieved by representing data in a space where training distributions are indistinguishable and relevant information is preserved.

Study finds exposure bias distortion is limited and not incremental in open-ended text generation.

problem Exposure bias in auto-regressive language models causing incremental distortion.
method Proposed metrics to quantify exposure bias impact, used ground-truth prefixes instead of model-generated prefixes.
result Exposure bias distortion is limited and not incremental during generation.

Paper introduces a method for supervised hierarchical clustering with Exponential Linkage.

problem Discrepancy between training and clustering objectives in supervised clustering.
method Tightly couples supervised training of dissimilarity function with hierarchical clustering, using Exponential Linkage.
result Joint training procedure consistently matches or outperforms other methods, improving dendrogram purity by up to 8 points.

We consider Lagrangian-like submanifolds in certain even-dimensional 'symplectic-like' Poisson manifolds. We show, under suitable transversality hypotheses, that the pair consisting of the ambient Poisson manifold and the submanifold has unobstructed deformations and that the deformations automatically preserve the Lag…

2013-11-12abs ↗pdf ↗

We generalize the notions of dual pair and polarity introduced by S. Lie and A. Weinstein in order to accommodate very relevant situations where the application of these ideas is desirable. The new notion of polarity is designed to deal with the loss of smoothness caused by the presence of singularities that are encoun…

2002-01-21abs ↗pdf ↗

Study assesses market simulation metrics to highlight discrepancies between real and simulated markets.

problem Lack of fidelity in market simulation methods leads to discrepancies between real and simulated market data.
method Surveyed and applied a set of reference metrics to real and simulated market data.
result Significant discrepancies remain between real and simulated markets.

Robust hypothesis testing designs a test for worst-case distributions using kernel methods.

problem Design a robust test for hypothesis testing under uncertainty sets.
method Data-driven uncertainty sets constructed using kernel mean embeddings and maximum mean discrepancy (MMD). Bayesian and Neyman-Pearson settings investigated.
result Proposed robust kernel tests are exponentially consistent and asymptotically optimal.

Novel upper bound for unsupervised domain adaptation considers joint error.

problem Addressing the issue of mixing samples from different classes when matching marginal distributions.
method Proposes a general upper bound that penalizes undesirable joint error, uses constrained hypothesis space, and introduces cross margin discrepancy.
result Our proposal outperforms related approaches in image classification error rates on domain adaptation benchmarks.

DPC uses physics and neural nets to solve SDEs.

problem Solving stochastic differential equations with missing physics.
method Physics-data fusion with conditional maximum mean discrepancy (CMMD) loss.
result DPC achieves highly accurate solutions on benchmark examples.

Proposes a semi-Bayesian nonparametric estimator for MMD in GOF tests and GANs.

problem Challenges in goodness-of-fit testing for intractable models.
method Semi-Bayesian nonparametric estimator of MMD.
result Outperforms frequentist MMD-based methods in false rejection and acceptance rates.

We consider Lagrangian Floer cohomology for a pair of Lagrangian submanifolds in a symplectic manifold M. Suppose that M carries a symplectic involution, which preserves both submanifolds. Under various topological hypotheses, we prove a localization theorem for Floer cohomology, which implies a Smith-type inequality f…

2010-02-12abs ↗pdf ↗

This paper introduces an information theoretic co-training objective for unsupervised learning. We consider the problem of predicting the future. Rather than predict future sensations (image pixels or sound waves) we predict "hypotheses" to be confirmed by future sensations. More formally, we assume a population distri…

2018-02-21abs ↗pdf ↗

Almost-isometries are quasi-isometries with multiplicative constant one. Lifting a pair of metrics on a compact space gives quasi-isometric metrics on the universal cover. Under some additional hypotheses on the metrics, we show that there is no almost-isometry between the universal covers. We show that Riemannian mani…

2014-09-10abs ↗pdf ↗

Approximate Markov chain Monte Carlo (MCMC) offers the promise of more rapid sampling at the cost of more biased inference. Since standard MCMC diagnostics fail to detect these biases, researchers have developed computable Stein discrepancy measures that provably determine the convergence of a sample to its target dist…

2017-03-06abs ↗pdf ↗

Generative model for TPPs using signatures and distributional discrepancies.

problem Limitations of signature methods for TPPs and lack of global sequence-level loss in neural models.
method Introduce interarrival embedding to lift jump paths to continuous paths of bounded variation, enabling signature methods for discrete event sequences. Develop sigTPP, a signature-based generative model trained on path-level loss.
result sigTPP achieves the best average rank across multiple metrics and outperforms or is within a standard error of the strongest baseline in 64% of dataset-metric pairs.

Computer Vision and machine learning methods were previously used to reveal screen presence of genders in TV and movies. In this work, using head pose, gender detection, and skin color estimation techniques, we demonstrate that the gender disparity in TV in a South Asian country such as Bangladesh exhibits unique chara…

2017-11-14abs ↗pdf ↗

CFR-Pro enhances treatment effect estimation by incorporating local proximity.

problem Treatment selection bias in HTE estimation from observational data.
method Proximity-enhanced CounterFactual Regression (CFR-Pro) with pair-wise proximity regularizer and subspace projector.
result Significantly outperforms competitors in HTE estimation accuracy.

It was proved by Mineyev and Yaman that, if (Γ,Γ)(Γ, Γ') is a relatively hyperbolic pair, the comparison map Hbk(Γ,Γ;V)Hk(Γ,Γ;V) H_b^k(Γ, Γ'; V) \to H^k(Γ, Γ'; V) is surjective for every k2k \ge 2, and any bounded ΓΓ--module VV. By exploiting results of Groves and Manning, we give another proof of this result. Moreover, we prove the …

2015-05-17abs ↗pdf ↗

A new metric compares true and learned causal graphs considering data and graph structure.

problem Comparing true and learned causal graphs accurately.
method Continuous Structural Intervention Distance (CSID) using conditional mean embeddings and maximum mean discrepancy.
result Validated the CSID with synthetic data, showing its effectiveness in comparing causal graphs.

The possibilities for new or unusual kinds of topological, locally linear periodic maps of non-prime order on closed, simply connected 4-manifolds with positive definite intersection pairings are explored. On the one hand, certain permutation representations on homology are ruled out under appropriate hypotheses. On th…

2002-05-10abs ↗pdf ↗

Unified framework for imitating tasks across domains with discrepancies.

problem Learning tasks across domains with embodiment, viewpoint, and dynamics mismatches.
method Two-step approach: alignment followed by adaptation. Alignment uses Generative Adversarial MDP Alignment (GAMA) for state and action correspondences from unpaired, unaligned demonstrations. Adaptation leverages these correspondences for zero-shot imitation.
result Effectiveness of the proposed approach in embodiment, viewpoint, and dynamics mismatch scenarios.

New discrepancy function compares discrete probability measures considering space geometry.

problem Comparing discrete probability measures in a geometrically meaningful way.
method Proposes the Fourier Discrepancy Function, proving convexity, differentiability, and providing gradient formula.
result Proves the Fourier Discrepancy is convex, twice differentiable, and provides an explicit gradient formula.

A new approach to learning in brain-like networks using adversarial algorithms.

problem Complex inter-dependencies in brain-like networks not compatible with conditional independence assumptions.
method Adversarial algorithm for learning models of perceptual processing.
result The approach can mimic known neural phenomena and yields testable hypotheses.

W2S FT often outperforms weak teachers due to low intrinsic dimensionality.

problem Understanding why weak-to-strong finetuning outperforms weak models.
method Analyzing W2S in ridgeless regression setting, focusing on variance reduction.
result Weak teacher's variance is inherited by strong student in shared feature subspace, reduced in discrepancy subspace.

Novel MBRL method for large-scale RL with reduced posterior complexity.

problem Theoretical guarantees for MBRL in large spaces with complex models.
method Kernelized Stein Discrepancy for compression of posterior estimate.
result Sublinear Bayesian regret and up to 50% reduction in training time.

The article introduces practical estimators for kernel discrepancies.

problem Estimating kernel discrepancies accurately and efficiently.
method Presented various estimators for MMD, HSIC, and KSD, including V-statistics, U-statistics, and incomplete U-statistics. Stressed the importance of kernel bandwidth and introduced adaptive estimators.
result Adaptive estimators combining multiple estimators with various kernels address the problem of kernel selection.

MASC balances dataset representation using affinity clustering and distribution discrepancies.

problem Representation bias in datasets due to group imbalance.
method MASC uses affinity clustering and pairwise distribution discrepancies to balance non-protected and protected groups.
result MASC effectively debiases target datasets, comparable to existing methods.

Paper defines class discrepancy for machine learning problems and provides coresets.

problem Addressing discrepancies in machine learning models.
method Defines class discrepancy, provides techniques for bounding discrepancy, and develops coresets and streaming sketches.
result Establishes coresets of size O(sqrt{d}/epsilon) for various machine learning problems.

Paper develops a consistent model selection framework for learning Hypotheses Space from data.

problem Avoiding overfitting in complex spaces with limited data.
method Develops a model selection framework based on Learning Spaces, selecting a Hypotheses Space from data.
result The method converges with probability one to a target Hypotheses Space, providing a consistent framework for model selection.

A new method for comparing measures in different spaces using flow alignment.

problem Comparing probability measures in different metric spaces with computational efficiency.
method Flow-based Alignment (\FlowAlign) and Depth-based Alignment (\DepthAlign) using tree structures.
result Flow-based Alignment and Depth-based Alignment are pseudo-distances and scalable for large-scale applications.