Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

59118177236 · May 202619922001200920182026
48 results for identity matches

RLINK uses deep reinforcement learning to improve user identity linkage across social networks.

problem Recognizing the same user across different social networks.
method Converts user identity linkage into a sequence decision problem and uses deep reinforcement learning to optimize the linkage strategy.
result Achieves better performance than state-of-the-art methods in experiments on various datasets.

Lower bounds on private estimation of Gaussian covariance matrices.

problem Private estimation of Gaussian covariance matrices under various parameter regimes.
method Stein-Haff identity and fingerprinting lemma extensions.
result Lower bounds match existing upper bounds in the widest known parameters.

New private identity testers for high-dimensional distributions with improved sample complexity.

problem Testing goodness-of-fit for high-dimensional product distributions under differential privacy.
method Developed novel differentially private testers for multivariate product distributions, including Gaussians and binary product distributions.
result Achieved sample complexity matching the minimax sample complexity of O(d1/2/α2)O(d^{1/2}/α^2) in many parameter regimes.

Given a planar curve singularity, we prove a conjecture of Oblomkov-Shende, relating the geometry of its Hilbert scheme of points to the HOMFLY polynomial of the associated algebraic link. More generally, we prove an extension of this conjecture, due to Diaconescu-Hua-Soibelman, relating stable pair invariants on the c…

2012-10-23abs ↗pdf ↗

Study decomposes market portfolio into body and tail legs, revealing systematic differences.

problem Understanding the relationship between body and tail components in market portfolios.
method Decomposes CRSP market portfolio into body and tail legs, analyzes their recombination identity.
result Recombination identity holds for all models but not for all, indicating systematic differences.

New theory shows neural networks learn similar representations under certain conditions.

problem Understanding how different neural networks learn similar representations.
method Developed a theory based on neuron activation subspace match model, characterized maximum match and simple match.
result Representations learned by networks with identical architecture but different initializations are not as similar as previously thought.

The paper models market dynamics using a limit order book system to explain slippage and inefficiency.

problem Inefficiency in matching markets due to structural liquidity constraints and slippage.
method Introduces a market microstructure framework with a latent preference state matrix and a dynamic discrete choice execution model.
result Persistent slippage and regional invariance of preference orderings are explained by liquidity thresholds.

In high dimensions, the mean and geometric median are nearly identical.

problem Understanding the relationship between mean and geometric median in high-dimensional spaces.
method Analytical derivation and simulation of the distance between mean and geometric median.
result The distance between mean and geometric median vanishes with dimensionality in high dimensions.

Generative models improve CECT template matching reliability.

problem Insufficient template matching for accurate CECT structure assessment.
method Image-derived generative adversarial network for pseudo-macromolecular structures.
result Statistical credibility of CECT template matching significantly improved.

New method synchronizes partial permutations using non-negative factorizations.

problem Synchronizing partial multi-matchings in a cycle-consistent manner.
method Non-negative factorization approach with spectral relaxation and rotation scheme.
result Guaranteed cycle-consistent results compared to existing methods.

Matching correlated VAR time series databases by recovering matching permutations.

problem Matching perturbed and permuted correlated VAR time series.
method Probabilistic framework modeling, maximum likelihood estimator (MLE), linear assignment, convex relaxations.
result Recovery guarantees for perfect or partial recovery of matching permutations, thresholds for σσ.

Develops methods for estimating volatility models in high dimensions.

problem Estimating volatility in high-dimensional settings with heavy-tailed data.
method Uses Stein's identities for variance index estimation in high-dimensional settings.
result Matches minimax optimal rate for mean index estimation in high-dimensional settings.

ITF improves DSR but inflates curvature, while marginal likelihood reduces it, affecting QoIs.

problem Curvature mismatch between teacher forcing and marginal likelihood in chaotic dynamical systems.
method Comparing objective-induced curvatures of ITF and marginal likelihood in a probabilistic switching augmentation of AL-RNNs.
result Curvature inflation by ITF and reduction by marginal likelihood affect dynamical quantities of interest.

Unified framework for training diffusion and flow models to sample from target distributions.

problem Training diffusion and flow models to sample from target distributions defined by exponential tilting.
method Unified framework combining stochastic optimal control and non-equilibrium thermodynamics perspectives.
result Unified bias-variance decompositions and theoretical support for adjoint-based methods.

The paper shows how shared random seeds can reduce variance in machine learning evaluations.

problem The statistical structure of comparative evaluation under shared random seeds is not well understood.
method An extended learning-based multi-agent economic simulator was used to demonstrate the effects of shared random seeds on variance reduction.
result Pairing seeds can reduce variance in machine learning evaluations, especially when outcomes are positively correlated at the seed level.

New method debiases counterfactual distributions using observational data.

problem Estimating counterfactual distributions under interventions without relying on observational data.
method Flow-matching approach to learn counterfactual distributions from observational data.
result Deconfounding flows outperform existing debiased counterfactual distribution estimators.

A new method handles mismatched data in multivariate regression.

problem Handling mismatched data in multivariate linear regression.
method Two-stage approach: first stage estimates parameters, second stage estimates permutation.
result Permutation recovery conditions become less stringent with increasing number of responses.

Stein Variational Gradient Descent optimizes particle sets to match distribution expectations.

problem Efficiently approximating complex distributions in machine learning.
method Evolve particle sets to match the expectations of a given distribution using Stein operators and kernels.
result Particles can be used to exactly estimate expectations of functions on distributions, providing insights into kernel choice.

Score matching errors are not sufficient for measuring diffusion model quality.

problem The L2L^2 score matching error is not a reliable measure of diffusion model performance.
method Decomposed score errors into gradient and solenoidal components and analyzed their geometric properties.
result Only the gradient component of the score error affects the marginal distributional quality.

Unsupervised ensemble classification for dependent data.

problem Classifying data with dependencies using multiple classifiers.
method Developed algorithms for sequential and networked data dependencies, using moment matching and Expectation Maximization.
result Improved classification performance on synthetic and real datasets.

Efficiently distills pretrained text-to-image models without real data, improving FID and CLIP scores.

problem Slow iterative refinement process of diffusion-based text-to-image models.
method Guided Score identity Distillation with Long and Short Classifier-Free Guidance.
result Achieves state-of-the-art FID performance with competitive CLIP score.

DynBRO learns robustly from dynamic Byzantine workers.

problem Fault-tolerant distributed learning with dynamic Byzantine workers.
method Multi-level Monte Carlo (MLMC) gradient estimation and adaptive learning rate.
result DynaBRO nearly matches static setting's convergence rate with O(T)\mathcal{O}(\sqrt{T}) Byzantine worker changes.

Study the connection between supersymmetry and geometric flows in supergravity.

problem Relate supersymmetry to geometric flows in supergravity.
method Derive flow equations from a functional of squares of supersymmetry operators, match with mathematics anomaly flow, generalize to higher dimensions.
result Flow equations match known mathematics anomaly flow and simplify to scalar equations on torus fibrations.

The study examines property testing and estimation under non-identically distributed samples, finding necessary and sufficient sample complexities.

problem Property testing and estimation under non-identically distributed samples.
method Analysis of distributional property testing and estimation in settings with heterogeneous entities.
result Necessary and sufficient sample complexities for property testing and estimation under non-identically distributed samples.

New geometric analysis shows L2L^2 score error is flawed for diffusion models.

problem Score matching errors in diffusion models do not fully capture distributional quality.
method Decomposed score errors into gradient and solenoidal components, focusing on gradient's role in Fokker-Planck dynamics.
result Only gradient component affects marginal distributional quality; solenoidal component is structurally invisible.

EC method calibrates neural networks by matching average confidence to correct label proportion.

problem Overoptimism in neural network prediction confidence.
method Expectation consistency (EC) post-training rescaling of weights.
result EC achieves similar calibration performance to temperature scaling (TS) but is based on a principled Bayesian principle.

A new method uses a product of experts with Dirichlet variables to approximate complex distributions.

problem Approximating complex distributions with tractable models.
method A product of experts with auxiliary Dirichlet variables, using a Feynman identity to sample and optimize.
result The method efficiently approximates complex distributions using a product of experts and Dirichlet variables.