Bayesian algorithm detects image matches and fraud.
problem Detecting identity matches and fraud in image databases.
method Generative model of image graph trained with matching algorithm.
result Bayesian approach improves detection accuracy.
RLINK uses deep reinforcement learning to improve user identity linkage across social networks.
problem Recognizing the same user across different social networks.
method Converts user identity linkage into a sequence decision problem and uses deep reinforcement learning to optimize the linkage strategy.
result Achieves better performance than state-of-the-art methods in experiments on various datasets.
Improves score estimation for noised targets using known clean scores.
problem Poor score estimation at low noise levels in Denoising Score Matching.
method Introduces Target Score Identity and Target Score Matching loss.
result Score estimates are more accurate at low noise levels.
Lower bounds on private estimation of Gaussian covariance matrices.
problem Private estimation of Gaussian covariance matrices under various parameter regimes.
method Stein-Haff identity and fingerprinting lemma extensions.
result Lower bounds match existing upper bounds in the widest known parameters.
Procedure tests if unknown Markov chain matches a reference chain.
problem Testing if an unknown Markov chain matches a reference chain.
method An efficient procedure based on a single long state sequence.
result Nearly matching upper and lower sample complexity bounds for total variation distance.
A new filter reduces density fitting to a linear solve, improving performance on nonlinear systems.
problem Nonlinear Bayesian filtering challenges in representing belief distributions.
method Combines score matching with Stein's identity to avoid partition function evaluation.
result The Score Kalman Filter (SKF) outperforms existing methods on nonlinear systems.
A method for matching vertices in large networks using seeds.
problem Matching vertices in large, overlapping networks.
method Identify seeds in local neighborhoods, match induced subgraphs, rank matches.
result Principled approach for large networks, demonstrated through simulations and real data.
New private identity testers for high-dimensional distributions with improved sample complexity.
problem Testing goodness-of-fit for high-dimensional product distributions under differential privacy.
method Developed novel differentially private testers for multivariate product distributions, including Gaussians and binary product distributions.
result Achieved sample complexity matching the minimax sample complexity of O ( d 1 / 2 / α 2 ) O(d^{1/2}/α^2) O ( d 1/2 / α 2 ) in many parameter regimes. Market portfolio decomposed into body and tail legs
problem Separation of market portfolio into body and tail legs
method Dynamic value-weighted body and tail legs
result Recombination identity holds for all models
Given a planar curve singularity, we prove a conjecture of Oblomkov-Shende, relating the geometry of its Hilbert scheme of points to the HOMFLY polynomial of the associated algebraic link. More generally, we prove an extension of this conjecture, due to Diaconescu-Hua-Soibelman, relating stable pair invariants on the c…
Study decomposes market portfolio into body and tail legs, revealing systematic differences.
problem Understanding the relationship between body and tail components in market portfolios.
method Decomposes CRSP market portfolio into body and tail legs, analyzes their recombination identity.
result Recombination identity holds for all models but not for all, indicating systematic differences.
New obstructions show some 4-manifold homeomorphisms are pseudo-isotopic but not isotopic.
problem Tackling the difference between pseudo-isotopy and isotopy for 4-manifolds.
method Defining and showing realizable obstructions in Whitehead groups.
result Found homeomorphisms that are pseudo-isotopic but not isotopic.
New theory shows neural networks learn similar representations under certain conditions.
problem Understanding how different neural networks learn similar representations.
method Developed a theory based on neuron activation subspace match model, characterized maximum match and simple match.
result Representations learned by networks with identical architecture but different initializations are not as similar as previously thought.
The paper models market dynamics using a limit order book system to explain slippage and inefficiency.
problem Inefficiency in matching markets due to structural liquidity constraints and slippage.
method Introduces a market microstructure framework with a latent preference state matrix and a dynamic discrete choice execution model.
result Persistent slippage and regional invariance of preference orderings are explained by liquidity thresholds.
Algorithm generates diverse images of the same subject while maintaining specific aspects.
problem Training GANs to generate realistic, identity-matched images.
method Pairwise training scheme with Siamese discriminators.
result Algorithm produces convincing, identity-matched photographs.
In high dimensions, the mean and geometric median are nearly identical.
problem Understanding the relationship between mean and geometric median in high-dimensional spaces.
method Analytical derivation and simulation of the distance between mean and geometric median.
result The distance between mean and geometric median vanishes with dimensionality in high dimensions.
Generative models improve CECT template matching reliability.
problem Insufficient template matching for accurate CECT structure assessment.
method Image-derived generative adversarial network for pseudo-macromolecular structures.
result Statistical credibility of CECT template matching significantly improved.
New method synchronizes partial permutations using non-negative factorizations.
problem Synchronizing partial multi-matchings in a cycle-consistent manner.
method Non-negative factorization approach with spectral relaxation and rotation scheme.
result Guaranteed cycle-consistent results compared to existing methods.
New algorithm speeds up causal discovery for network data.
problem Scalability issues in score-matching for temporal network data.
method Developed a new parent-finding subroutine for DAGs, improving score matching efficiency.
result Efficiency-lifted score matching for both i.i.d. and temporal data on networks.
BNEM improves Boltzmann sampler efficiency.
problem Generating IID samples from Boltzmann distributions efficiently.
method Bootstrapped Noised Energy Matching (NEM) combined with diffusion-based learning and bootstrapping.
result BNEM achieves state-of-the-art performance with improved robustness.
Matching correlated VAR time series databases by recovering matching permutations.
problem Matching perturbed and permuted correlated VAR time series.
method Probabilistic framework modeling, maximum likelihood estimator (MLE), linear assignment, convex relaxations.
result Recovery guarantees for perfect or partial recovery of matching permutations, thresholds for σ σ σ . IVON optimizes large neural networks, matching or outperforming Adam.
problem The inefficacy of variational learning in large neural networks.
method Improved Variational Online Newton (IVON) optimizer.
result IVON consistently matches or outperforms Adam for large networks.
Develops methods for estimating volatility models in high dimensions.
problem Estimating volatility in high-dimensional settings with heavy-tailed data.
method Uses Stein's identities for variance index estimation in high-dimensional settings.
result Matches minimax optimal rate for mean index estimation in high-dimensional settings.
New method improves SBI efficiency and scalability.
problem Scalability issues in SBI methods for large datasets.
method Langevin dynamics with score matching, exploiting likelihood structure.
result Structured score network enhances statistical efficiency and scalability.
ITF improves DSR but inflates curvature, while marginal likelihood reduces it, affecting QoIs.
problem Curvature mismatch between teacher forcing and marginal likelihood in chaotic dynamical systems.
method Comparing objective-induced curvatures of ITF and marginal likelihood in a probabilistic switching augmentation of AL-RNNs.
result Curvature inflation by ITF and reduction by marginal likelihood affect dynamical quantities of interest.
Neural network converts speech from one language to multiple languages.
problem Creating speech in multiple languages without parallel data.
method Polyglot neural network with multiple per-language sub-networks and loss terms to preserve speaker identity.
result Convincing conversion capabilities demonstrated across three languages.
Unified framework for training diffusion and flow models to sample from target distributions.
problem Training diffusion and flow models to sample from target distributions defined by exponential tilting.
method Unified framework combining stochastic optimal control and non-equilibrium thermodynamics perspectives.
result Unified bias-variance decompositions and theoretical support for adjoint-based methods.
The paper shows how shared random seeds can reduce variance in machine learning evaluations.
problem The statistical structure of comparative evaluation under shared random seeds is not well understood.
method An extended learning-based multi-agent economic simulator was used to demonstrate the effects of shared random seeds on variance reduction.
result Pairing seeds can reduce variance in machine learning evaluations, especially when outcomes are positively correlated at the seed level.
DPA preserves data distribution in reduced dimensions.
problem Loss of data distribution in dimension reduction.
method DPA combines encoder and decoder to match data distribution.
result DPA successfully reconstructs data distribution.
UPM clusters product titles for efficient unsupervised matching.
problem Matching product titles accurately in e-commerce.
method Clustering-based approach using title combinations.
result UPM outperforms state-of-the-art methods in efficiency and effectiveness.
New method debiases counterfactual distributions using observational data.
problem Estimating counterfactual distributions under interventions without relying on observational data.
method Flow-matching approach to learn counterfactual distributions from observational data.
result Deconfounding flows outperform existing debiased counterfactual distribution estimators.
A new method handles mismatched data in multivariate regression.
problem Handling mismatched data in multivariate linear regression.
method Two-stage approach: first stage estimates parameters, second stage estimates permutation.
result Permutation recovery conditions become less stringent with increasing number of responses.
Stein Variational Gradient Descent optimizes particle sets to match distribution expectations.
problem Efficiently approximating complex distributions in machine learning.
method Evolve particle sets to match the expectations of a given distribution using Stein operators and kernels.
result Particles can be used to exactly estimate expectations of functions on distributions, providing insights into kernel choice.
Score matching errors are not sufficient for measuring diffusion model quality.
problem The L 2 L^2 L 2 score matching error is not a reliable measure of diffusion model performance. method Decomposed score errors into gradient and solenoidal components and analyzed their geometric properties.
result Only the gradient component of the score error affects the marginal distributional quality.
Authors propose GRI-CNN systems for rotationally invariant processing.
problem Creating rotationally invariant CNN systems.
method Designed geared rotationally identical CNN systems (GRI-CNN) with a small step angle.
result GRI-CNN produces quantitatively identical output results under rotation.
Unsupervised ensemble classification for dependent data.
problem Classifying data with dependencies using multiple classifiers.
method Developed algorithms for sequential and networked data dependencies, using moment matching and Expectation Maximization.
result Improved classification performance on synthetic and real datasets.
Neural network predicts electron-ionization mass spectra quickly.
problem Identifying unknown molecules not in existing libraries.
method Lightweight neural network model for predicting mass spectra.
result High accuracy predictions of small molecule mass spectra.
Efficiently distills pretrained text-to-image models without real data, improving FID and CLIP scores.
problem Slow iterative refinement process of diffusion-based text-to-image models.
method Guided Score identity Distillation with Long and Short Classifier-Free Guidance.
result Achieves state-of-the-art FID performance with competitive CLIP score.
DynBRO learns robustly from dynamic Byzantine workers.
problem Fault-tolerant distributed learning with dynamic Byzantine workers.
method Multi-level Monte Carlo (MLMC) gradient estimation and adaptive learning rate.
result DynaBRO nearly matches static setting's convergence rate with O ( T ) \mathcal{O}(\sqrt{T}) O ( T ) Byzantine worker changes. New theorem connects distant points and identical points on manifolds.
problem Continuous maps and distant points on manifolds.
method Qualitative extension of Hopf theorem, using topological 'distant' points.
result Existence of connected component containing both distant and identical points.
DSM on manifolds removes singularities and computes small-noise expansions.
problem DSM on manifolds with singular noise.
method Rao-Blackwellized score matching, nearest-point projection, intrinsic Riemannian score.
result Canonical target equals intrinsic Riemannian score up to a small correction.
Public trading wallets reveal price information not captured by anonymous data.
problem Informed traders' anonymity in public exchanges.
method Reconstructed full-depth limit order book from 17.1 billion messages.
result Wallets' aggressive orders predict returns, with a 13.2% gain over anonymous benchmarks.
Study the connection between supersymmetry and geometric flows in supergravity.
problem Relate supersymmetry to geometric flows in supergravity.
method Derive flow equations from a functional of squares of supersymmetry operators, match with mathematics anomaly flow, generalize to higher dimensions.
result Flow equations match known mathematics anomaly flow and simplify to scalar equations on torus fibrations.
The study examines property testing and estimation under non-identically distributed samples, finding necessary and sufficient sample complexities.
problem Property testing and estimation under non-identically distributed samples.
method Analysis of distributional property testing and estimation in settings with heterogeneous entities.
result Necessary and sufficient sample complexities for property testing and estimation under non-identically distributed samples.
New geometric analysis shows L 2 L^2 L 2 score error is flawed for diffusion models.
problem Score matching errors in diffusion models do not fully capture distributional quality.
method Decomposed score errors into gradient and solenoidal components, focusing on gradient's role in Fokker-Planck dynamics.
result Only gradient component affects marginal distributional quality; solenoidal component is structurally invisible.
We propose a general purpose variational inference algorithm that forms a natural counterpart of gradient descent for optimization. Our method iteratively transports a set of particles to match the target distribution, by applying a form of functional gradient descent that minimizes the KL divergence. Empirical studies…
EC method calibrates neural networks by matching average confidence to correct label proportion.
problem Overoptimism in neural network prediction confidence.
method Expectation consistency (EC) post-training rescaling of weights.
result EC achieves similar calibration performance to temperature scaling (TS) but is based on a principled Bayesian principle.
A new method uses a product of experts with Dirichlet variables to approximate complex distributions.
problem Approximating complex distributions with tractable models.
method A product of experts with auxiliary Dirichlet variables, using a Feynman identity to sample and optimize.
result The method efficiently approximates complex distributions using a product of experts and Dirichlet variables.