KDM unifies feature matching in neural networks.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
FMI uses matching to mimic interventions for causal feature learning.
We introduce new families of Integral Probability Metrics (IPM) for training Generative Adversarial Networks (GAN). Our IPMs are based on matching statistics of distributions embedded in a finite dimensional feature space. Mean and covariance feature matching IPMs allow for stable training of GANs, which we will call M…
Paper quantifies label shift with robustness guarantees using distribution feature matching.
Proposes DWMD for better matching of hidden representations across domains.
Improves deep learning training by matching mini-batch distributions.
Domain adaptation is an important technique to alleviate performance degradation caused by domain shift, e.g., when training and test data come from different domains. Most existing deep adaptation methods focus on reducing domain shift by matching marginal feature distributions through deep transformations on the inpu…
L2M learns to match distributions for domain adaptation without relying on hand-crafted priors.
Paper tackles distribution matching by partially matching distributions, achieving robust results.
Unsupervised domain adaptation methods aim to alleviate performance degradation caused by domain-shift by learning domain-invariant representations. Existing deep domain adaptation methods focus on holistic feature alignment by matching source and target holistic feature distributions, without considering local feature…
Feature Quantization improves GAN training stability.
New algorithm guarantees domain generalization with few environments.
A faster method for density estimation using denoising score matching with random Fourier features.
A new method for adapting to label shifts using class probability matching.
A new model combines diffusion and random features for better interpretability and comparable performance.
New method learns disentangled representations using Gromov-Monge maps.
A broad range of cross--domain generation researches boil down to matching a joint distribution by deep generative models (DGMs). Hitherto algorithms excel in pairwise domains while as increases, remain struggling to scale themselves to fit a joint distribution. In this paper, we propose a domain-scalable DGM, i…
Centroids Matching tackles catastrophic forgetting by matching feature vectors to class centroids.
HTFM improves mode coverage and tail-statistic recovery for heavy-tailed data.
A new method generates mixed-type features in tabular data with improved realism and accuracy.
In this work, we formulate the fixed-length distribution matching as a Bayesian inference problem. Our proposed solution is inspired from the compressed sensing paradigm and the sparse superposition (SS) codes. First, we introduce sparsity in the binary source via position modulation (PM). We then present a simple and …
We consider a covariate shift problem where one has access to several different training datasets for the same learning problem and a small validation set which possibly differs from all the individual training distributions. This covariate shift is caused, in part, due to unobserved features in the datasets. The objec…
Unified framework for distribution shift estimation, explanation, and improvement.
New method debiases counterfactual distributions using observational data.
We study graph matching with correlated Gaussian features and find thresholds for exact recovery.
GANs mode collapse solved with Bures distance.
The paper matches features in images using centro-affine invariants and heat flow.
Stein variational gradient descent (SVGD) is a non-parametric inference algorithm that evolves a set of particles to fit a given distribution of interest. We analyze the non-asymptotic properties of SVGD, showing that there exists a set of functions, which we call the Stein matching set, whose expectations are exactly …
Domain adaptation aims to leverage the supervision signal of source domain to obtain an accurate model for target domain, where the labels are not available. To leverage and adapt the label information from source domain, most existing methods employ a feature extracting function and match the marginal distributions of…
The Generative Adversarial Network (GAN) has achieved great success in generating realistic (real-valued) synthetic data. However, convergence issues and difficulties dealing with discrete data hinder the applicability of GAN to text. We propose a framework for generating realistic text via adversarial training. We emp…
Deep neural networks, trained with large amount of labeled data, can fail to generalize well when tested with examples from a \emph{target domain} whose distribution differs from the training data distribution, referred as the \emph{source domain}. It can be expensive or even infeasible to obtain required amount of lab…
Study connects database alignment and planted matching using Gaussian features.
We show that kernel-based quadrature rules for computing integrals can be seen as a special case of random feature expansions for positive definite kernels, for a particular decomposition that always exists for such kernels. We provide a theoretical analysis of the number of required samples for a given approximation e…
New models for analyzing microbiome data with interactions.
The study examines how the number of noise samples affects diffusion models' performance.
Method matches noisy remote sensing images robustly.
This paper focuses on the problem of query by example spoken term detection (QbE-STD) in zero-resource scenario. State-of-the-art approaches primarily rely on dynamic time warping (DTW) based template matching techniques using phone posterior or bottleneck features extracted from a deep neural network (DNN). We use bot…
Correctly detecting the semantic type of data columns is crucial for data science tasks such as automated data cleaning, schema matching, and data discovery. Existing data preparation and analysis systems rely on dictionary lookups and regular expression matching to detect semantic types. However, these matching-based …
Posterior Matching enables VAEs to model arbitrary conditional densities.
Paper uses Gaussian mixture models and Wasserstein distance for schema matching.
Paper learns data-driven organ matching rules from observational data.
Map matching of GPS trajectories from a sequence of noisy observations serves the purpose of recovering the original routes in a road network. In this work in progress, we attempt to share our experience of feature construction in a spatial database by reporting our ongoing experiment of feature extrac-tion in Conditio…
Study shows how diffusion models learn on low-dimensional manifolds.
When pre-processing observational data via matching, we seek to approximate each unit with maximally similar peers that had an alternative treatment status--essentially replicating a randomized block design. However, as one considers a growing number of continuous features, a curse of dimensionality applies making asym…
Proposes CDTD, a diffusion model for mixed-type tabular data.
Causal inference often relies on the counterfactual framework, which requires that treatment assignment is independent of the outcome, known as strong ignorability. Approaches to enforcing strong ignorability in causal analyses of observational data include weighting and matching methods. Effect estimates, such as the …
The learning of domain-invariant representations in the context of domain adaptation with neural networks is considered. We propose a new regularization method that minimizes the discrepancy between domain-specific latent feature representations directly in the hidden activation space. Although some standard distributi…
Proposes a non-adversarial method for distribution matching.