Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

2815618421,122 · Jun 202019922001200920182026
48 results for data mismatch

Paper shows using unlabeled mismatched images can improve KD for image classification.

problem Improving Knowledge Distillation for image classification with limited labeled data.
method Used unlabeled mismatched images as stimulus for KD, focusing on stimulus complexity.
result Stimulus complexity is crucial for KD's effectiveness, as demonstrated on MNIST and CIFAR datasets.

Paper addresses linear regression with partially mismatched data using local search with theoretical guarantees.

problem Linear regression with partially mismatched data.
method Optimization formulation and greedy local search algorithm with theoretical guarantees.
result Local search algorithm converges to nearly-optimal solution at a linear rate under certain conditions.

Paper proposes a method to handle linear regression with partially shuffled data.

problem Linear regression with mismatched predictors and responses.
method Pseudo-likelihood approach based on two-component mixture densities with EM optimization.
result The method can tolerate larger fractions of mismatches and estimate noise level.

Proposes a method to improve probabilistic models by reweighting data to correct for model assumptions.

problem Data mismatches between model assumptions and reality undermine probabilistic model inference and prediction quality.
method Bayesian data reweighting to identify and down-weight observations that do not match model assumptions.
result Improves predictive accuracy and robustness of probabilistic models through systematic detection and mitigation of data mismatches.

A new method handles mismatched data in multivariate regression.

problem Handling mismatched data in multivariate linear regression.
method Two-stage approach: first stage estimates parameters, second stage estimates permutation.
result Permutation recovery conditions become less stringent with increasing number of responses.

The paper addresses score-mismatched diffusion models and zero-shot conditional samplers.

problem Theoretical guarantees for score-mismatched diffusion models in zero-shot conditional sampling.
method Theoretical analysis of score-mismatched diffusion models and zero-shot conditional samplers.
result Theoretical performance guarantees with explicit dimensional dependencies for score-mismatched diffusion samplers.

Paper identifies objective mismatch in MBRL, affecting control task performance.

problem Objective mismatch in MBRL framework affects control task performance.
method Proposes re-weighting dynamics model training to mitigate mismatch.
result Likelihood of one-step ahead predictions is not always correlated with control performance.

Deep learning corrects GRACE TWSA mismatch in NOAH models.

problem Improving hydrological model predictive performance with GRACE data.
method Developed and applied deep convolutional neural network (CNN) models to learn and correct TWSA mismatch.
result Significant improvement in correlation coefficient and Nash-Sutcliff efficiency over original NOAH TWSA.

Deep learning predicts mismatched ratings in Amazon reviews.

problem Identifying reviews with mismatched ratings on Amazon.
method Converted reviews to vectors using paragraph vector, trained a recurrent neural network with gated recurrent unit, incorporated semantic relationships.
result Model accurately predicts rating mismatches and provides feedback.

New algorithm reduces performance loss in IRL with mismatched transition dynamics.

problem Performance degradation in inverse reinforcement learning due to mismatched transition dynamics.
method Proposed a robust Maximum Causal Entropy (MCE) IRL algorithm leveraging robust reinforcement learning insights.
result Empirically demonstrated stable performance improvement under transition dynamics mismatches.

A new method improves data generation quality by correcting score mismatches.

problem Score mismatch issue in conditional score-based data generation methods.
method Denoising Likelihood Score Matching (DLSM) loss for classifier training.
result The proposed method outperforms previous methods on Cifar-10 and Cifar-100 benchmarks.

New method reduces state distribution mismatch in off-policy RL.

problem State distribution mismatch in off-policy RL algorithms.
method Develops a novel constrained off-policy gradient objective to minimize state distribution shift.
result Minimizing state distribution shift improves performance in off-policy RL algorithms.

Study addresses covariate mismatch in federated learning, improving model accuracy.

problem Learning from clients with different feature sets in federated learning.
method Developed two approaches for linear prediction under covariate mismatch: plug-in estimator and impute-then-regress strategy.
result Proposed methods provide asymptotic and finite-sample learning rates, improving model accuracy.

A new BN method corrects size mismatch in WLF for imbalanced data.

problem Learning from imbalanced datasets in neural networks.
method Proposes weighted batch normalization (WBN) to correct size mismatch between BN and WLF.
result WBN corrects size mismatch and improves classification performance in imbalanced datasets.

MixMOOD improves SSDL by selecting unlabelled data based on deep feature similarity.

problem Class distribution mismatch in semi-supervised learning.
method MixMOOD uses deep dataset dissimilarity measures to select unlabelled data.
result MixMOOD selects unlabelled data based on strong correlation with MixMatch accuracy.

KalmanNet uses neural networks to improve state estimation in systems with unknown dynamics.

problem State estimation of systems with non-linear dynamics and partial information.
method KalmanNet integrates a recurrent neural network with the Kalman filter to handle non-linearities and model mismatches.
result KalmanNet outperforms classic filtering methods in systems with both mismatched and accurate domain knowledge.

BinaryDuo improves BNNs by coupling binary activations, outperforming state-of-the-art models.

problem Gradient mismatch in BNNs due to binarizing activations.
method Using gradient of smoothed loss function to estimate gradient mismatch, proposing BinaryDuo scheme with coupled ternary activations.
result BinaryDuo outperforms state-of-the-art BNNs on various benchmarks.

Proposes a generative model using scaled-Bregman divergences to handle support mismatch in training.

problem Support mismatch between model and data distributions during training.
method Augments the base measure of the problematic divergence (scaled-Bregman) to resolve the support mismatch problem.
result Demonstrates promising results on MNIST, CelebA, and CIFAR-10 datasets.

Thompson Sampling shows polynomial regret for combinatorial semi-bandits with subgaussian rewards.

problem Finding optimal solutions in combinatorial semi-bandits with suboptimal sampling.
method Proposes Thompson Sampling with polynomial regret for linear combinatorial semi-bandits.
result Demonstrates 'mismatched sampling paradox' where knowing distributions can lead to worse performance.

New research shows no trade-off between fairness and accuracy in machine learning.

problem The trade-off between fairness and accuracy in machine learning is a widely accepted belief.
method Using mismatched hypothesis testing and Chernoff information, the study demonstrates that optimal fairness and accuracy can be achieved simultaneously.
result There is no inherent trade-off between fairness and accuracy in ideal distributions, but it exists when measured with respect to biased datasets.

New method improves spatial prediction validation accuracy.

problem Validation methods fail for spatial prediction tasks due to mismatch between validation and test locations.
method Proposes a new validation method that adapts existing covariate-shift ideas to spatial settings.
result Proves and demonstrates the new method's superiority in spatial prediction validation.

PACMAN provides bounds for classification tasks considering accuracy vs. negative log-loss mismatch.

problem Mismatch between accuracy and negative log-loss in classification tasks.
method Point-wise PAC approach over generalization gap, using likelihood ratio and concentration inequalities.
result PACMAN provides point-wise PAC bounds for the generalization problem.

Learning reward functions can lead to poor policy performance despite low error.

problem Low error in learned reward functions does not guarantee low regret in policy performance.
method Mathematical analysis of reward learning and policy optimization.
result A low expected test error of the reward model guarantees low worst-case regret, but error-regret mismatch can occur with certain data distributions.

Voice conversion methods degrade in noisy conditions, with BLFWAS outperforming others.

problem Robustness of voice conversion techniques under mismatched conditions.
method Comparative analysis of five VC techniques on CMU ARCTIC corpus, exploring speech enhancement techniques.
result Bilinear frequency warping with amplitude scaling (BLFWAS) outperforms other methods in noisy conditions.

A new method for inventory control using in-context learning and generative models.

problem Inventory control with decision-dependent censoring, focusing on the censored newsvendor problem.
method In-context generative posterior sampling (ICGPS) combining modern generative models and in-context autoregressive generation.
result ICGPS achieves sublinear Bayesian regret for the censored newsvendor problem, outperforming existing methods.

Unified framework for imitating tasks across domains with discrepancies.

problem Learning tasks across domains with embodiment, viewpoint, and dynamics mismatches.
method Two-step approach: alignment followed by adaptation. Alignment uses Generative Adversarial MDP Alignment (GAMA) for state and action correspondences from unpaired, unaligned demonstrations. Adaptation leverages these correspondences for zero-shot imitation.
result Effectiveness of the proposed approach in embodiment, viewpoint, and dynamics mismatch scenarios.

Subspace models play an important role in a wide range of signal processing tasks, and this paper explores how the pairwise geometry of subspaces influences the probability of misclassification. When the mismatch between the signal and the model is vanishingly small, the probability of misclassification is determined b…

2015-07-15abs ↗pdf ↗

Synthetic augmentation helps but not always in imbalanced learning.

problem Imbalanced learning causes poor performance on rare classes.
method Developed a statistical framework for synthetic augmentation in imbalanced learning.
result Synthetic augmentation is not always beneficial and depends on the imbalance regime.

We study the tracking problem, namely, estimating the hidden state of an object over time, from unreliable and noisy measurements. The standard framework for the tracking problem is the generative framework, which is the basis of solutions such as the Bayesian algorithm and its approximation, the particle filters. Howe…

2012-03-15abs ↗pdf ↗

Training on mixed distributions improves test performance even when components are unrelated.

problem Improving test performance with mismatched training and test distributions.
method Analyzing mixture distributions with different training and test proportions.
result Distribution shift can be beneficial, improving test performance even when components are unrelated.

Empirical study compares finite- and infinite-width BNNs, revealing performance differences under model mismatch.

problem Comparing BNNs with different widths due to conflicting model properties and inference intractability.
method Empirical comparison of finite- and infinite-width BNNs, analyzing performance under model mismatch.
result Increasing width can hurt BNN performance when the model is mis-specified, and finite-width BNNs generalize better under model mismatch.

We characterize the performance of sequential information guided sensing, Info-Greedy Sensing, when there is a mismatch between the true signal model and the assumed model, which may be a sample estimate. In particular, we consider a setup where the signal is low-rank Gaussian and the measurements are taken in the dire…

2015-01-26abs ↗pdf ↗

RFC enhances humanoid control to imitate complex human motions.

problem Dynamics mismatch between humanoid models and real humans.
method Residual Force Control (RFC) augments control policies with external forces.
result RFC outperforms state-of-the-art methods in convergence speed and motion quality.