OMD and DA perform similarly in static settings but OMD is inferior under dynamic learning rates.
problem Proving and understanding the performance difference between OMD and DA under dynamic learning rates.
method Introducing stabilization to OMD and modifying its convergence analysis.
result OMD with stabilization and DA have the same performance guarantees under dynamic learning rates.
Paper tackles multi-source domain adaptation for regression.
problem Predicting HDL cholesterol levels using gut microbiome data.
method Two-step procedure: 1) Extend a flexible single-source DA algorithm for classification to regression. 2) Augment with ensemble learning for multi-source DA.
result Consistent improvement in HDL cholesterol level prediction performance over existing methods.
Domain adaptation (DA) is an important and emerging field of machine learning that tackles the problem occurring when the distributions of training (source domain) and test (target domain) data are similar but different. Current theoretical results show that the efficiency of DA algorithms depends on their capacity of …
Single-layer GCN model improves recommendation performance with less complexity.
problem Severe computational burden and excessive model parameters in existing GCN models.
method Proposes a single-layer GCN architecture with a simplified aggregation step using DA similarity.
result Significantly outperforms existing GCN models and achieves up to a few orders of magnitude speedup.
We study few-shot supervised domain adaptation (DA) for regression problems, where only a few labeled target domain data and many labeled source domain data are available. Many of the current DA methods base their transfer assumptions on either parametrized distribution shift or apparent distribution similarities, e.g.…
Ozsvath and Szabo recently constructed an algebraically defined invariant of tangles which takes the form of a DA bimodule. This invariant is expected to compute knot Floer homology. The authors have a similar construction for open braids and their plat closures which can be viewed as a filtered DA bimodule over the sa…
CAD-DA controls anomaly detection under domain adaptation.
problem Valid statistical inference after domain adaptation.
method Conditional Selective Inference to handle domain adaptation effects.
result Valid statistical inference under domain adaptation achieved.
SFS-DA method statistically tests FS reliability under domain adaptation.
problem Feature selection reliability under domain adaptation with limited target data.
method Selective Inference framework to control false positive rate and enhance true positive rate.
result SFS-DA method controls FPR below a pre-specified level α (e.g., 0.05) while maximizing true positive rate. New DA method CIRM outperforms existing methods under structural causal model assumptions.
problem Improving prediction performance in domain adaptation with perturbed source and target data.
method Theoretical framework based on structural causal models to analyze and compare DA methods.
result CIRM method outperforms existing methods when covariates and label distributions are perturbed in target data.
Deep neural networks suffer from performance decay when there is domain shift between the labeled source domain and unlabeled target domain, which motivates the research on domain adaptation (DA). Conventional DA methods usually assume that the labeled data is sampled from a single source distribution. However, in prac…
There are a variety of Domain Adaptation (DA) scenarios subject to label sets and domain configurations, including closed-set and partial-set DA, as well as multi-source and multi-target DA. It is notable that existing DA methods are generally designed only for a specific scenario, and may underperform for scenarios th…
Data augmentation (DA) is commonly used during model training, as it significantly improves test error and model robustness. DA artificially expands the training set by applying random noise, rotations, crops, or even adversarial perturbations to the input data. Although DA is widely used, its capacity to provably impr…
In ultrasound (US) imaging, individual channel RF measurements are back-propagated and accumulated to form an image after applying specific delays. While this time reversal is usually implemented using a hardware- or software-based delay-and-sum (DAS) beamformer, the performance of DAS decreases rapidly in situations w…
ADDA framework speeds up data augmentation in massive data settings.
problem Slow data augmentation in massive data settings.
method Develops asynchronous and distributed data augmentation (ADDA) framework.
result ADDA significantly speeds up data augmentation compared to parent DA algorithms.
TransCal calibrates DA models with lower bias and variance.
problem Calibrating DA models to estimate accurate predictive uncertainty.
method Transferable Calibration (TransCal) in a unified hyperparameter-free optimization framework.
result TransCal achieves more accurate calibration with lower bias and variance.
EnFF uses flows to speed up DA in high dimensions.
problem Efficiently assimilating noisy data in high-dimensional systems.
method Flow Matching (FM) for training-free, scalable data assimilation.
result EnFF accelerates DA with improved cost-accuracy tradeoffs and scalability.
For machine learning task, lacking sufficient samples mean the trained model has low confidence to approach the ground truth function. Until recently, after the generative adversarial networks (GAN) had been proposed, we see the hope of small samples data augmentation (DA) with realistic fake data, and many works valid…
GGDA simplifies DA for large models, speeding up attribution by up to 50x.
problem Computational intensity of existing DA methods limits their applicability to large-scale models.
method Generalized Group Data Attribution (GGDA) framework attributing to groups of training points.
result GGDA achieves up to 50x speedups over standard DA methods while maintaining effectiveness.
This review article surveys data augmentation MCMC algorithms.
problem Sampling from intractable probability distributions.
method Comprehensive study of DA MCMC algorithms, their convergence properties, and acceleration strategies.
result Synthesizes recent developments and provides insights for researchers.
STAND-DA improves AD in DA target domains with limited data.
problem Statistical validity of AD after DA with limited data.
method Selective Inference framework for GPU-accelerated p-value computation. result Valid p-values and controlled false positive rate. DA improves solar wind forecasts by updating model boundary conditions.
problem Improving solar wind forecasting accuracy.
method Variational Data Assimilation with solar wind model and in-situ observations.
result DA forecasts are more accurate than non-DA forecasts, especially when STEREO-B's latitude is offset from Earth.
The paper studies geometric properties of soliton surfaces using an extended Darboux frame field.
problem Geometric analysis of soliton surfaces associated with the Betchov-Da Rios equation.
method Derivative formulas of an extended Darboux frame field, geometric invariants, curvature calculations.
result Construction of curvature ellipse and Wintgen ideal soliton surfaces.
In dialogues, an utterance is a chain of consecutive sentences produced by one speaker which ranges from a short sentence to a thousand-word post. When studying dialogues at the utterance level, it is not uncommon that an utterance would serve multiple functions. For instance, "Thank you. It works great." expresses bot…
Domain adaptation (DA) addresses the real-world image classification problem of discrepancy between training (source) and testing (target) data distributions. We propose an unsupervised DA method that considers the presence of only unlabelled data in the target domain. Our approach centers on finding matches between sa…
There are few papers about the consumption pattern of the Portuguese wine, using econometrics techniques. This work, pretend to analyze the consumers behavior of the wine produced in Portugal, determining the demand equation with panel data methods. There were used statistical data available in the Alentejo Regional Wi…
This work extends ME-RL using diffusion models to sample optimal policies.
problem Sampling from the optimal policy trajectory distribution in ME-RL.
method Introducing Diffusion-Augmented Markov Decision Processes (DA-MDPs) to minimize reverse KL divergence.
result DA-MDPs enable seamless integration into various ME-RL methods and outperform baselines.
Paper analyzes gradient descent with noisy data copies for linear regression, showing regularization and acceleration effects.
problem Improving generalization in machine learning through data augmentation with noise.
method Gradient descent with on-line noisy copies for linear regression analysis.
result Training with on-line noisy copies is equivalent to ridge regularization with a specific regularization parameter.
New method uses data augmentation to improve causal effect estimation.
problem Improving causal effect estimation in the presence of hidden confounders.
method Introduces IV-like regression and data augmentation techniques.
result Data augmentation can simulate worst-case scenarios for causal estimation.
Voice activity detection (VAD), which classifies frames as speech or non-speech, is an important module in many speech applications including speaker verification. In this paper, we propose a novel method, called self-adaptive soft VAD, to incorporate a deep neural network (DNN)-based VAD into a deep speaker embedding …
Deep Neural Networks (DNNs) have recently been achieving state-of-the-art performance on a variety of computer vision related tasks. However, their computational cost limits their ability to be implemented in embedded systems with restricted resources or strict latency constraints. Model compression has therefore been …
Deep learning improves chaotic dynamics filtering without ensemble.
problem Discovering efficient DA schemes for chaotic dynamics.
method Residual Convolutional Neural Network for the analysis step.
result Deep learning achieves ensemble filtering accuracy without an ensemble.
Bayesian model selection optimizes data augmentation for improved machine learning robustness.
problem Choosing optimal data augmentation parameters is challenging and often done through trial and error.
method Interprets augmentation parameters as model hyperparameters and uses Bayesian model selection to optimize them.
result Our approach improves calibration and robust performance on various tasks.
(Unsupervised) Domain Adaptation (DA) seeks for classifying target instances when solely provided with source labeled and target unlabeled examples for training. Learning domain-invariant features helps to achieve this goal, whereas it underpins unlabeled samples drawn from a single or multiple explicit target domains …
Study of discrete analogues of Atiyah sequence in principal bundles.
problem Discrete analogues of vector bundles and connections in principal bundles.
method Analysis in two categories: fiber bundles with sections and local Lie groupoids, defining discrete curvature and splittings.
result Correspondence between splittings of discrete Atiyah sequence and discrete connections with trivial curvature.
PLOT uses optimal transport to find neural site handles for causal abstraction.
problem Finding the relevant neural site for causal analysis is computationally challenging.
method PLOT employs optimal transport to localize causal variables from neural network outputs.
result PLOT efficiently finds intervention handles for causal abstraction in neural networks.
The standard Gibbs sampler of Mixed Multinomial Logit (MMNL) models involves sampling from conditional densities of utility parameters using Metropolis-Hastings (MH) algorithm due to unavailability of conjugate prior for logit kernel. To address this non-conjugacy concern, we propose the application of Pólygamma data a…
Paper optimizes energy trading on DA markets using RL.
problem Volatility and randomness in renewable energy sources.
method Markov Decision Process, reinforcement learning, evolutionary algorithm.
result RL-based strategy generates highest market profits.
DAS-PINNs uses deep learning to solve complex PDEs more accurately.
problem Solving high-dimensional PDEs with high accuracy.
method Deep neural networks and generative models for adaptive sampling.
result DAS-PINNs significantly improves solution accuracy for low regularity and high-dimensional problems.
CoDATS improves DA on time series data with weak supervision.
problem Improving domain adaptation for time series data with limited labeled data.
method CoDATS model for Time Series data, DA-WS method with weak supervision.
result Significant accuracy improvements over state-of-the-art methods.
A new method for handling imbalanced data in regression models.
problem Imbalanced data in regression models with continuous or discrete covariates.
method Combines weighted resampling and data augmentation procedures.
result Improves the accuracy of model estimates by addressing imbalanced data.
New algorithm uses conditionally invariant components to improve domain adaptation performance.
problem Improving domain adaptation performance when source and target data distributions differ.
method Conditionally invariant components (CICs) and importance-weighted conditional invariant penalty (IW-CIP) algorithm.
result New algorithm provides target risk guarantees and addresses label-flipping features.
Machine learning improves model forecasts by correcting errors.
problem Improving short- to mid-range forecasts by correcting model errors.
method Iterative method combining data assimilation and machine learning.
result Hybrid models outperform original models in forecasts.
Regularization and data augmentation can be class-dependent, leading to poor performance on some classes.
problem Class-dependent effects of regularization and data augmentation.
method Evaluation of regularization and data augmentation techniques on Imagenet and INaturalist datasets.
result Regularization and data augmentation can lead to significant performance drops on some classes.
For any Lie algebroid A, its 1-jet bundle JA is a Lie algebroid naturally and there is a representation π: JA ->DA. Denote by dJ the corresponding coboundary operator. In this paper, we realize the deformation cohomology of a Lie algebroid A introduced by M. Crainic and I. Moerdijk as the cohomology of a subcomplex (Γ(…
DA-GNN improves robustness of GNNs by modeling noise dependencies.
problem Real-world graph node features often contain noise, leading to performance degradation in GNNs.
method DA-GNN captures noise dependencies using variational inference and new benchmark datasets.
result DA-GNN consistently outperforms existing baselines across various noise scenarios.
The paper proposes a method to test features selected by SeqFS-DA with controlled FPR.
problem Ensuring reliability of feature selection after domain adaptation in high-dimensional regression.
method Proposes a novel method to test features selected by SeqFS-DA with controlled FPR.
result The proposed method controls FPR below a significance level α (e.g., 0.05) and enhances statistical power. For an entire mapping f:C↦C and a triple (p,α,r)∈(0,∞)×(−∞,∞)×(0,∞], the Gaussian integral means of f (with respect to the area measure dA) is defined by $$ {\mathsf M}_{p,α}(f,r)=\Big({\int_{|z|<r}e^{-α|z|^2}dA(z)}\Big)^{-1}{\int_{|z|<r}|f(z)|^p{e^{-α|z|^…
Proposes a method to adapt to new classes in a domain shift.
problem Learning new classes in a domain shift without labeled supervision.
method Inspired by prototypical networks, the method classifies target samples into shared and novel classes.
result Superior performance compared to DA and CI methods in the CIDA paradigm.