AdaSearch improves adaptive sensing in noisy environments.
problem Adaptive source seeking in environments with variable background signals.
method Combines global trajectory planning with principled confidence intervals.
result AdaSearch outperforms uniform sampling and other methods in simulations and hardware tests.
CT compares two distributions using Bayes' theorem and chain rule.
problem Measuring the difference between two probability distributions.
method Conditional transport (CT) using chain rule and Bayes' theorem.
result CT strikes a good balance between mode-covering and mode-seeking behaviors.
We introduce a copula mixture model to perform dependency-seeking clustering when co-occurring samples from different data sources are available. The model takes advantage of the great flexibility offered by the copulas framework to extend mixtures of Canonical Correlation Analysis to multivariate data with arbitrary c…
In machine learning, Domain Adaptation (DA) arises when the distribution gen- erating the test (target) data differs from the one generating the learning (source) data. It is well known that DA is an hard task even under strong assumptions, among which the covariate-shift where the source and target distributions diver…
TASFAR adapts regression models without labeled source data.
problem Lack of labeled source data for domain adaptation.
method Uses prediction confidence to estimate target label distribution and calibrate source model.
result Substantially reduces errors in various regression tasks.
AR-Flow VAE improves blind source separation with flexible autoregressive priors.
problem Unsupervised blind source separation of latent signals from mixtures.
method AR-Flow VAE uses autoregressive flows to model latent sources, enhancing flexibility and capturing complex dependencies.
result AR-Flow VAE effectively separates latent sources, demonstrating improved performance over conventional methods.
This research improves representation learning for new domains with limited new supervision.
problem Learning representations that generalize well to new domains with minimal new data.
method Encourages linearity of factors of variation through learned linear transformations called latent canonicalizers.
result Reduces the number of observations needed to generalize to a similar target domain compared to supervised baselines.
A new heuristic strategy improves sparse BSS performance.
problem Efficiently solving non-convex penalized matrix factorization problems.
method Combining heuristic approach with PALM algorithm.
result Significantly improved separation results on astrophysical data.
Improves domain adaptation by clustering target representations.
problem Learning invariant and discriminative representations for unlabeled target domains.
method Simultaneously learns tightly clustered target representations and assigns each cluster to a unique class from the source.
result Achieves state-of-the-art performance in balanced, imbalanced, and partial domain adaptation.
Framework for few-shot relation classification with minimal training data.
problem Few-shot relation classification with limited training data.
method Meta-learning framework that combines instance and support knowledge.
result Framework outperforms state-of-the-art results and achieves competitive performance with large training data.
The paper tackles the trade-off between fairness and accuracy in machine learning models.
problem Ensuring fairness in machine learning often reduces model accuracy.
method The paper introduces formal tools for reconciling the fairness-accuracy tension using Pareto optimality from multi-objective optimization.
result The Chebyshev scalarization scheme is superior for finding Pareto optimal solutions compared to the linear scalarization scheme.
Proposes using Wasserstein barycenters for robust optimization with multiple data sources.
problem Distributionally robust optimization with multiple heterogeneous data sources.
method Construct nominal distribution through Wasserstein barycenter of multiple data samples, reformulates as a finite convex program.
result Proposed scheme outperforms other estimators in sparse inverse covariance matrix estimation.
CRCCA framework improves non-linear CCA with compressed representations.
problem Non-linear CCA for multi-view data with limited samples.
method Information-theoretic compressed representation framework (CRCCA) based on lattice quantization.
result The CRCCA framework provides theoretical bounds and optimality conditions, offering a flexible and computationally efficient solution.
Optimal schedules improve transport map approximation and learning.
problem Improving the approximation and learning of transport maps.
method Optimal scheduling of transport maps to minimize spatial Lipschitz constant.
result The optimal schedule can be computed in closed form and results in a significantly smaller Lipschitz constant.
Paper proposes a novel online transfer learning method to reduce domain discrepancy.
problem Online transfer learning with online distribution discrepancy minimization.
method The method seeks a new feature representation to simultaneously reduce marginal and conditional distribution discrepancies.
result The proposed method outperforms state-of-the-art methods in comprehensive experiments.
Paper tackles domain adaptation without labeled target data.
problem Performing well on an unlabeled target domain using only labeled source data.
method Learning self-supervised tasks on both source and target domains simultaneously.
result Successfully generalizes to the unlabeled target domain.
Deep learning approximates Bayesian posteriors for gravitational-wave data.
problem Efficiently estimating posterior probabilities for gravitational-wave signals.
method Train a neural network to approximate the posterior distribution from signal + noise data.
result The neural network produces a parametrized approximation of the posterior distribution.
Metric anomalies arising from a distribution of point defects (intrinsic interstitials, vacancies, point stacking faults), thermal deformation, biological growth, etc. are well known sources of material inhomogeneity and internal stress. By emphasizing the geometric nature of such anomalies we seek their representation…
The paper tackles MSDA by learning dictionary atoms in Wasserstein space.
problem Mitigating data distribution shifts across multiple source domains to target domain.
method Dictionary learning and optimal transport in Wasserstein space; DaDiL algorithm for learning.
result Improved classification performance by 3.15%, 2.29%, and 7.71% in benchmarks.
LADDER improves DG by reweighting domain-specific classifiers.
problem Challenges of domain generalization when causal mechanisms vary across domains.
method LADDER learns causal and style representations, reweights classifiers at inference.
result LADDER achieves gains in accuracy on various DG tasks.
Estimates vaccine effectiveness and immune correlates in TND studies with missing data.
problem Confounding and missing data in TND studies of vaccine effectiveness and immune correlates.
method Targeted maximum likelihood estimation using a semiparametric logistic regression model.
result Valid causal inference of vaccine effectiveness and immune correlates in TND studies with missing exposure data.
Study optimal reinsurance pricing under model uncertainty for multiple insurers.
problem Optimal reinsurance pricing in the presence of multiple sources of model uncertainty.
method Solves a continuous-time Stackelberg game for general reinsurance contracts, considering entropy penalties and ambiguity in insurers' models.
result Reinsurer prices under a distortion of the barycentre of insurers' models, maximizing expected wealth with an entropy penalty.
The paper analyzes phase transitions in transfer learning for perceptrons.
problem Understanding when transfer learning from a source task to a target task is beneficial.
method Theoretical analysis of a pair of related perceptron learning tasks.
result Reveals a phase transition from negative to positive transfer as task similarity changes.
Proposes a method to improve few-shot transfer in off-dynamics RL.
problem Traditional RL struggles with transferring policies between environments with different dynamics.
method Introduces a penalty to regulate source-trained policies in target environments with limited data.
result Improves performance in various off-dynamics RL scenarios compared to existing methods.
Study benchmarks embedding-based entity alignment methods for KGs.
problem Align entities across different KGs using embeddings.
method Survey and categorize 23 embedding-based methods, propose new KG sampling algorithm, develop open-source library.
result Understood strengths and limitations of embedding-based methods.
New method predicts target genres from source genres, unifying music tag systems.
problem Automatic genre inference fails to handle music genre diversity and subjectivity.
method Knowledge-based, statistical, and hybrid translation models.
result Hybrid translation model is most effective for multilabel classification.
Develops a weighting framework to generalize ITRs from source to target populations.
problem Challenges in generalizing ITRs from a source population to a target population with differing characteristics.
method A robust sample weighting framework using a reproducing kernel Hilbert space to balance covariates and improve ITR learning methods.
result Improves ITR estimation for the target population compared to other weighting methods.
Boosting improves ICA for better component recovery.
problem Improving ICA's reliance on prior knowledge of sources.
method Maximizing likelihood via boosting and fixed-point unmixing.
result Boosting-based ICA outperforms existing methods.
SMiRL learns to minimize surprise in unstable environments, improving agent performance.
problem Learning useful behaviors in unpredictable, unstable environments.
method Alternates between learning a density model and improving policy to seek more predictable stimuli.
result SMiRL agents can play games, control robots, and navigate mazes without task-specific rewards.
An asymmetric information model is introduced for the situation in which there is a small agent who is more susceptible to the flow of information in the market than the general market participant, and who tries to implement strategies based on the additional information. In this model market participants have access t…
In our previous works, we proposed a physically-inspired rule to organize the data points into an in-tree (IT) structure, in which some undesired edges are allowed to occur. By removing those undesired or redundant edges, this IT structure is divided into several separate parts, each representing one cluster. In this w…
Cryptocurrency traders increase stock risk-seeking behavior.
problem Understanding the motivations behind cryptocurrency trading.
method Individual-level brokerage data analysis of stock trading behavior.
result Cryptocurrency traders increase risk-seeking behavior in stocks when engaging in cryptocurrency trading.
New method for robustly estimating barycenters in data aggregation.
problem Outliers and noise in data measures hinder traditional OT barycenter estimation.
method Proposes a novel scalable approach using semi-unbalanced neural optimal transport.
result Demonstrates robustness to outliers and class imbalance.
Wide-AdGraph detects ads and trackers using a graph of resource requests.
problem Detecting and blocking ad trackers to protect user privacy.
method Combining a large-scale graph of resource requests from multiple websites to train a machine learning algorithm.
result High accuracy (96.1% biased, 90.9% unbiased) in detecting ads and trackers.
The paper explores optimal regularizers for data sources, linking them to star bodies.
problem Understanding optimal regularizers for data sources.
method Investigates optimal regularizers for data distributions using star bodies and dual Brunn-Minkowski theory.
result Identifies optimal regularizers and assesses amenability to convex regularization.
A model predicts user movie preferences based on novelty-seeking traits.
problem Accurately predicting user movie preferences for competitive websites.
method DFNSM model uses demographic, genre, and novelty-seeking data.
result DFNSM outperforms previous models in movie recommendation accuracy.
Method quantifies sensitivity of reliability analysis to uncertainty sources.
problem Computational expense in reliability analysis of complex models.
method Gaussian process surrogate model, active learning, sensitivity analysis.
result Reduces main source of error in estimating rare event probabilities.
Two semi-supervised manifold alignment methods improve cross-domain classification.
problem Aligning data from multiple sources for better analysis.
method SPUD and MASH methods using graph integration and diffusion.
result SPUD and MASH methods outperform existing methods in cross-domain classification.
This work studies the statistical performance of Sinkhorn iterations in estimating Schrödinger bridges.
problem Estimating Schrödinger bridges with limited samples.
method Intermediate Sinkhorn iterations applied to the time-dependent drifts of SDEs.
result Established a statistical bound on the squared total variation error of Sinkhorn bridge iterations.
Analyzes how learning algorithms affect and are affected by data manipulation.
problem Characterizing the closed-loop behavior of learning algorithms in the presence of decision-dependent data.
method Analyzes repeated risk minimization as perturbed gradient flows of performative risk minimization, considering multiple local minimizers.
result Characterizes the region of attraction for various equilibria and introduces performative alignment.
Paper uses Super-App data to improve income estimation models.
problem Improving accuracy of income estimation models.
method TreeSHAP method for Stochastic Gradient Boosting Interpretation.
result Alternative data from Super-Apps capture more information than traditional financial data.
A new method improves ICA performance by approximating MDI.
problem Improving F astICA's performance with nonlinear functions.
method Second-order approximation of MDI for joint maximization.
result Efficiency validated through experiments compared to other ICA algorithms.
EKG-based models show better stability across patient populations than EHR-based models.
problem Model generalization issues in EHR and EKG-based predictive models.
method Two tests to measure model generalization, comparing EHR and EKG data.
result EKG-based models are more stable across different patient populations.
Domain adaptation is transfer learning which aims to generalize a learning model across training and testing data with different distributions. Most previous research tackle this problem in seeking a shared feature representation between source and target domains while reducing the mismatch of their data distributions.…
New method uses limited labeled data and multiple starts to adapt models across domains.
problem Accurate predictions in target domain with few labeled data.
method Fine-tuning from multiple adaptive starts, extending UDA methods.
result Minimax-optimal target performance with limited labeled target data.
GO-OED maximizes predictive information gain on nonlinear QoIs.
problem Maximizing information gain on nonlinear predictive quantities.
method Nested Monte Carlo estimator, Markov chain Monte Carlo, kernel density estimation, Bayesian optimization.
result GO-OED outperforms conventional OED in nonlinear settings.
ALICE combines feature selection and inter-rater agreeability for ML model insights.
problem Improving interpretability of black box machine learning models.
method Integrates feature selection and inter-rater agreeability into a user-friendly Python library.
result Initial experiments on customer churn modeling show promising insights.
Examines optimal risk sharing with realistic risk attitudes, finding risk seeking in certain subdomains.
problem Optimal risk sharing with empirically realistic risk attitudes.
method Allows for risk-seeking agents, generalizes expected utility, and uses counter-monotonic improvement theorem.
result First empirical results on optimal risk sharing with realistic risk attitudes.