Paper proposes a method to identify negative transfers in multitask learning using surrogate models.
problem Identifying subsets of source tasks that improve target task performance in multitask learning.
method Surrogate modeling to precompute multitask learning performances and approximate them with a linear regression model.
result The approach predicts negative transfers from multiple source tasks to target tasks more accurately than existing methods.
This paper reduces graph-Laplacian dimensionality for spectral clustering of subsets.
problem High computational cost of graph-Laplacian spectral embedding for large data sets.
method Develops two algorithms for low-dimensional graph-Laplacian representation of target subsets.
result Ensures consistency of target subset clustering with full data set spectral clustering.
New method finds subsets of vertices to minimize interventions in causal discovery.
problem Learning only part of the causal graph from interventional data.
method Introducing Meek separators and an algorithm to find them efficiently.
result First known average-case provable guarantees for subset search and causal matching.
This paper uses MIO to select features for kernel SVM classification.
problem Feature selection for kernel SVM classification.
method Mixed-integer optimization (MIO) for feature subset selection.
result The MIO approach can often outperform linear-SVM-based methods in prediction performance.
Two-parameter models can learn high-dimensional targets via gradient flow.
problem Learning high-dimensional targets with limited parameters.
method Gradient flow approach for W<d models. result Two-parameter models can learn targets with arbitrarily high success probability.
Sharp sample complexity for learning bounded Lp subsets.
problem Learning bounded subsets of Lp with p>4. method Heavy-tailed learning procedure.
result Sharp sample complexity estimate for any p>4. Study on images and singularities of pseudoholomorphic maps.
problem Characterize images and singularities of pseudoholomorphic maps.
method Analyzes pseudoholomorphic maps in domains and targets of dimension four.
result Proves properties of images and singularities of pseudoholomorphic maps.
Study efficient algorithms for identifying minimum interventional sets to learn causal relationships.
problem Identify the smallest set of interventions to learn causal relationships between a subset of edges.
method Develop algorithms for subset verification and search problems under assumptions of faithfulness, causal sufficiency, and ideal interventions.
result For subset verification, an efficient algorithm is provided to compute a minimum sized interventional set.
This paper defines a generalized column subset selection problem which is concerned with the selection of a few columns from a source matrix A that best approximate the span of a target matrix B. The paper then proposes a fast greedy algorithm for solving this problem and draws connections to different problems that ca…
The paper finds symplectic compactifications of coadjoint orbits.
problem Understanding symplectic structures on coadjoint orbits.
method Defined real analytic symplectomorphisms on subsets of coadjoint orbits.
result Coadjoint orbits of compact Lie algebras are symplectic compactifications of domains of cotangent bundles.
Parallel decoding improves machine translation efficiency and accuracy.
problem Efficiently generating translations from left to right.
method Conditional masked language modeling for parallel decoding.
result Improves translation performance by over 4 BLEU points.
ViTaX provides formal guarantees for targeted explanations in safety-critical systems.
problem Need trustworthy explanations for safety-critical deep neural networks.
method Formal reachability analysis for targeted, semifactual explanations.
result First method to provide formally guaranteed explanations of model resilience.
A single pre-trained agent guides feature selection using knockoffs.
problem Feature selection challenges in AI-readiness of data.
method Generates knockoff features and uses reinforcement learning.
result Optimal feature subset identified with reduced dependency on target variable.
Markov random field (MRF) learning is intractable, and its approximation algorithms are computationally expensive. We target a small subset of MRF that is used frequently in computer vision. We characterize this subset with three concepts: Lattice, Homogeneity, and Inertia; and design a non-markov model as an alternati…
Charities can increase donations by targeting optimal recipients.
problem Ineffective fundraising leads to lower resources for goods.
method Combines field experiment and causal machine-learning approach.
result Machine-learning-based optimal targeting increases donations significantly.
Efficient algorithm for estimating target mean under known sampling distribution.
problem Statistical estimation under known sampling distribution without distributional assumptions.
method Worst-case analysis of weighted combination of sample values.
result Worst-case expected error is at most a π/2 factor worse than optimal.
Adaptive algorithm samples k items from a DPP without seeing all items.
problem Efficiently sample k items from a DPP without preprocessing all n items. method Adaptive uniform sampling of a subset of data, followed by k-DPP sampling on this subset. result Produces a k-DPP sample after observing only a small fraction of all elements, significantly faster than state-of-the-art. Finding an informative subset of a large collection of data points or models is at the center of many problems in computer vision, recommender systems, bio/health informatics as well as image and natural language processing. Given pairwise dissimilarities between the elements of a `source set' and a `target set,' we co…
New estimators improve efficiency in two-phase designs with coarsened data.
problem Efficient estimation in two-phase designs with incomplete data.
method Developed new estimators within the TMLE framework.
result New estimators are asymptotically equivalent and more efficient.
Method leverages data transfer for estimating CATE with KRR.
problem Leveraging findings from one study to estimate CATE in a different population.
method Overlap-adaptive transfer learning of CATE using kernel ridge regression.
result The method achieves superior efficiency and adaptability in estimating CATE.
Causal invariance can improve finite-sample domain adaptation, but only when the target risk margins are large.
problem Finite-sample domain adaptation
method Linear regression with causal knowledge
result Adaptive aggregation can match best candidate predictor while avoiding negative transfer
We show how to train a quantum network of pairwise interacting qubits such that its evolution implements a target quantum algorithm into a given network subset. Our strategy is inspired by supervised learning and is designed to help the physical construction of a quantum computer which operates with minimal external cl…
In standard graph clustering/community detection, one is interested in partitioning the graph into more densely connected subsets of nodes. In contrast, the "search" problem of this paper aims to only find the nodes in a "single" such community, the target, out of the many communities that may exist. To do so , we are …
Proposes a few-shot learning method for feature selection without labeled data.
problem Feature selection in unlabeled data with limited instances.
method Uses Concrete random variables and permutation-invariant neural networks to select features from multiple source tasks.
result Outperforms existing methods in feature selection performance.
We prove that a sequence of Fueter sections of a bundle of compact hyperkahler manifolds X over a 3-manifold M with bounded energy converges (after passing to a subsequence) outside a 1-dimensional closed rectifiable subset S⊂M. The non-compactness along S has two sources: (1) Bubbling-off…
In this paper, we consider active information acquisition when the prediction model is meant to be applied on a targeted subset of the population. The goal is to label a pre-specified fraction of customers in the target or test set by iteratively querying for information from the non-target or training set. The number …
Domain adaptation (DA) addresses the real-world image classification problem of discrepancy between training (source) and testing (target) data distributions. We propose an unsupervised DA method that considers the presence of only unlabelled data in the target domain. Our approach centers on finding matches between sa…
Method identifies unknown intervention targets in structural causal models from diverse data.
problem Identifying unknown intervention targets in structural causal models from heterogeneous data.
method Two-phase approach: first recovers exogenous noises, second matches with endogenous variables.
result Proposed method uniquely identifies intervention targets under causal sufficiency assumption.
New theoretical insights improve feature selection using mutual information.
problem Improving feature selection methods with theoretical guarantees.
method Proposed novel stopping condition for greedy feature selection methods.
result Ideal prediction error remains bounded by a threshold.
We present a general framework for accelerating a large class of widely used Markov chain Monte Carlo (MCMC) algorithms. Our approach exploits fast, iterative approximations to the target density to speculatively evaluate many potential future steps of the chain in parallel. The approach can accelerate computation of t…
Deep neural networks require a large amount of labeled training data during supervised learning. However, collecting and labeling so much data might be infeasible in many cases. In this paper, we introduce a source-target selective joint fine-tuning scheme for improving the performance of deep learning tasks with insuf…
TWINs tackles PDA by weighting inconsistency between two networks.
problem PDA where target classes are a subset of source classes.
method Two Weighted Inconsistency-reduced Networks (TWINs) with two classification networks and weighted classification loss.
result TWINs outperforms other methods in PDA datasets.
Efficiently approximates integrals using a subset of samples from a target distribution in RKHS.
problem Approximating integrals with a target distribution using limited pointwise evaluations.
method Proposes a procedure using a small random subset of samples from the target distribution, either uniformly or using approximate leverage scores.
result Upper bound on approximation error for both sampling strategies, achieving optimal rate with reduced evaluations.
Paper proposes scalable algorithm to estimate intervention targets in linear models.
problem Estimating intervention targets in linear models from observational and interventional data.
method The paper proposes a scalable algorithm that estimates intervention sites from the difference between precision matrices of observational and interventional datasets.
result The algorithm consistently identifies all intervention targets and updates observational Markov equivalence classes to interventional ones.
SNAP efficiently identifies causal effects without needing full graph learning.
problem Efficiently estimating causal effects on a subset of variables.
method Sequential Non-Ancestor Pruning (SNAP) framework.
result SNAP reduces independence tests and computation time without sacrificing causal effect estimations.
The paper describes flows of MMD functionals with distance kernel and quantile functions.
problem Wasserstein gradient flows of MMD functionals with negative distance kernel.
method Characterization via Cauchy problem on L2(0,1), solution via subdifferential construction. result Flow invariance and smoothing properties on subsets of C(0,1), absolute continuity of initial measures. A new portfolio model DEWSP improves Sharpe ratio by 0.24% to 5.15%.
problem High sensitivity of optimized portfolios to estimation errors.
method Deep learning algorithms predict returns for top-N ranked assets, then equally weight them.
result DEWSPs provide an improvement rate of 0.24% to 5.15% in terms of monthly Sharpe ratio compared to HEWSPs.
A method for finding most influential sets reduces a complex problem to a sequence of simpler top-k problems.
problem Identifying most influential subsets in complex models.
method Reduces the problem to a sequence of top-k problems using Dinkelbach's method. result The method returns a globally optimal set for the univariate ratio objective, including partial linear models.
Model predicts diverse chemical reactions for target compounds.
problem Making generalizable and diverse retrosynthetic reaction predictions.
method Transformer architecture with novel pre-training methods and a latent variable model.
result Improves performance on USPTO-50k dataset, generating more diverse predictions.
New method for valid prediction intervals in counterfactual outcomes with runtime confounding.
problem Valid prediction intervals for counterfactual outcomes under runtime confounding.
method Debiased machine learning framework grounded in semiparametric efficiency theory.
result Prediction intervals achieve desired coverage rates with faster convergence compared to standard methods.
Study on infinitely-wide CNNs and their adaptability to function spatial scales.
problem Understanding how CNNs efficiently learn high-dimensional functions and their adaptability to function spatial scales.
method Study infinitely-wide deep CNNs in the kernel regime, characterizing their spectrum and using generalisation bounds to prove adaptability.
result Deep CNNs adapt to the spatial scale of the target function, with error decay controlled by the effective dimensionality of function subsets.
Methods of transfer learning try to combine knowledge from several related tasks (or domains) to improve performance on a test task. Inspired by causal methodology, we relax the usual covariate shift assumption and assume that it holds true for a subset of predictor variables: the conditional distribution of the target…
The MAXENT principle helps merge datasets to infer causal effects.
problem Inferring causal effects from unobserved variables.
method Using the maximum entropy principle with causal sufficiency and faithfulness assumptions.
result Identifies causal edges among variables from merged datasets.
Deep transfer learning improves automatic sleep staging accuracy.
problem Small cohort data variability and inefficiency in sleep studies.
method Deep transfer learning approach using a large dataset to a small cohort.
result Significant performance improvement on automatic sleep staging.
METASET selects diverse unit cells for efficient data-driven metamaterial design.
problem Imbalanced datasets in unit cells can bias data-driven metamaterial design.
method METASET uses similarity metrics and DPPs to select diverse subsets of unit cells.
result Smaller, diverse subsets improve search process and structural performance.
Proposes a new method for finding non-redundant, standout subgroups in numeric datasets.
problem Mining large numbers of redundant subgroups in numeric datasets.
method Dispersion-aware problem formulation based on MDL principle for subgroup set discovery.
result Empirically demonstrates SSD++ returns outstanding subgroup lists.
BiKaehler geometry is characterized by a Riemannian metric g_{ab} and two covariantly constant generally non commuting complex structures K_+^a_b, K_-^a_b, with respect to which g_{ab} is Hermitian. It is a particular case of the biHermitian geometry of Gates, Hull and Roceck, the most general sigma model target space …
DELTA improves transfer learning by aligning feature maps of target networks.
problem Limited accuracy in fine-tuning pre-trained networks for new tasks.
method DELTA preserves outer layer outputs of target networks through constrained feature maps learned by attention.
result DELTA outperforms state-of-the-art methods in accuracy for new tasks.