New method improves IL from imperfect demos using confidence scores.
problem Learning optimal policies from imperfect demonstrations is challenging.
method Proposes two confidence-based IL methods: 2IWIL and IC-GAIL.
result Confidence scores from sub-optimal demos significantly improve IL performance.
A new one-step method for covariate shift adaptation.
problem Real-world data often violates the assumption of same distribution for training and test samples.
method Proposes a one-step optimization approach to jointly learn the model and weights.
result The proposed method achieves a generalization error bound and is empirically effective.
New estimator improves off-policy evaluation for large action spaces.
problem Conventional importance-weighting approaches suffer from excessive variance in off-policy evaluation for large discrete action spaces.
method Proposes OffCEM estimator based on conjunct effect model (CEM), applying importance weighting only to action clusters and using model-based reward estimation for residual effects.
result Proposed estimator is unbiased under local correctness condition, providing substantial improvements in OPE especially with many actions.
Unified framework for sequence models using test-time regression.
problem Designing efficient sequence models with associative memory.
method Formalizing associative recall as regression over input tokens, deriving various sequence models.
result Clarifies the effectiveness of query-key normalization in softmax attention and offers new generalizations.
New method estimates spatial weights matrix for lattice data, improving prediction accuracy.
problem Estimating spatial dependence structure for regular lattice data.
method Adaptive lasso with cross-sectional resampling to estimate sparse spatial weights matrix.
result Improves prediction accuracy of nitrogen dioxide concentrations.
Adaptive importance sampling (AIS) uses past samples to update the \textit{sampling policy} qt at each stage t. Each stage t is formed with two steps : (i) to explore the space with nt points according to qt and (ii) to exploit the current amount of information to update the sampling policy. The very funda…
The weighted and directed network of countries based on the number of overseas banks is analyzed in terms of its fragility to the banking crisis of one country. We use two different models to describe transmission of shocks, one local and the other global. Depending on the original source of the crisis, the overall siz…
New method identifies parameters of wider shallow neural networks with biases.
problem Identifying parameters of wide shallow neural networks with biases from finite samples.
method Two-step pipeline: direction of weights via second order information, signs via algebraic evaluations, biases via gradient descent.
result Constructive methods and theoretical guarantees of finite sample identification for wider shallow networks with biases.
A new method for clustering functional data outperforms existing methods.
problem Clustering heterogeneous functional linear regression data.
method funWeightClust, a family of parsimonious models based on cluster weighted models.
result funWeightClust outperforms existing methods in simulations and real-world traffic analysis.
Survey on importance weighting in machine learning applications.
problem Distribution shift in supervised learning.
method Weighting objective function or probability distribution based on instance importance.
result Importance weighting can guarantee desirable statistical properties in distribution shift scenarios.
Efficiently optimizes constrained problems with two-step lookahead BO.
problem Optimizing constrained problems with limited computational resources.
method Two-step lookahead Bayesian optimization with inequality constraints, using a novel unbiased gradient estimator.
result Significantly improves query efficiency over previous methods.
A new method uses nearest neighbors for importance weighting.
problem Data covariate shift problems in machine learning.
method Nearest neighbor classification scheme for determining importance weights.
result Demonstrated effectiveness through comparative experiments on various classification tasks.
The paper extends two-step homogeneous geodesics to homogeneous Finsler spaces.
problem Extending two-step homogeneous geodesics to Finsler spaces.
method Providing sufficient conditions for (α,β) spaces and decomposable cubic spaces to have two-step Finsler geodesic orbit spaces. result Presented examples of two-step Finsler geodesic orbit spaces.
DFA trains deep networks by aligning weights then memorizing data.
problem Understanding why DFA works for some networks but not others.
method Two-step learning process: alignment followed by memorization.
result DFA aligns weights to maximize gradient alignment, breaking degeneracy.
Classifies two-step solvable Lie groups with SKT structures.
problem Classifying Lie groups with SKT structures.
method Shear construction and analysis of SKT shear data on Abelian Lie algebras.
result Large part of the classification for two-step solvable SKT algebras of dimension six.
We consider a method popular in the literature of associating a two-step nilpotent Lie algebra with a finite simple graph. We prove that the two-step nilpotent Lie algebras associated with two graphs are Lie isomorphic if and only if the graphs from which they arise are isomorphic.
A new method forecasts financial tail risks by combining and weighting quantiles.
problem Reducing uncertainty in financial tail risk forecasting.
method Two-step procedure: quantile combination followed by ES computation.
result The proposed framework outperforms individual models and simple approaches.
Paper introduces new actuarial-consistent valuations for insurance liabilities.
problem Valuation of insurance liabilities considering both financial and actuarial risks.
method Proposes two-step actuarial valuations and actuarial-consistent procedures.
result Actuarial-consistent valuations are equivalent to two-step actuarial valuations under coherence.
PS^2 selects assets then weights for high-dimensional investing.
problem High-dimensional mean--variance investing challenges.
method Two-step framework: Lasso screening followed by standard portfolio estimation.
result FPS^2 with defactored returns improves performance.
New loss function restores importance weighting in overparameterized models.
problem Restoring importance weighting in overparameterized neural networks.
method Introduced polynomially-tailed losses to restore effects of importance weighting.
result Polynomially-tailed losses improve performance in correcting distribution shift.
New algorithm improves accuracy of importance weights for diverse applications.
problem Improving accuracy of importance weights for various applications.
method Formulated multicalibrated partitions and developed an efficient algorithm.
result Algorithm significantly improves accuracy of importance weights.
Cross-validation under sample selection bias can, in principle, be done by importance-weighting the empirical risk. However, the importance-weighted risk estimator produces sub-optimal hyperparameter estimates in problem settings where large weights arise with high probability. We study its sampling variance as a funct…
A new SBM for non-negative zero-inflated edge weights in networks.
problem Modeling international trading networks with non-negative zero-inflated edge weights.
method Restricted Tweedie distribution and nodal information accounting.
result Efficient two-step algorithm for estimating covariate effects.
Proves conjecture about compatible SKT and balanced metrics on compact solvmanifolds.
problem Compact complex manifolds with both SKT and balanced metrics.
method Shear construction and classification of two-step solvable Lie algebras.
result Proves conjecture for compact two-step solvmanifolds with invariant complex structures.
Importance-weighted risk minimization is a key ingredient in many machine learning algorithms for causal inference, domain adaptation, class imbalance, and off-policy reinforcement learning. While the effect of importance weighting is well-characterized for low-capacity misspecified models, little is known about how it…
Improved neural spike inference from calcium imaging data.
problem Neural spike inference from calcium imaging data.
method Importance weighted adversarial variational autoencoders (IWAE) with adversarial training.
result Adversarial IWAE methods outperform VAEs in inferring neural spikes.
Kernel methods accurately predict Hamiltonian systems from data.
problem Data-driven simulation of Hamiltonian systems.
method Two-step and one-step kernel-based methods for identifying and forecasting Hamiltonian systems.
result Framework achieves accurate, data-efficient predictions across various benchmark systems.
Importance-weighting is a popular and well-researched technique for dealing with sample selection bias and covariate shift. It has desirable characteristics such as unbiasedness, consistency and low computational complexity. However, weighting can have a detrimental effect on an estimator as well. In this work, we empi…
Proposes novel wSVMs for sparse learning and accurate probability estimation.
problem Sparse features with redundant noise limit the performance of existing wSVMs.
method Develops ℓ1-norm and elastic net regularized wSVMs for automatic variable selection and probability estimation. result Elastic net regularized wSVMs achieve superior performance in variable selection and probability estimation.
Optimizes weights for better model performance in shifting data.
problem Improper importance weighting leads to poor model performance in data shifts.
method Interprets weights as a bias-variance trade-off and optimizes them simultaneously with model parameters.
result Optimizing weights significantly improves model generalization performance.
We prove that two-step analytic sub-Riemannian structures on a compact analytic manifold equipped with a smooth measure and Lipschitz Carnot groups satisfy measure contraction properties.
A new method combines experts' opinions to train regression models with noisy labels.
problem Training regression models with noisy labels from multiple experts.
method Estimate each labeler's expertise and combine opinions using learned weights.
result Empirically outperforms existing techniques on simulated and real data.
A Riemannian Einstein solvmanifold (possibly, any noncompact homogeneous Einstein space) is almost completely determined by the nilradical of its Lie algebra. A nilpotent Lie algebra, which can serve as the nilradical of an Einstein metric solvable Lie algebra, is called an Einstein nilradical. Despite a substantial pr…
Improves transfer learning by weighting importance based on test-over-training density.
problem Distribution shift in training and test data.
method Joint and dynamic importance-predictor estimation, causal mechanism transfer.
result Enhanced transfer learning performance in complex, high-dimensional tasks.
Direct learning framework for integrating multi-source causal data.
problem Conditional average treatment effects inference from heterogeneous data.
method Direct learning framework, double robustness, causal information-aware weighting function.
result Effective causal data fusion in both homogeneous and heterogeneous scenarios.
Sharp analysis of out-of-distribution error in overparameterized models with importance weights.
problem Understanding and quantifying the degradation of performance in overparameterized models when faced with underrepresented data.
method Sharp analysis of an overparameterized Gaussian mixture model with spurious features and cost-sensitive interpolating solutions incorporating importance weights.
result Characterization of a novel tradeoff between worst-case robustness and average accuracy as a function of importance weight magnitude.
This paper compares gradient estimators in importance-weighted VI and justifies the superiority of DREP over REP.
problem Understanding the impact of gradient estimators on importance-weighted VI algorithms.
method Unified theoretical comparison of reparameterized and doubly-reparameterized gradient estimators tied to IWAE, VR, and VR-IWAE bounds.
result Formally justifies the superiority of doubly-reparameterized gradient estimators over reparameterized ones in importance-weighted VI.
An intrinsic problem of classifiers based on machine learning (ML) methods is that their learning time grows as the size and complexity of the training dataset increases. For this reason, it is important to have efficient computational methods and algorithms that can be applied on large datasets, such that it is still …
A novel Bayesian computation method using importance weighting improves numerical stability and performance.
problem Bayesian computation stability and performance issues.
method Nonparametric approach via feature means, importance weighting, and kernel Bayes' rule.
result Importance weighted kernel Bayes' rule yields superior numerical stability and performance.
Corrects distribution shift in target shift scenarios using importance weighting.
problem Analyzes importance weighting for correcting distribution shift under target shift.
method Analyzed importance-weighted kernel ridge regression under target shift.
result Shows that importance weighting corrects the train-test mismatch without altering input-space complexity.
This work explores the importance of model weights and Hessian bias in pruning.
problem Understanding the relative importance of model weights for efficient pruning.
method A principled exploration of pruning, focusing on linear models and neural networks.
result Asymptotic formulas reveal the performance of different pruning methods.
New framework forecasts ES using weighted quantiles.
problem Forecasting Expected Shortfall (ES) in financial markets.
method Two-step procedure: VaR estimation through quantile regressions, ES computation as weighted average.
result Proposed models outperform other methods in stock market indices forecasting.
The standard interpretation of importance-weighted autoencoders is that they maximize a tighter lower bound on the marginal likelihood than the standard evidence lower bound. We give an alternate interpretation of this procedure: that it optimizes the standard variational lower bound, but using a more complex distribut…
Importance sampling has become an important tool for the computation of tail-based risk measures. Since such quantities are often determined mainly by rare events standard Monte Carlo can be inefficient and importance sampling provides a way to speed up computations. This paper considers moderate deviations for the wei…
We associate a two-step nilpotent Lie algebra to an arbitrary Schreier graph. We then use properties of the Schreier graph to determine necessary and sufficient conditions for this Lie algebra to extend to a three-step nilpotent Lie algebra. As an application, if we start with pairs of non-isomorphic Schreier graphs co…
Two-stage recommender systems struggle with exploration, leading to linear regret.
problem Linear regret in two-stage recommender systems due to exploration issues.
method Proposed a method to synchronize exploration strategies between the ranker and nominators using LinUCB.
result Demonstrated the effectiveness of the proposed algorithm experimentally.
Unified framework for analyzing pessimism in off-policy learning with regularized importance sampling.
problem High variance in importance weighting for off-policy learning.
method Unified PAC-Bayesian study of pessimism with regularized importance sampling.
result Derivation of a tractable PAC-Bayesian generalization bound for common importance weight regularizations.
Corrects bias in learned generative models using likelihood-free importance weighting.
problem Bias in learned generative models relative to true data distribution.
method Estimate likelihood ratio using a classifier, apply importance weighting.
result Consistently improves goodness-of-fit metrics for deep generative models.