This paper addresses GE estimation in non-standard settings using various resampling methods.
problem Biased GE estimates in non-standard settings like clustered data and concept drift.
method Tailored resampling methods for clustered, spatial, unequal sampling, concept drift, and hierarchically structured outcomes.
result Standard resampling methods often yield biased GE estimates in non-standard settings.
Paper shows adversarial training can be fooled by new type of noise.
problem Adversarial training can be fooled by new types of noise.
method Designing ADVIN, a new type of inducing noise.
result ADVIN can degrade adversarial training robustness by 99.9%.
We propose iSCMs to standardize SCM data, improving causal inference.
problem Artifacts in SCM data can lead to misleading conclusions.
method We introduce internally-standardized structural causal models (iSCMs).
result iSCMs are not Var-sortable and mostly not R2-sortable. Adversarial training can degrade standard accuracy even when optimal for robust accuracy.
problem Tradeoff between standard and robust accuracy in adversarial training.
method Analyzes adversarial training's impact on standard accuracy, even when optimal for robust accuracy.
result Even with optimal predictors, adversarial training can still degrade standard accuracy.
Interestingness measures provide information that can be used to prune or select association rules. A given value of an interestingness measure is often interpreted relative to the overall range of the values that the interestingness measure can take. However, properties of individual association rules restrict the val…
Many unsupervised kernel methods rely on the estimation of the kernel covariance operator (kernel CO) or kernel cross-covariance operator (kernel CCO). Both kernel CO and kernel CCO are sensitive to contaminated data, even when bounded positive definite kernels are used. To the best of our knowledge, there are few well…
Proposes standards for evaluating online machine learning methods in evolving data streams.
problem Difficulty in evaluating online machine learning methods under realistic conditions.
method Proposes comprehensive evaluation standards, performance measures, and evaluation strategies.
result Provides a new Python framework (float) for modular integration of libraries and custom code.
Historical (Stressed-) Value-at-Risk ((S)VAR), and Expected Shortfall (ES), are widely used risk measures in regulatory capital and Initial Margin, i.e. funding, computations. However, whilst the definitions of VAR and ES are unambiguous, they depend on input distributions that are data-cleaning- and Data-Model-depende…
LOCA learns standardized data coordinates from measurements.
problem Learning invariant data coordinates from non-linearly deformed manifolds.
method LOCA, a LOcal Conformal Autoencoder, learns an isometric embedding.
result LOCA preserves geometric information while learning invariant coordinates.
ProSper learns data components with non-standard priors and superpositions.
problem Learning complex data components with non-standard priors and superpositions.
method Probabilistic algorithms for sparse coding with non-standard priors and superpositions.
result Library supports scalable and parallelizable dictionary learning for large-scale applications.
New method balances multivariate model fitting for mixed likelihoods.
problem Multivariate models often fit only a subset of observed variables.
method Lipschitz standardization for data preprocessing.
result Lipschitz standardization leads to more accurate multivariate models.
More data can widen the gap between robust and standard models against adversarial attacks.
problem The gap between adversarially robust and standard machine learning models' generalization error increases with more data.
method Theoretical analysis of Gaussian and Bernoulli models under ℓ∞ attacks, and experiments on linear regression models. result Additional data can increase the generalization gap between adversarially robust and standard models.
Meta-Router optimizes LLM selection using gold-standard and preference-based data.
problem Training a high-quality LLM router with combined data sources is challenging due to bias and scarcity.
method Developed an integrative causal router training framework to correct bias and improve routing accuracy.
result Our approach delivers more accurate routing and improves the trade-off between cost and quality.
Paper explores tradeoff between standard and robust accuracy for latent models.
problem Tradeoff between standard accuracy and robust accuracy in adversarial training.
method Revisits adversarial training for latent models, considering Gaussian mixture and generalized linear models.
result Low-dimensional manifold structure mitigates the tradeoff between standard and robust accuracy.
A new method for CT using graph-based regularization.
problem Transfer calibrations between instruments without suitable transfer standards.
method Employing manifold regularization of PLS objective to enforce invariant projections in latent variable space.
result Implicit removal of inter-device variation in predictive directions.
Paper discusses methods to measure privacy in synthetic tabular data.
problem Lack of standard methods to quantify privacy in synthetic data.
method Discusses proposed quantification approaches for synthetic data privacy.
result Contributes to SD privacy standards and stimulates discussion.
ProteinNet provides a standardized data set for protein structure prediction.
problem Lack of standardized data sets for protein structure prediction.
method Created high-quality sequence alignments, multiple data splits, and validation sets.
result Facilitates fair assessment of machine learning models for protein structure.
We study solutions to the static vacuum Einstein equations on exterior domains with prescribed metric and mean curvature on the inner boundary. It is proved that for any such boundary data near the standard round boundary data in Euclidean space, there exists a unique AF solution to the static vacuum equations realizin…
Paper classifies pseudomanifolds over stratified spaces.
problem Classifying pseudomanifolds over stratified spaces.
method Introducing locally standard T-pseudomanifolds and using characteristic data. result Locally standard T-pseudomanifolds over topological stratified pseudomanifolds are classified by their characteristic data. The standard taxonomy of predictive uncertainty is inconsistent with standard measures.
problem Uncertainty taxonomy and measure inconsistency
method Proof of inconsistency
result Uncertainty is not reducible to data collection
Advances in Python and data standards boost reproducibility in neuroimaging.
problem Lack of reproducibility in scientific research.
method Open data sharing resources, data standards, open-source Python, and software engineering advances.
result Reproducibility in neuroimaging has improved with new tools and practices.
Authors classify 3D locally standard T-pseudomanifolds under weaker conditions.
problem Classifying equivariant homeomorphism types of 3D locally standard T-pseudomanifolds.
method Introduced and classified by characteristic data under homotopy equivalence condition.
result Condition for classification can be removed when dimension is at most three.
Study on TVL computation in DeFi protocols, proposing verifiable metrics.
problem Lack of standardization and verifiability in TVL computation.
method Systematic study of 939 DeFi projects, analyzing methodologies and proposing vTVL.
result 240 protocols use repeated balance queries, limiting verifiability.
We propose a method for building an interpretable recommender system for personalizing online content and promotions. Historical data available for the system consists of customer features, provided content (promotions), and user responses. Unlike in a standard multi-class classification setting, misclassification cost…
Paper extends learning theory to dependent data with uniform risk bounds.
problem Learning with dependent data sequences.
method Derives uniform risk bounds for dependent data using VC-dimension and Rademacher complexity.
result Standard classification risk bounds hold for dependent data, same as for independent data.
New method corrects biased comparisons in two-group data.
problem Unreliable inferences from biased sampling.
method Developed an inference method resilient to sampling biases.
result Controls false positives under moderate bias levels.
Improves NF for complex data distributions with multiple modes.
problem Difficulty in handling data distributions with multiple isolated modes.
method Proposes a new framework using variational latent representation to improve NF.
result Significantly more powerful for generating data distributions with multiple modes.
We observe standard transfer learning can improve prediction accuracies of target tasks at the cost of lowering their prediction fairness -- a phenomenon we named discriminatory transfer. We examine prediction fairness of a standard hypothesis transfer algorithm and a standard multi-task learning algorithm, and show th…
Study on adversarial examples from data size, task, and model factors.
problem Understanding adversarial examples from data size, task, and model perspectives.
method Systematic study on adversarial examples from three aspects: data size, task-dependent, and model-specific factors.
result Adversarial generalization requires more data than standard generalization.
Inference models are a key component in scaling variational inference to deep latent variable models, most notably as encoder networks in variational auto-encoders (VAEs). By replacing conventional optimization-based inference with a learned model, inference is amortized over data examples and therefore more computatio…
This work introduces a tensor-based method to perform supervised classification on spatiotemporal data processed in an echo state network. Typically when performing supervised classification tasks on data processed in an echo state network, the entire collection of hidden layer node states from the training dataset is …
New trees-based models handle correlated data better.
problem Standard trees-based models ignore correlation structure.
method Explicitly accounts for correlation structure in splitting criterion, stopping rules, and fitted values.
result New approach superior to standard models in simulations and real data.
Private PGB boosts synthetic data quality using GANs and privacy techniques.
problem Differentially private GANs struggle with convergence and poor output quality.
method Combines reweighted samples from GAN training using Private Multiplicative Weights method.
result Improves synthetic data quality across various datasets and tasks.
Study proves existence of global solutions for Standard Model on expanding spacetimes.
problem Existence of global solutions for the Standard Model on expanding spacetimes.
method Gauge-invariant energy estimate for the Euler-Lagrange equations.
result Existence of global solutions for the Standard Model under specific conditions.
Paper improves risk estimation for extreme events.
problem Estimating extreme risks accurately.
method Modified Bayes risk for expectiles, asymptotic expansions, efficient estimators.
result Asymptotic normality of estimators proved.
New estimator improves ATT estimation efficiency with external controls.
problem Reduced efficiency when incorporating external controls into ATT estimation.
method Proposes a novel doubly robust estimator for ATT that maintains higher efficiency than standard approaches.
result Demonstrates improved efficiency of the new estimator compared to standard approaches, even under model misspecification.
A method for constructing tight prediction intervals for multiple numerical outputs.
problem Constructing tight prediction intervals for multiple related numerical outputs.
method A novel coordinate-wise standardization procedure that makes residuals comparable across output dimensions, estimating suitable scaling parameters using calibration data.
result The method produces tighter prediction intervals than existing baselines while maintaining valid simultaneous coverage.
Enhances classification accuracy on low data sets using synthetic data.
problem Low sample size in data augmentation.
method Variational Autoencoder and manifold sampling.
result Significant improvement in classification accuracy (e.g., 88.6% vs 80.7%).
Study compares new audio representation methods for limited data music retrieval.
problem Improving machine learning for audio data with limited training data.
method Investigated mel-spectrogram and Mel scattering representations, and augmented target loss function.
result All proposed methods outperform standard mel-spectrogram when using limited data.
In standard generative adversarial network (SGAN), the discriminator estimates the probability that the input data is real. The generator is trained to increase the probability that fake data is real. We argue that it should also simultaneously decrease the probability that real data is real because 1) this would accou…
Reinterprets classifiers as energy-based models for joint distributions.
problem Improving classifier performance and calibration.
method Interprets discriminative classifiers as energy-based models, trains on unlabeled data, and improves model quality.
result Improves calibration, robustness, and out-of-distribution detection.
We study inference and learning based on a sparse coding model with `spike-and-slab' prior. As in standard sparse coding, the model used assumes independent latent sources that linearly combine to generate data points. However, instead of using a standard sparse prior such as a Laplace distribution, we study the applic…
MIMIC-Extract transforms EHR data for reproducible healthcare machine learning.
problem Lack of accessible, standardized healthcare data for machine learning.
method Open-source pipeline for converting raw EHR data into usable dataframes.
result Demonstrates utility through benchmark tasks and baseline results.
We show that there may exist an inherent tension between the goal of adversarial robustness and that of standard generalization. Specifically, training robust models may not only be more resource-consuming, but also lead to a reduction of standard accuracy. We demonstrate that this trade-off between the standard accura…
The paper analyzes adversarial training effects on classification accuracy.
problem Understanding adversarial training's impact on standard and robust accuracy.
method Derived precise statistical analysis for binary classification problems with Gaussian data.
result Theoretical explanation of standard and robust accuracy trends for adversarial training.
X-VAE uses data-adaptive Gaussian priors to improve latent space modeling.
problem Limitations of standard Gaussian priors in complex datasets.
method Data-adaptive Gaussian prior derived from pretrained autoencoder latent codes.
result Improved latent space modeling and generation quality.
Many applications of machine learning, for example in health care, would benefit from methods that can guarantee privacy of data subjects. Differential privacy (DP) has become established as a standard for protecting learning results. The standard DP algorithms require a single trusted party to have access to the entir…
We present new algorithms for learning Bayesian networks from data with missing values using a data augmentation approach. An exact Bayesian network learning algorithm is obtained by recasting the problem into a standard Bayesian network learning problem without missing data. To the best of our knowledge, this is the f…