This paper addresses GE estimation in non-standard settings using various resampling methods.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper shows adversarial training can be fooled by new type of noise.
We propose iSCMs to standardize SCM data, improving causal inference.
Interestingness measures provide information that can be used to prune or select association rules. A given value of an interestingness measure is often interpreted relative to the overall range of the values that the interestingness measure can take. However, properties of individual association rules restrict the val…
Many unsupervised kernel methods rely on the estimation of the kernel covariance operator (kernel CO) or kernel cross-covariance operator (kernel CCO). Both kernel CO and kernel CCO are sensitive to contaminated data, even when bounded positive definite kernels are used. To the best of our knowledge, there are few well…
Proposes standards for evaluating online machine learning methods in evolving data streams.
Historical (Stressed-) Value-at-Risk ((S)VAR), and Expected Shortfall (ES), are widely used risk measures in regulatory capital and Initial Margin, i.e. funding, computations. However, whilst the definitions of VAR and ES are unambiguous, they depend on input distributions that are data-cleaning- and Data-Model-depende…
LOCA learns standardized data coordinates from measurements.
New method balances multivariate model fitting for mixed likelihoods.
More data can widen the gap between robust and standard models against adversarial attacks.
Meta-Router optimizes LLM selection using gold-standard and preference-based data.
Paper explores tradeoff between standard and robust accuracy for latent models.
Paper discusses methods to measure privacy in synthetic tabular data.
A new method for CT using graph-based regularization.
We study solutions to the static vacuum Einstein equations on exterior domains with prescribed metric and mean curvature on the inner boundary. It is proved that for any such boundary data near the standard round boundary data in Euclidean space, there exists a unique AF solution to the static vacuum equations realizin…
Paper classifies pseudomanifolds over stratified spaces.
The standard taxonomy of predictive uncertainty is inconsistent with standard measures.
Authors classify 3D locally standard T-pseudomanifolds under weaker conditions.
Study on TVL computation in DeFi protocols, proposing verifiable metrics.
We propose a method for building an interpretable recommender system for personalizing online content and promotions. Historical data available for the system consists of customer features, provided content (promotions), and user responses. Unlike in a standard multi-class classification setting, misclassification cost…
Paper extends learning theory to dependent data with uniform risk bounds.
In many applications, different populations are compared using data that are sampled in a biased manner. Under sampling biases, standard methods that estimate the difference between the population means yield unreliable inferences. Here we develop an inference method that is resilient to sampling biases and is able to …
Most previous works usually explained adversarial examples from several specific perspectives, lacking relatively integral comprehension about this problem. In this paper, we present a systematic study on adversarial examples from three aspects: the amount of training data, task-dependent and model-specific factors. Pa…
Improves NF for complex data distributions with multiple modes.
We observe standard transfer learning can improve prediction accuracies of target tasks at the cost of lowering their prediction fairness -- a phenomenon we named discriminatory transfer. We examine prediction fairness of a standard hypothesis transfer algorithm and a standard multi-task learning algorithm, and show th…
Rapid progress in deep learning has spurred its application to bioinformatics problems including protein structure prediction and design. In classic machine learning problems like computer vision, progress has been driven by standardized data sets that facilitate fair assessment of new methods and lower the barrier to …
Inference models are a key component in scaling variational inference to deep latent variable models, most notably as encoder networks in variational auto-encoders (VAEs). By replacing conventional optimization-based inference with a learned model, inference is amortized over data examples and therefore more computatio…
Convolutional neural network (CNN) architectures have originated and revolutionized machine learning for images. In order to take advantage of CNNs in predictive modeling with audio data, standard FFT-based signal processing methods are often applied to convert the raw audio waveforms into an image-like representations…
This work introduces a tensor-based method to perform supervised classification on spatiotemporal data processed in an echo state network. Typically when performing supervised classification tasks on data processed in an echo state network, the entire collection of hidden layer node states from the training dataset is …
New trees-based models handle correlated data better.
Private PGB boosts synthetic data quality using GANs and privacy techniques.
Study proves existence of global solutions for Standard Model on expanding spacetimes.
Paper improves risk estimation for extreme events.
Robust machine learning relies on access to data that can be used with standardized frameworks in important tasks and the ability to develop models whose performance can be reasonably reproduced. In machine learning for healthcare, the community faces reproducibility challenges due to a lack of publicly accessible data…
New estimator improves ATT estimation efficiency with external controls.
A method for constructing tight prediction intervals for multiple numerical outputs.
Enhances classification accuracy on low data sets using synthetic data.
In standard generative adversarial network (SGAN), the discriminator estimates the probability that the input data is real. The generator is trained to increase the probability that fake data is real. We argue that it should also simultaneously decrease the probability that real data is real because 1) this would accou…
ProSper is a python library containing probabilistic algorithms to learn dictionaries. Given a set of data points, the implemented algorithms seek to learn the elementary components that have generated the data. The library widens the scope of dictionary learning approaches beyond implementations of standard approaches…
We study inference and learning based on a sparse coding model with `spike-and-slab' prior. As in standard sparse coding, the model used assumes independent latent sources that linearly combine to generate data points. However, instead of using a standard sparse prior such as a Laplace distribution, we study the applic…
We show that there may exist an inherent tension between the goal of adversarial robustness and that of standard generalization. Specifically, training robust models may not only be more resource-consuming, but also lead to a reduction of standard accuracy. We demonstrate that this trade-off between the standard accura…
The paper analyzes adversarial training effects on classification accuracy.
Many applications of machine learning, for example in health care, would benefit from methods that can guarantee privacy of data subjects. Differential privacy (DP) has become established as a standard for protecting learning results. The standard DP algorithms require a single trusted party to have access to the entir…
We present new algorithms for learning Bayesian networks from data with missing values using a data augmentation approach. An exact Bayesian network learning algorithm is obtained by recasting the problem into a standard Bayesian network learning problem without missing data. To the best of our knowledge, this is the f…
X-VAE uses data-adaptive Gaussian priors to improve latent space modeling.
This paper introduces a generalization of Convolutional Neural Networks (CNNs) from low-dimensional grid data, such as images, to graph-structured data. We propose a novel spatial convolution utilizing a random walk to uncover the relations within the input, analogous to the way the standard convolution uses the spatia…
New method improves statistical inference using machine learning-imputed data.
Deep ensembles don't necessarily improve calibration in low data regimes.