Adaptive feature normalization improves model robustness to extraneous variables.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Electronic Medical Records (EMR) are a rich source of patient information, including measurements reflecting physiologic signs and administered therapies. Identifying which variables are useful in predicting clinical outcomes can be challenging. Advanced algorithms such as deep neural networks were designed to process …
New metric improves latent dynamics inference from neural data.
Both the median-based classifier and the quantile-based classifier are useful for discriminating high-dimensional data with heavy-tailed or skewed inputs. But these methods are restricted as they assign equal weight to each variable in an unregularized way. The ensemble quantile classifier is a more flexible regularize…
Paper introduces a method to generate stable shapes using Grassmann manifolds.
We address the online linear optimization problem with bandit feedback. Our contribution is twofold. First, we provide an algorithm (based on exponential weights) with a regret of order for any finite action set with actions, under the assumption that the instantaneous loss is bounded by 1. This…
Prompted by a recent experiment by Victor Haghani and Richard Dewey, this note generalises the Kelly strategy (optimal for simple investment games with log utility) to a large class of practical utility functions and including the effect of extraneous wealth. A counterintuitive result is proved : for any continuous, co…
Despite its omnipresence in robotics application, the nature of spatial knowledge and the mechanisms that underlie its emergence in autonomous agents are still poorly understood. Recent theoretical work suggests that the concept of space can be grounded by capturing invariants induced by the structure of space in an ag…
Despite its omnipresence in robotics application, the nature of spatial knowledge and the mechanisms that underlie its emergence in autonomous agents are still poorly understood. Recent theoretical works suggest that the Euclidean structure of space induces invariants in an agent's raw sensorimotor experience. We hypot…
Novel approach reduces DNN redundancy and enhances computational efficiency.
We consider the problem of precision matrix estimation where, due to extraneous confounding of the underlying precision matrix, the data are independent but not identically distributed. While such confounding occurs in many scientific problems, our approach is inspired by recent neuroscientific research suggesting that…
Curating labeled training data has become the primary bottleneck in machine learning. Recent frameworks address this bottleneck with generative models to synthesize labels at scale from weak supervision sources. The generative model's dependency structure directly affects the quality of the estimated labels, but select…
Gaussian processes are rich distributions over functions, with generalization properties determined by a kernel function. When used for long-range extrapolation, predictions are particularly sensitive to the choice of kernel parameters. It is therefore critical to account for kernel uncertainty in our predictive distri…
Unified framework for deriving generalization bounds in supervised learning.
Cluster analysis is a field of data analysis that extracts underlying patterns in data. One application of cluster analysis is in text-mining, the analysis of large collections of text to find similarities between documents. We used a collection of about 30,000 tweets extracted from Twitter just before the World Cup st…
A new method for generating replay samples on the fly, optimizing for not forgetting.
Meta framework generates noise to improve multi-attack robustness.
An asymmetric information model is introduced for the situation in which there is a small agent who is more susceptible to the flow of information in the market than the general market participant, and who tries to implement strategies based on the additional information. In this model market participants have access t…
Unified framework detects changes in complex system models.
Graph Convolutional Networks improve prosthetic sensation interpretation.
Alternative wavelet analysis method for financial signals.
Isometry regularizer improves autoencoder performance on manifold learning.
We propose a novel technique for analyzing adaptive sampling called the {\em Simulator}. Our approach differs from the existing methods by considering not how much information could be gathered by any fixed sampling strategy, but how difficult it is to distinguish a good sampling strategy from a bad one given the limit…
The Gaussian curvature of a two-dimensional Riemannian manifold is uniquely determined by the choice of the metric. The formulas for computing the curvature in terms of components of the metric, in isothermal coordinates, involve the Laplacian operator and therefore, the problem of finding a Riemannian metric for a giv…
Various psychological factors affect how individuals express emotions. Yet, when we collect data intended for use in building emotion recognition systems, we often try to do so by creating paradigms that are designed just with a focus on eliciting emotional behavior. Algorithms trained with these types of data are unli…
Method interprets deep learning models using topological data analysis.
There has been great interest recently in applying nonparametric kernel mixtures in a hierarchical manner to model multiple related data samples jointly. In such settings several data features are commonly present: (i) the related samples often share some, if not all, of the mixture components but with differing weight…
New optimal rates for score estimation improve diffusion model performance.
Estimates network structure from node potentials and edge flows under Gaussian injection statistics.
Statistical inference is considered for variables of interest, called primary variables, when auxiliary variables are observed along with the primary variables. We consider the setting of incomplete data analysis, where some primary variables are not observed. Utilizing a parametric model of joint distribution of prima…
VC-PCR improves prediction by clustering correlated variables.
Study on inequalities for multinomial variables.
In this paper, we propose multi-variable LSTM capable of accurate forecasting and variable importance interpretation for time series with exogenous variables. Current attention mechanism in recurrent neural networks mostly focuses on the temporal aspect of data and falls short of characterizing variable importance. To …
Variable importance is central to scientific studies, including the social sciences and causal inference, healthcare, and other domains. However, current notions of variable importance are often tied to a specific predictive model. This is problematic: what if there were multiple well-performing predictive models, and …
A neural network finds causal relationships among latent variables.
Derives derivatives and geometric framework for functions with non-independent variables.
A new distance for mixed-variable, hierarchical datasets with meta variables.
Unified Bayesian Optimisation for mixed variables improves performance.
In this paper, we propose an interpretable LSTM recurrent neural network, i.e., multi-variable LSTM for time series with exogenous variables. Currently, widely used attention mechanism in recurrent neural networks mostly focuses on the temporal aspect of data and falls short of characterizing variable importance. To th…
Random Forest variable importance is improved by class balancing techniques.
Variable selection for Gaussian process models is often done using automatic relevance determination, which uses the inverse length-scale parameter of each input variable as a proxy for variable relevance. This implicitly determined relevance has several drawbacks that prevent the selection of optimal input variables i…
Knoop enhances variable selection with over-parameterization and knockoffs.
A serious problem in learning probabilistic models is the presence of hidden variables. These variables are not observed, yet interact with several of the observed variables. Detecting hidden variables poses two problems: determining the relations to other variables in the model and determining the number of states of …
New method for fitting graphical models with latent variables using regularized conditional likelihood.
CIB compresses variables causally, preserving key causal interactions.
We generalize to the finite-state case the notion of the extreme effect variable that accumulates all the effect of a variant variable observed in changes of another variable . We conduct theoretical analysis and turn the problem of finding of an effect variable into a problem of a simultaneous decomposition…
Discond-VAE separates continuous and discrete factors in data.
A new method selects important variables for clustering from dependency networks.