Current economic theories miss most of economic dynamics.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New methods explain NE embeddings by identifying key variables.
Neurons in higher cortical areas, such as the prefrontal cortex, are known to be tuned to a variety of sensory and motor variables. The resulting diversity of neural tuning often obscures the represented information. Here we introduce a novel dimensionality reduction technique, demixed principal component analysis (dPC…
New methods reveal colored Jones polynomials from quantum R-matrices and knot invariants.
Dual random fields improve mineral potential predictions.
AI systems that explain their decisions can be monitored for harmful intentions.
Medical imaging machine learning algorithms are usually evaluated on a single dataset. Although training and testing are performed on different subsets of the dataset, models built on one study show limited capability to generalize to other studies. While database bias has been recognized as a serious problem in the co…
Adversarial attacks found to be effective on code models.
Machine learning uncovers hidden patterns in Calabi-Yau hypersurfaces.
Early detection of breast cancer can increase treatment efficiency. Architectural Distortion (AD) is a very subtle contraction of the breast tissue and may represent the earliest sign of cancer. Since it is very likely to be unnoticed by radiologists, several approaches have been proposed over the years but none using …
IUPM monitors machine learning models under gradual shifts using optimal transport and active labeling.
Supervised training of deep learning models requires large labeled datasets. There is a growing interest in obtaining such datasets for medical image analysis applications. However, the impact of label noise has not received sufficient attention. Recent studies have shown that label noise can significantly impact the p…
We present a supervised-learning algorithm from graph data (a set of graphs) for arbitrary twice-differentiable loss functions and sparse linear models over all possible subgraph features. To date, it has been shown that under all possible subgraph features, several types of sparse learning, such as Adaboost, LPBoost, …
NUTS mixing time scales as d^(1/4) for Gaussian distributions.
Celiac Disease (CD) is a chronic autoimmune disease that affects the small intestine in genetically predisposed children and adults. Gluten exposure triggers an inflammatory cascade which leads to compromised intestinal barrier function. If this enteropathy is unrecognized, this can lead to anemia, decreased bone densi…
Network operators are generally aware of common attack vectors that they defend against. For most networks the vast majority of traffic is legitimate. However new attack vectors are continually designed and attempted by bad actors which bypass detection and go unnoticed due to low volume. One strategy for finding such …
The state price density of a basket, even under uncorrelated Black-Scholes dynamics, does not allow for a closed from density. (This may be rephrased as statement on the sum of lognormals and is especially annoying for such are used most frequently in Financial and Actuarial Mathematics.) In this note we discuss short …
Sales data in a commodity market (supermarket sales to consumers) has been analysed by studying the fluctuation spectrum and noise correlations. Three related products (ketchup, mayonnaise and curry sauce) have been analysed. Most noise in sales is caused by promotions, but here we focus on the fluctuations in baseline…
Researchers explore valuations on polyhedra and topological arrangements without imposing algebraic structures.
In knot concordance three genera arise naturally, g(K), g_4(K), and g_c(K): these are the classical genus, the 4-ball genus, and the concordance genus, defined to be the minimum genus among all knots concordant to K. Clearly 0 <= g_4(K) <= g_c(K) <= g(K). Casson and Nakanishi gave examples to show that g_4(K) need not …
Polyhedra can mimic constant curvature surfaces, even with self-intersections.
We show that the cost of market orders and the profit of infinitesimal market-making or -taking strategies can be expressed in terms of directly observable quantities, namely the spread and the lag-dependent impact function. Imposing that any market taking or liquidity providing strategies is at best marginally profita…
SPECTRE defends against backdoor attacks by amplifying corrupted data's spectral signature.
The paper explores symplectic connections on homogeneous spaces, finding a unique invariant connection.
The paper shows how ignoring temporal context in recommender systems evaluation leads to false confidence, proposing a method to embed temporal context.
New Bayesian method for sparse multidimensional item response theory.
This work studies the statistical performance of Sinkhorn iterations in estimating Schrödinger bridges.
In the noisy tensor completion problem we observe entries (whose location is chosen uniformly at random) from an unknown tensor . We assume that is entry-wise close to being rank . Our goal is to fill in its missing entries using as few observations as possible. Let $n = \max(n…
Graph auto-encoders improve financial clustering using news and stock data.
This work uncovers how model and data biases interact to cause unfairness in fraud detection.
Bayesian inference reconstructs external potentials in DFT for many-particle systems.
Statistical inference is considered for variables of interest, called primary variables, when auxiliary variables are observed along with the primary variables. We consider the setting of incomplete data analysis, where some primary variables are not observed. Utilizing a parametric model of joint distribution of prima…
VC-PCR improves prediction by clustering correlated variables.
Study on inequalities for multinomial variables.
In this paper, we propose multi-variable LSTM capable of accurate forecasting and variable importance interpretation for time series with exogenous variables. Current attention mechanism in recurrent neural networks mostly focuses on the temporal aspect of data and falls short of characterizing variable importance. To …
Variable importance is central to scientific studies, including the social sciences and causal inference, healthcare, and other domains. However, current notions of variable importance are often tied to a specific predictive model. This is problematic: what if there were multiple well-performing predictive models, and …
A neural network finds causal relationships among latent variables.
Derives derivatives and geometric framework for functions with non-independent variables.
A new distance for mixed-variable, hierarchical datasets with meta variables.
Extends effect variable concept to finite states for web search evaluation.
Unified Bayesian Optimisation for mixed variables improves performance.
A method to assess variable importance in complex predictive models.
In this paper, we propose an interpretable LSTM recurrent neural network, i.e., multi-variable LSTM for time series with exogenous variables. Currently, widely used attention mechanism in recurrent neural networks mostly focuses on the temporal aspect of data and falls short of characterizing variable importance. To th…
Random Forest variable importance is improved by class balancing techniques.
Variable selection for Gaussian process models is often done using automatic relevance determination, which uses the inverse length-scale parameter of each input variable as a proxy for variable relevance. This implicitly determined relevance has several drawbacks that prevent the selection of optimal input variables i…
Knoop enhances variable selection with over-parameterization and knockoffs.
A serious problem in learning probabilistic models is the presence of hidden variables. These variables are not observed, yet interact with several of the observed variables. Detecting hidden variables poses two problems: determining the relations to other variables in the model and determining the number of states of …
New method for fitting graphical models with latent variables using regularized conditional likelihood.