Countries tend to diversify their exports by entering products that are related to their current exports. Yet this average behavior is not representative of every diversification path. In this paper, we introduce a method to identify periods when countries enter unrelated products. We analyze the economic diversificati…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study shows volume and genus unrelated for hyperbolic fibred knots.
Develops a contrastive framework for data-efficient multimodal learning.
Study on fibered knots in 3-manifolds, proving unrelated volume and genus.
Training on mixed distributions improves test performance even when components are unrelated.
New model improves inference on asset market durations.
Quasi-orthonormal encoding reduces high dimensionality for categorical data.
We present a novel technique based on deep learning and set theory which yields exceptional classification and prediction results. Having access to a sufficiently large amount of labelled training data, our methodology is capable of predicting the labels of the test data almost always even if the training data is entir…
In this paper we show that two seemingly unrelated problems in economics, the hypothesis of integrability and the hypothesis of additive separability are linked by the absence of curvature of connections on webs naturally associated with each problem.
DERWENT learns paths for distant transfer learning via deep random walk.
We show how binary classification methods developed to work on i.i.d. data can be used for solving statistical problems that are seemingly unrelated to classification and concern highly-dependent time series. Specifically, the problems of time-series clustering, homogeneity testing and the three-sample problem are addr…
We present an axiomatic modification of quaternionic quantum mechanics with a possible-worlds semantics capable of predicting essential "nonquantum" features of an observable universe model - the dimensionality and topology of spacetime, the existence, the signature and a specific form of a metric on it, and certain na…
Confounding variables are a well known source of nuisance in biomedical studies. They present an even greater challenge when we combine them with black-box machine learning techniques that operate on raw data. This work presents two case studies. In one, we discovered biases arising from systematic errors in the data g…
Reviews six finance topics, including 'radical complexity'.
We propose a general-purpose approach to discovering active learning (AL) strategies from data. These strategies are transferable from one domain to another and can be used in conjunction with many machine learning models. To this end, we formalize the annotation process as a Markov decision process, design universal s…
Novel framework for quantifying data distribution values.
Correlated component analysis as proposed by Dmochowski et al. (2012) is a tool for investigating brain process similarity in the responses to multiple views of a given stimulus. Correlated components are identified under the assumption that the involved spatial networks are identical. Here we propose a hierarchical pr…
We find maximal representatives within equivalence classes of metric spheres. For Ahlfors regular spheres these are uniquely characterized by satisfying the seemingly unrelated notions of Sobolev-to-Lipschitz property, or volume rigidity. We also apply our construction to solutions of the Plateau problem in metric spac…
We introduce generative adversarial models in which the discriminator is replaced by a calibrated (non-differentiable) classifier repeatedly enhanced by domain relevant features. The role of the classifier is to prove that the actual and generated data differ over a controlled semantic space. We demonstrate that such m…
Recently, a number of statistical problems have found an unexpected solution by inspecting them through a "modal point of view". These include classical tasks such as clustering or regression. This has led to a renewed interest in estimation and inference for the mode. This paper offers an extensive survey of the tradi…
Investor attention is an important concept in behavioral finance. Many articles have conducted cross-disciplinary research leading by this concept. In this paper, we use data extraction technology to collect a large number of Baidu Index keyword search volume data. After analyzing the data, we draw a conclusion that ha…
Boosts barely robust learners to be more adversarially robust.
Most generative models for clustering implicitly assume that the number of data points in each cluster grows linearly with the total number of data points. Finite mixture models, Dirichlet process mixture models, and Pitman--Yor process mixture models make this assumption, as do all other infinitely exchangeable cluste…
Most generative models for clustering implicitly assume that the number of data points in each cluster grows linearly with the total number of data points. Finite mixture models, Dirichlet process mixture models, and Pitman--Yor process mixture models make this assumption, as do all other infinitely exchangeable cluste…
New method for rolling bodies on inclined planes, with applications to rescue operations.
Neural recordings are nonstationary time series, i.e. their properties typically change over time. Identifying specific changes, e.g. those induced by a learning task, can shed light on the underlying neural processes. However, such changes of interest are often masked by strong unrelated changes, which can be of physi…
The application of deep learning (DL) models to the decoding of cognitive states from whole-brain functional Magnetic Resonance Imaging (fMRI) data is often hindered by the small sample size and high dimensionality of these datasets. Especially, in clinical settings, where patient data are scarce. In this work, we demo…
TQFT signatures linked to trace fields of knots.
Mind wandering (MW) is a ubiquitous phenomenon which reflects a shift in attention from task-related to task-unrelated thoughts. There is a need for intelligent interfaces that can reorient attention when MW is detected due to its detrimental effects on performance and productivity. In this paper, we propose a deep lea…
A generalized cusp is diffeomorphic to times a closed Euclidean manifold. Geometrically is the quotient of a properly convex domain by a lattice, , in one of a family of affine groups , parameterized by a point in the (dual closed) Weyl chamber for , and determi…
Powerful generative models, particularly in Natural Language Modelling, are commonly trained by maximizing a variational lower bound on the data log likelihood. These models often suffer from poor use of their latent variable, with ad-hoc annealing factors used to encourage retention of information in the latent variab…
This paper diagnoses factor-model pricing errors using a new method.
We introduce a new class of quantum enhancements we call biquandle brackets, which are customized skein invariants for biquandle colored links.Quantum enhancements of biquandle counting invariants form a class of knot and link invariants that includes biquandle cocycle invariants and skein invariants such as the HOMFLY…
When labeled data is scarce for a specific target task, transfer learning often offers an effective solution by utilizing data from a related source task. However, when transferring knowledge from a less related source, it may inversely hurt the target performance, a phenomenon known as negative transfer. Despite its p…
UMFI improves feature importance methods by reducing runtime and enhancing performance.
Accurate goodness-of-fit tests for the extreme tails of empirical distributions is a very important issue, relevant in many contexts, including geophysics, insurance, and finance. We have derived exact asymptotic results for a generalization of the large-sample Kolmogorov-Smirnov test, well suited to testing these extr…
Local explanation frameworks aim to rationalize particular decisions made by a black-box prediction model. Existing techniques are often restricted to a specific type of predictor or based on input saliency, which may be undesirably sensitive to factors unrelated to the model's decision making process. We instead propo…
Survey classifies Clustered Federated Learning into three types of approaches.
GRASP removes spurious correlations in fine-tuned models, improving task performance and reducing bias.
New product structures encode superintegrable Hamiltonian systems in Euclidean spaces.
Online shopping caters to the needs of millions of users daily. Search, recommendations, personalization have become essential building blocks for serving customer needs. Efficacy of such systems is dependent on a thorough understanding of products and their representation. Multiple information sources and data types p…
Cluster analysis is a field of data analysis that extracts underlying patterns in data. One application of cluster analysis is in text-mining, the analysis of large collections of text to find similarities between documents. We used a collection of about 30,000 tweets extracted from Twitter just before the World Cup st…
Over the years, there has been growing interest in using Machine Learning techniques for biomedical data processing. When tackling these tasks, one needs to bear in mind that biomedical data depends on a variety of characteristics, such as demographic aspects (age, gender, etc) or the acquisition technology, which migh…
K-priors enable quick adaptation with minimal retraining.
Neural networks struggle with extrapolation, but a new framework allows them to learn counterfactual invariances.
Linking facts across documents is a challenging task, as the language used to express the same information in a sentence can vary significantly, which complicates the task of multi-document summarization. Consequently, existing approaches heavily rely on hand-crafted features, which are domain-dependent and hard to cra…
Neural score matching improves high-dimensional causal inference by using neural networks for balancing scores.
HCL learns shared and modality-specific latent representations for multimodal data.