Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

2805598391,118 · Jun 202019922001200920172026
48 results for unrelated data

Develops a contrastive framework for data-efficient multimodal learning.

problem Expensive training of multimodal generative models requiring related multimodal data.
method Contrastive framework for multimodal learning, distinguishing related from unrelated data.
result Data-efficient multimodal learning on challenging datasets for various VAE models.

Training on mixed distributions improves test performance even when components are unrelated.

problem Improving test performance with mismatched training and test distributions.
method Analyzing mixture distributions with different training and test proportions.
result Distribution shift can be beneficial, improving test performance even when components are unrelated.

We present a novel technique based on deep learning and set theory which yields exceptional classification and prediction results. Having access to a sufficiently large amount of labelled training data, our methodology is capable of predicting the labels of the test data almost always even if the training data is entir…

2019-01-17abs ↗pdf ↗

In this paper we show that two seemingly unrelated problems in economics, the hypothesis of integrability and the hypothesis of additive separability are linked by the absence of curvature of connections on webs naturally associated with each problem.

2009-01-01abs ↗pdf ↗

We show how binary classification methods developed to work on i.i.d. data can be used for solving statistical problems that are seemingly unrelated to classification and concern highly-dependent time series. Specifically, the problems of time-series clustering, homogeneity testing and the three-sample problem are addr…

2012-10-22abs ↗pdf ↗

We present an axiomatic modification of quaternionic quantum mechanics with a possible-worlds semantics capable of predicting essential "nonquantum" features of an observable universe model - the dimensionality and topology of spacetime, the existence, the signature and a specific form of a metric on it, and certain na…

2007-02-28abs ↗pdf ↗

We propose a general-purpose approach to discovering active learning (AL) strategies from data. These strategies are transferable from one domain to another and can be used in conjunction with many machine learning models. To this end, we formalize the annotation process as a Markov decision process, design universal s…

2018-10-09abs ↗pdf ↗

Correlated component analysis as proposed by Dmochowski et al. (2012) is a tool for investigating brain process similarity in the responses to multiple views of a given stimulus. Correlated components are identified under the assumption that the involved spatial networks are identical. Here we propose a hierarchical pr…

2018-02-07abs ↗pdf ↗

We find maximal representatives within equivalence classes of metric spheres. For Ahlfors regular spheres these are uniquely characterized by satisfying the seemingly unrelated notions of Sobolev-to-Lipschitz property, or volume rigidity. We also apply our construction to solutions of the Plateau problem in metric spac…

2019-09-23abs ↗pdf ↗

We introduce generative adversarial models in which the discriminator is replaced by a calibrated (non-differentiable) classifier repeatedly enhanced by domain relevant features. The role of the classifier is to prove that the actual and generated data differ over a controlled semantic space. We demonstrate that such m…

2019-10-07abs ↗pdf ↗

Recently, a number of statistical problems have found an unexpected solution by inspecting them through a "modal point of view". These include classical tasks such as clustering or regression. This has led to a renewed interest in estimation and inference for the mode. This paper offers an extensive survey of the tradi…

2018-07-08abs ↗pdf ↗

Investor attention is an important concept in behavioral finance. Many articles have conducted cross-disciplinary research leading by this concept. In this paper, we use data extraction technology to collect a large number of Baidu Index keyword search volume data. After analyzing the data, we draw a conclusion that ha…

2019-11-02abs ↗pdf ↗

Boosts barely robust learners to be more adversarially robust.

problem Learning predictors robust to small perturbations on a small fraction of data.
method Oracle-efficient algorithm for robustness with larger perturbation set.
result Qualitative and quantitative equivalence between strongly robust and barely robust learning.

New method for rolling bodies on inclined planes, with applications to rescue operations.

problem Constructing solid bodies rolling along curves on inclined planes.
method Rigorous existence theorems and connections to maritime rescue operations.
result Comprehensive existence theorems for rolling bodies on inclined planes.

Neural recordings are nonstationary time series, i.e. their properties typically change over time. Identifying specific changes, e.g. those induced by a learning task, can shed light on the underlying neural processes. However, such changes of interest are often masked by strong unrelated changes, which can be of physi…

2013-01-25abs ↗pdf ↗

The application of deep learning (DL) models to the decoding of cognitive states from whole-brain functional Magnetic Resonance Imaging (fMRI) data is often hindered by the small sample size and high dimensionality of these datasets. Especially, in clinical settings, where patient data are scarce. In this work, we demo…

2019-07-02abs ↗pdf ↗

A generalized cusp CC is diffeomorphic to [0,)[0,\infty) times a closed Euclidean manifold. Geometrically CC is the quotient of a properly convex domain by a lattice, ΓΓ, in one of a family of affine groups G(ψ)G(ψ), parameterized by a point ψψ in the (dual closed) Weyl chamber for SL(n+1,R)SL(n+1,\mathbb{R}), and ΓΓ determi…

2017-10-09abs ↗pdf ↗

Powerful generative models, particularly in Natural Language Modelling, are commonly trained by maximizing a variational lower bound on the data log likelihood. These models often suffer from poor use of their latent variable, with ad-hoc annealing factors used to encourage retention of information in the latent variab…

2018-06-12abs ↗pdf ↗

This paper diagnoses factor-model pricing errors using a new method.

problem Measuring pricing errors in factor models with general characteristic axes.
method Developed a method to measure factor-model pricing errors as bridge-alpha curves, using a predetermined characteristic order and prefix portfolios.
result Adding a counterpart factor flips the curve's sign on every axis, but only HML and CMA overcorrect enough to be rejected.

We introduce a new class of quantum enhancements we call biquandle brackets, which are customized skein invariants for biquandle colored links.Quantum enhancements of biquandle counting invariants form a class of knot and link invariants that includes biquandle cocycle invariants and skein invariants such as the HOMFLY…

2015-08-26abs ↗pdf ↗

When labeled data is scarce for a specific target task, transfer learning often offers an effective solution by utilizing data from a related source task. However, when transferring knowledge from a less related source, it may inversely hurt the target performance, a phenomenon known as negative transfer. Despite its p…

2018-11-24abs ↗pdf ↗

UMFI improves feature importance methods by reducing runtime and enhancing performance.

problem Improving feature importance methods to better explain causal and associative relationships in data.
method Introducing UMFI, which uses dependence removal techniques from AI fairness literature.
result UMFI outperforms MCI, especially in complex data scenarios, and reduces runtime from exponential to super-linear.

Accurate goodness-of-fit tests for the extreme tails of empirical distributions is a very important issue, relevant in many contexts, including geophysics, insurance, and finance. We have derived exact asymptotic results for a generalization of the large-sample Kolmogorov-Smirnov test, well suited to testing these extr…

2012-07-31abs ↗pdf ↗

GRASP removes spurious correlations in fine-tuned models, improving task performance and reducing bias.

problem Fine-tuned models can latch onto spurious correlations, leading to bias and reduced generalization.
method GRASP identifies and removes spurious correlations from model weights without removing latent factors.
result GRASP significantly reduces bias and improves task performance in various fine-tuning tasks.

New product structures encode superintegrable Hamiltonian systems in Euclidean spaces.

problem Encoding superintegrable Hamiltonian systems using product structures.
method Introducing commutative and associative product structures on Euclidean spaces of dimension at least three, satisfying specific conditions.
result All abundant superintegrable Hamiltonian systems on Euclidean space of dimension at least three arise from these product structures.

Online shopping caters to the needs of millions of users daily. Search, recommendations, personalization have become essential building blocks for serving customer needs. Efficacy of such systems is dependent on a thorough understanding of products and their representation. Multiple information sources and data types p…

2019-06-28abs ↗pdf ↗

Neural networks struggle with extrapolation, but a new framework allows them to learn counterfactual invariances.

problem Neural networks' inability to extrapolate beyond training data distribution.
method Introduces a learning framework that allows neural networks to extrapolate over group transformations based on counterfactual invariances.
result Neural networks can learn counterfactual invariances from a single environment, overcoming their limitations in extrapolation.

Neural score matching improves high-dimensional causal inference by using neural networks for balancing scores.

problem Impracticality of traditional matching methods in high-dimensional datasets due to the curse of dimensionality.
method Develops neural networks to create non-trivial, multivariate balancing scores for high-dimensional causal inference.
result Neural score matching outperforms other methods in treatment effect estimation and reducing imbalance on high-dimensional datasets.

HCL learns shared and modality-specific latent representations for multimodal data.

problem Binary shared-private decomposition inadequately represents shared information across subsets of modalities.
method Hierarchical Contrastive Learning framework combining latent-variable formulation, structural sparsity, and contrastive objective.
result HCL accurately recovers hierarchical structure and improves predictive performance on multimodal data.