Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

3517021,0521,403 · Jun 202019922001200920172026
48 results for data space

The paper explores transferring functions from one data space to another.

problem Approximating a function on a new data set using a learned function from an old data set.
method Transfer learning from one data space to another, focusing on subsets of the target data space.
result Local smoothness of the function and its lifting are related.

We provide the proof that the space of time series data is a Kolmogorov space with T0T_{0}-separation axiom using the loop space of time series data. In our approach we define a cyclic coordinate of intrinsic time scale of time series data after empirical mode decomposition. A spinor field of time series data comes fro…

2016-06-10abs ↗pdf ↗

Empirical mode modeling improves state-space analysis of noisy data.

problem Analyzing nonlinear systems with noisy data.
method Combining empirical mode decomposition with empirical dynamic modeling.
result Empirical mode modeling enhances state-space representations in noisy data.

Develops G-MLKM for better data-target association in constrained spaces.

problem Data-target association problem in constrained spaces with limited sensor information.
method Graph-based multi-layer k-means++ (G-MLKM) method, including MLKM for local space and G-MLKM for general constrained space.
result Improves data-target association accuracy through error correction mechanisms.

The paper proposes a Gaussian mixture model for Hilbert-space-valued data.

problem Challenges in characterizing probability measures for infinite-dimensional random objects.
method Gaussian mixture framework based on kernel mean embeddings.
result The proposed algorithm yields a dense class of approximations in infinite-dimensional spaces.

Linear classifiers in product space forms improve scRNA-seq data classification.

problem Linear classification in products of Euclidean, spherical, and hyperbolic spaces.
method Novel formulations of linear classifiers on Riemannian manifolds, proving expressive power, and formalizing perceptron and SVM classifiers.
result Linear classifiers in product space forms have the same expressive power as in Euclidean space of the same dimension.

A new method improves semi-supervised learning by handling tasks with different attribute spaces.

problem Existing methods assume tasks share the same attribute space, limiting their applicability.
method Meta-learning approach that embeds labeled and unlabeled data in task-specific spaces using neural networks.
result Improves test performance on tasks with small labeled data using unlabeled and various task data.

Generative models for function-valued data in infinite dimensions.

problem Lack of semantics relating discretized data to underlying functional forms.
method Generalized diffusion models to function space, using Gaussian measures on Hilbert spaces.
result Explicit specification of function space allows unconditional and conditional generation of function-valued data.

New method for reducing dimensions of distributional data.

problem Nonlinear sufficient dimension reduction for distribution-on-distribution regression.
method Building universal kernels on metric spaces to characterize conditional independence.
result Method outperforms competing methods in synthetic and real data applications.

Continuous family of elliptic operators' projections maintain Cauchy data spaces.

problem Maintaining Cauchy data spaces for a continuous family of elliptic operators.
method Elementary tools and classical results applied to operator graphs, Sobolev spaces, and Green's formula.
result Orthogonalized Calderón projections form a continuous family of projections.

This paper proposes a spectral clustering algorithm for hyperbolic spaces, improving efficiency over Euclidean methods.

problem Inefficient clustering in Euclidean spaces for complex data structures.
method Developed a spectral clustering algorithm using hyperbolic similarity matrices.
result The algorithm converges at least as fast as Euclidean spectral clustering and performs better on complex datasets.

The paper introduces exponential-wrapped distributions on symmetric spaces for better data modeling.

problem Challenges in statistical modeling due to curvature of data spaces.
method Construction and use of exponential-wrapped distributions on affine locally symmetric spaces.
result Exponential-wrapped distributions on symmetric spaces have useful properties for practical use.

We develop methods to efficiently approximate data in metric spaces without additional assumptions.

problem Efficiently approximating data in metric spaces without imposing structural assumptions.
method Identify discrete modulus of continuity, investigate consistency, propose algorithm, and develop approximation theory.
result Consistent approximation of data in metric spaces without structural assumptions.

Study proves well-posedness and scattering for wave equations on hyperbolic spaces with singular data.

problem Proving well-posedness and scattering for wave equations on hyperbolic spaces with singular initial data.
method Using weak-LpL^{p} spaces and dispersive estimates on Lorentz spaces, the study establishes global well-posedness and exponential asymptotic stability.
result Developed a scattering theory and constructed wave operators in a singular framework.

A large class of vacuum space-times is constructed in dimension 4+1 from hyperboloidal initial data sets which are not small perturbations of empty space data. These space-times are future geodesically complete, smooth up to their future null infinity, and extend as vacuum space-times through their Cauchy horizon. Dime…

2001-06-20abs ↗pdf ↗

Archetypal analysis is a data decomposition method that describes each observation in a dataset as a convex combination of "pure types" or archetypes. These archetypes represent extrema of a data space in which there is a trade-off between features, such as in biology where different combinations of traits provide opti…

2019-01-25abs ↗pdf ↗

Representation costs in data science: Unifying function-space views of parametric methods

problem Analyzing representation costs of parametric data-fitting methods
method Developing a general framework for analyzing representation costs through parameter-space regularizers
result Proving that many natural results hold in this abstract setting, including representer theorems for parametric methods on their native spaces

M-CaStLe discovers causal structures in multivariate space-time data.

problem Challenges in causal graph discovery for high-dimensional gridded data.
method Generalizes CaStLe to multivariate analyses, using local embeddings and pooling spatial replicates.
result More accurately recovers multivariate causal structure and identifies physical dynamics.

Semi-automatic data annotation helps experts label unlabeled samples based on feature space projection.

problem Laborious manual data annotation for machine learning.
method Interactive semi-automatic approach using feature space projection and semi-supervised learning.
result Reduces user annotation effort and improves classification accuracy.

Faced with massive data, is it possible to trade off (statistical) risk, and (computational) space and time? This challenge lies at the heart of large-scale machine learning. Using k-means clustering as a prototypical unsupervised learning problem, we show how we can strategically summarize the data (control space) in …

2016-05-02abs ↗pdf ↗

New method uses Diffusion Maps for latent space modeling of dynamical systems.

problem Building reduced dynamical models from time series data.
method Two rounds of Diffusion Maps on latent coordinates, with lifting back to ambient space.
result Approximation of full state functions in reduced coordinates.

Control data constructed for smooth weak deformation retraction of stratified spaces.

problem Construct control data for smooth weak deformation retraction of stratified spaces.
method Show smooth local triviality with conical fibers, construct control data, use fiber-wise scalar multiplications.
result Obtain neighbourhood smooth weak deformation retraction of stratified spaces.

Kontsevich and Soibelman introduced a notion of orientation data on Calabi-Yau category. It can be viewed as a consistent choice of spin structure on moduli space of objects in the given category. The orientation data plays an important role in Donaldson-Thomas theory. Let X be a projective, simply connected and torsio…

2012-12-16abs ↗pdf ↗

New method for learning with non-Euclidean data using decomposable kernels.

problem Difficulty in using classical kernels for non-Euclidean data.
method Reproducing kernel Krein space (RKKS) methods for kernels that admit a positive decomposition.
result Invariant kernels can be used for learning in non-Euclidean spaces.

Finding rare information hidden in a huge amount of data from the Internet is a necessary but complex issue. Many researchers have studied this issue and have found effective methods to detect anomaly data in low dimensional space. However, as the dimension increases, most of these existing methods perform poorly in de…

2014-05-05abs ↗pdf ↗

In this paper we give an explicit parametrisation of the moduli space of equivariant harmonic maps from a 2-torus to the 3-sphere. As Hitchin proved, a harmonic map of a 2-torus is described by its spectral data, which consists of a hyperelliptic curve together with a pair of differentials and a line bundle. The space …

2020-01-27abs ↗pdf ↗

Machine learning improves planetary space physics by incorporating physical knowledge.

problem Improving performance and interpretability of machine learning models for planetary space physics.
method Building on a previous semi-supervised physics-based classification, the team used varying data and physical information to improve machine learning performance and interpretability.
result Incorporating physical knowledge improves machine learning performance and interpretability, essential for deriving scientific meaning.

Study shows latent space OOD detection isn't a reliable proxy for model performance.

problem Evaluating and interpreting deep learning systems on real-world data.
method Empirical investigation of latent space OOD detection and classification accuracy using SAR datasets.
result OOD detection cannot be used as a proxy measure for model performance.

Deep generative models have made tremendous advances in image and signal representation learning and generation. These models employ the full Euclidean space or a bounded subset as the latent space, whose flat geometry, however, is often too simplistic to meaningfully reflect the manifold structure of the data. In this…

2019-12-20abs ↗pdf ↗