Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

3336671,0001,333 · Jun 202019922001200920172026
48 results for Rich data structures

A fundamental challenge in developing high-impact machine learning technologies is balancing the need to model rich, structured domains with the ability to scale to big data. Many important problem areas are both richly structured and large scale, from social and biological networks, to knowledge graphs and the Web, to…

2015-05-17abs ↗pdf ↗

Generative Neuro-Symbolic model learns from raw data with rich conceptual representations.

problem Learning rich, general-purpose conceptual representations from raw perceptual inputs.
method Generative Neuro-Symbolic (GNS) model combining symbolic and neural network approaches.
result Model learns from raw data and generalizes to 4 unique tasks.

Develops a deep learning architecture for rich-item recommendations.

problem Rich data structures with multiple entity types and side-information.
method General formulation, multiple graph-CNN based architecture (AL-GCN), ranking metric pAp@k.
result 5-6% points more accurate than production models in real-world applications.

Compositional structures between parts and objects are inherent in natural scenes. Modeling such compositional hierarchies via unsupervised learning can bring various benefits such as interpretability and transferability, which are important in many downstream tasks. In this paper, we propose the first deep latent vari…

2019-10-21abs ↗pdf ↗

SentenceMIM learns rich latent representations for variable-length language data.

problem Challenges in learning VAEs for variable-length language data, especially posterior collapse.
method Probabilistic auto-encoder trained with Mutual Information Machine (MIM) learning.
result SentenceMIM learns informative latent representations with high mutual information.

Study binary choice with asymmetric loss, offering simple solutions.

problem Binary choice with asymmetric loss in data-rich environments.
method Loss-based reweighting of logistic regression or machine learning techniques.
result Valid decisions on binary outcomes with general loss functions.

Structured prediction provides a general framework to deal with supervised problems where the outputs have semantically rich structure. While classical approaches consider finite, albeit potentially huge, output spaces, in this paper we discuss how structured prediction can be extended to a continuous scenario. Specifi…

2018-06-26abs ↗pdf ↗

Graph neural network predicts new bank client interactions using transaction data.

problem Predicting new interactions in the network of bank clients.
method Proposes a graph neural network model that uses both network topology and time-series data.
result The model outperforms existing approaches in link prediction and credit scoring.

Many data-rich industries are interested in the efficient discovery and modelling of structures underlying large data sets, as it allows for the fast triage and dimension reduction of large volumes of data embedded in high dimensional spaces. The modelling of these underlying structures is also beneficial for the creat…

2019-09-27abs ↗pdf ↗

Understanding and interacting with everyday physical scenes requires rich knowledge about the structure of the world, represented either implicitly in a value or policy function, or explicitly in a transition model. Here we introduce a new class of learnable models--based on graph networks--which implement an inductive…

2018-06-04abs ↗pdf ↗

In many situations, we need to build and deploy separate models in related environments with different data qualities. For example, an environment with strong observation equipments (e.g., intensive care units) often provides high-quality multi-modal data, which are acquired from multiple sensory devices and have rich-…

2018-09-06abs ↗pdf ↗

Multi-task learning (MTL) improves prediction performance in different contexts by learning models jointly on multiple different, but related tasks. Network data, which are a priori data with a rich relational structure, provide an important context for applying MTL. In particular, the explicit relational structure imp…

2014-11-10abs ↗pdf ↗

The kernel exponential family is a rich class of distributions, which can be fit efficiently and with statistical guarantees by score matching. Being required to choose a priori a simple kernel such as the Gaussian, however, limits its practical applicability. We provide a scheme for learning a kernel parameterized by …

2018-11-20abs ↗pdf ↗

We present Blitzkriging, a new approach to fast inference for Gaussian processes, applicable to regression, optimisation and classification. State-of-the-art (stochastic) inference for Gaussian processes on very large datasets scales cubically in the number of 'inducing inputs', variables introduced to factorise the mo…

2015-10-27abs ↗pdf ↗

We present a first procedure that can estimate -- with statistical consistency guarantees -- any local-maxima of a density, under benign distributional conditions. The procedure estimates all such local maxima, or modal-sets\textit{modal-sets}, of any bounded shape or dimension, including usual point-modes. In practice, modal-…

2016-06-13abs ↗pdf ↗

We define hypersymplectic structures on Lie algebroids recovering, as particular cases, all the classical results and examples of hypersymplectic structures on manifolds. We prove a 1-1 correspondence theorem between hypersymplectic structures and (pseudo-)hyperkähler structures. We show that the hypersymplectic framew…

2013-04-15abs ↗pdf ↗

New framework extracts useful information from tensor data with structural properties.

problem Extract useful information from tensor data with structural properties.
method Proposed an additive tensor decomposition (ATD) framework and an ADMM algorithm to solve the high dimensional optimization problem.
result Versatile and effective framework demonstrated in simulations and real medical image analysis.

New metric measures dynamical richness without relying on accuracy.

problem Lack of a reliable metric for measuring dynamical richness.
method Developed a computationally efficient, performance-independent metric based on low-rank bias.
result Metric recovers neural collapse as a special case and captures known transitions without accuracy.

Exact solutions reveal how unbalanced initializations promote rapid feature learning in neural networks.

problem Understanding how neural networks efficiently extract features from data.
method Deriving exact solutions to a minimal model of neural networks transitioning between lazy and rich learning regimes.
result Unbalanced layer-specific initialization variances and learning rates determine the degree of feature learning.

Sparsity-based models and techniques have been exploited in many signal processing and imaging applications. Data-driven methods based on dictionary and sparsifying transform learning enable learning rich image features from data, and can outperform analytical models. In particular, alternating optimization algorithms …

2018-05-31abs ↗pdf ↗

Physics-informed GCRL tackles sparse feedback learning with hybrid dynamics.

problem Sparse feedback learning with high-dimensional, hybrid, or contact-dependent dynamics.
method Introduces physics-informed inductive biases into goal-conditioned value learning.
result Contact-rich manipulation tasks degrade existing Pi-GCRL methods.

We present a riemannian structure on the disk that has a remarkably rich structure. Geodesics are hypocycloids and the (negative of the) laplacian has integer spectrum with multiplicity the Dirichlet divisor function. Eigenfunctions of the laplacian are orthogonal polynomials naturally suited to the analysis of acousti…

2016-03-21abs ↗pdf ↗

Deep Discrete Encoders (DDEs) tackle interpretable generative models for rich data with discrete latent layers.

problem Overparametrized, non-identifiable, and uninterpretable deep generative models in high-stakes applications.
method Directed graphical model with multiple binary latent layers, transparent identifiability conditions, scalable estimation pipeline.
result Transparent identifiability conditions and scalable estimation pipeline for interpretable DDEs.

Quantum models can approximate any function if data encoding allows for a rich enough frequency spectrum.

problem Theoretical properties of quantum machine learning models, particularly their expressive power.
method Investigated how data encoding affects the expressive power of parametrized quantum circuits.
result Quantum models can access increasingly rich frequency spectra by repeating data encoding gates, potentially making them universal function approximators.

In this note, we unveil homotopy-rich algebraic structures generated by the Atiyah classes relative to a Lie pair (L,A)(L,A) of algebroids. In particular, we prove that the quotient L/AL/A of such a pair admits an essentially canonical homotopy module structure over the Lie algebroid AA, which we call Kapranov module.

2012-11-15abs ↗pdf ↗

Effective field theories with explicit Lorentz violation are intimately linked to Riemann-Finsler geometry. The quadratic single-fermion restriction of the Standard-Model Extension provides a rich source of pseudo-Riemann-Finsler spacetimes and Riemann-Finsler spaces. An example is presented that is constructed from a …

2011-04-28abs ↗pdf ↗

Deep generative models (DGMs) have shown promise in image generation. However, most of the existing work learn the model by simply optimizing a divergence between the marginal distributions of the model and the data, and often fail to capture the rich structures and relations in multi-object images. Human knowledge is …

2019-06-10abs ↗pdf ↗

We define a copula process which describes the dependencies between arbitrarily many random variables independently of their marginal distributions. As an example, we develop a stochastic volatility model, Gaussian Copula Process Volatility (GCPV), to predict the latent standard deviations of a sequence of random varia…

2010-06-07abs ↗pdf ↗

The Hilbert manifold ΣΣ consisting of positive invertible (unitized) Hilbert-Schmidt operators has a rich structure and geometry. The geometry of unitary orbits ΩΣΩ\subset Σ is studied from the topological and metric viewpoints: we seek for conditions that ensure the existence of a smooth local structure for the set $…

2008-08-07abs ↗pdf ↗