Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

11.1%22.1%33.2%44.2% · Jun 202019922001200920182026
48 results for data relationships

LoCEC classifies user relationships in large social networks, addressing sparsity issues.

problem Sparse relationship feature and label data in real social platforms.
method Local Community-based Edge Classification (LoCEC) framework with three-phase processing.
result Effective and efficient classification of user relationships in large-scale networks.

Paper discovers sub-interval relationships in time series data.

problem Finding complex patterns of relationships between time series data.
method Proposes a novel approach to find most interesting sub-interval relationships (SIR) in a pair of time series.
result Discovered statistically significant sub-interval relationships with physical interpretation.

Relational Autoencoder improves feature extraction by considering data relationships.

problem Feature extraction from high-dimensional data fails to consider data relationships.
method Proposes a Relation Autoencoder model that considers both features and relationships.
result Considering data relationships generates more robust features with lower error rates.

Proposes a method to generate realistic counterfactuals by learning relationships.

problem Counterfactual explanations often ignore intrinsic relationships between data attributes.
method Uses a variational auto-encoder to learn relationships and perturb the latent space.
result The model preserves relationships and generates realistic counterfactuals.

Bayesian method discovers local causal relationships among genes from gene expression data.

problem Discovering gene regulatory relationships from gene expression data.
method Bayesian approach scoring covariance structures for triplets of normally distributed variables, incorporating background knowledge as priors.
result Stable and conservative posterior probability estimates of local causal structures.

New method discovers useful structure in multi-view data for clinical applications.

problem Detecting slow bleeding in patients monitored for central venous pressure.
method Proposes a method to characterize globally nonlinear multi-view relationships using a mixture of linear relationships.
result Demonstrates the potential to find useful structure in data that is hard to find with single-view or current multi-view methods.

Modeling lead-lag relationship between two text corpora for improved topic modeling.

problem Recognizing the relationship between multiple text corpora for better topic modeling.
method Proposed a jointly dynamic topic model and embedding extension for large-scale text corpus.
result The proposed model can well recognize the lead-lag relationship between two text corpora and improve topic learning.

Algorithm finds significant sub-interval relationships in time series data.

problem Finding meaningful interactions in small sub-intervals of time series data.
method Fast-optimal guaranteed algorithm for sub-interval relationships (SIR).
result Algorithm identifies SIR relationships that are prominent in specific sub-intervals.

We employ the Bayesian framework to define a cointegration measure aimed to represent long term relationships between time series. For visualization of these relationships we introduce a dissimilarity matrix and a map based on the Sorting Points Into Neighborhoods (SPIN) technique, which has been previously used to ana…

2007-01-05abs ↗pdf ↗

Socio-economic inequality is characterized from data using various indices. The Gini (gg) index, giving the overall inequality is the most common one, while the recently introduced Kolkata (kk) index gives a measure of 1k1-k fraction of population who possess top kk fraction of wealth in the society. Here, we show t…

2016-06-10abs ↗pdf ↗

Rhino learns causal relationships from time series data with history-dependent noise.

problem Discovering causal relationships from time series data with non-linear relations, instantaneous effects, and history-dependent noise.
method Combines vector auto-regression, deep learning, and variational inference.
result Demonstrates better causal relationship discovery performance compared to baselines.

Capsule networks improve temporal data understanding, achieving 96.21% ECG accuracy.

problem Improving temporal data understanding with capsule networks.
method Generated capsules along temporal and channel dimensions, learning contrasting relationships.
result Achieved 96.21% accuracy on ECG signal beat categories, surpassing state-of-the-art.

An approach for learning ancestral causal relationships in high dimensions, validated on human genome-wide data.

problem Learning ancestral causal relationships in high-dimensional biological data.
method Supervised learning approach with discrete indicators treated as labels, scalable to large problems.
result The approach is highly effective and scalable to the human genome-wide setting, robust to perturbations of input information.

Study lead-lag relationships in foreign exchange markets using three approaches.

problem Lack of research on lead-lag relationships in foreign exchange markets.
method Three approaches: lagged correlations, lagged partial correlations, and Granger causality.
result Statistically significant lead-lag relationships found in some exchange rate pairs.

New algorithm discovers causal relationships from observational data efficiently.

problem Inferring direct causal parents from a large set of variables.
method Orthogonal structure search approach, scaling to large graphs, guarantees for nonlinear relationships.
result Significant improvements over existing methods in causal discovery from observational data.

Generative approach improves clustering with user-defined relationships.

problem Improving unsupervised clustering performance with user-defined relationships.
method Proposes a generative approach to model joint distribution of data points given user-defined pairwise relations.
result Proposed approach results in a closed-form solution for updated parameters in standard distribution forms.

Discover novel multivariate relationships in time series data.

problem Capturing novel relationships between time series in complex systems.
method Introducing multipoles as linear relationships among more than two time series, identifying them as cliques of negative correlations in a correlation network.
result Almost all multipoles can be efficiently found using a clique-enumeration approach.

A new method treats all variables equally in fitting data.

problem Fitting relationships to data with multiple variables, especially when dependent and independent variables are not clearly defined.
method A general method treating all variables impartially, using geometric mean functional relationships and correlation.
result The method provides coefficients that are easily calculated from covariances or correlations, making it scale-invariant and applicable to various units.

Recently the interest of researchers has shifted from the analysis of synchronous relationships of financial instruments to the analysis of more meaningful asynchronous relationships. Both of those analyses are concentrated only on Pearson's correlation coefficient and thus intraday lead-lag relationships associated wi…

2014-02-16abs ↗pdf ↗

NTP struggles to learn relationships without increased exploration.

problem NTP's performance in extracting true relationships among data is poor.
method Created synthetic logical datasets with injected relationships to test NTP's performance and identify algorithmic issues.
result Increasing exploration in NTP's algorithm improves its performance in recovering relationships.

Estimates lead-lag relationships in high-frequency financial markets without interpolation.

problem Lag relationships in high-frequency financial markets with non-synchronous data.
method Proposes a novel estimation procedure for scale-by-scale lead-lag relationships.
result Identifies two types of lead-lag relationships at different time scales.

New sampling strategy preserves relationships in multivariate scientific data.

problem Reducing storage and enabling efficient multivariate analyses on large scientific data.
method Uses principal component analysis for multivariate data and combines with existing univariate sampling algorithms.
result Efficacy demonstrated on real-world data sets, showing data reduction and multivariate analysis ease.

Method detects lead-lag relationships in multivariate time series.

problem Discovering lead-lag relationships in multivariate time series.
method Clustering-driven methodology using sliding window and various clustering techniques.
result Robust lead-lag estimates across clusters enhance consistent relationships identification.

This paper proposes a method to reveal task relationships in multi-task learning models using sparse graphs.

problem Understanding the underlying task relationships in multi-task learning models.
method Proposes a bilevel formulation of multi-task learning that induces sparse graphs.
result The method improves interpretability of multi-task learning models without sacrificing generalization performance.

This study examines lead-lag relationships in Chinese futures markets using high-frequency data.

problem Understanding high-frequency trading dynamics and information flow in futures markets.
method High-frequency tick-by-tick data analysis of lead-lag relationships between different maturity futures contracts.
result The near-month futures lead longer-dated contracts by one tick, with a negative feedback effect on the leading asset.

Proposes a method to create robust models that adapt to target domains.

problem Creating reliable models when training and target distributions differ.
method Integrates prior knowledge of data generating process to learn stable relationships.
result The Surgery Estimator finds stable relationships in more scenarios than previous methods.

Double autoencoder Ae2IAe^2I improves missing value imputation in recommender systems.

problem Imputing missing values in tables using row-row and column-column relationships.
method Simultaneously uses row-row and column-column relationships through a double autoencoder.
result Ae2IAe^2I outperforms state-of-the-art models in recommender systems.

Novel framework detects lead-lag relationships in Chinese A-share market.

problem Detecting lead-lag relationships in the Chinese A-share market.
method Two-stage framework: long-term coupling via correlation, dynamic time warping, and rank-based metrics; high-frequency data analysis via cross-correlation, Granger causality, and regression models.
result Strongly coupled stock pairs often exhibit lead-lag effects, especially at finer time scales.

Proposes a method to calibrate data for more accurate linear correlation testing.

problem Inaccurate Pearson's correlation coefficient due to sample size and data non-normality.
method Predictive data calibration using machine learning to condition data on expected linear relationship.
result Calibrated Pearson's correlation coefficient yields a calibrated p-value and r estimate for posterior probability interpretation.

Proposes a partially linear structure to capture nonlinear relationships in mixture of experts models.

problem Suboptimal estimates due to linearity assumption in mixture of experts models.
method Introduces a partially linear structure that incorporates unspecified functions to capture nonlinear relationships.
result Establishes the identifiability of the proposed model under mild conditions and introduces a practical estimation algorithm.

ReGENN improves time series forecasting by considering inter and intra-temporal relationships.

problem Achieving reliable predictions in real-world time series applications.
method ReGENN combines graph evolution with deep recurrent learning to model dynamic dependencies among multiple variables.
result Sound improvement of up to 64.87% over competing algorithms in time-series forecasting.