Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

2905808701,160 · Jun 202019922001200920182026
48 results for individual data

Deep learning models can infer individual trajectories from sparse data.

problem Learning individual dynamics from limited data points.
method Combining variational autoencoders (VAEs) with ordinary differential equations (ODEs) for dynamic modeling.
result Deep learning can recover individual trajectories from sparse data, but requires careful adaptation.

Paper proposes a method to optimize policies for diverse individuals using heterogeneous data.

problem Learning optimal policies for a heterogeneous population from pre-collected data.
method Individualized offline policy optimization framework for heterogeneous MDPs.
result The proposed P4L algorithm achieves a fast rate of average regret.

Paper introduces Functional Effects Models to account for individual heterogeneity in panel data.

problem Accounting for preference heterogeneity in panel data with machine learning.
method Functional Effects Models using gradient boosting decision trees and deep neural networks to learn individual-specific preference parameters.
result Functional Effects Models outperform traditional models in learning inter-individual heterogeneity and predictive performance.

Method estimates group structure in panel data using variance information.

problem Estimating group structure in panel data with unknown groups.
method Proposes a method to estimate unobserved groupings for panel data models using variance information.
result Superior performance compared to existing methods in simulations and empirical applications.

Paper presents a framework to infer individual data from aggregate data.

problem Inference of individual-level data from aggregate data due to privacy concerns.
method End-to-end pipeline for processing aggregate data, novel algorithm for reconstruction, machine learning models.
result Valid and usable answers derived from machine learning models using multiple candidate datasets.

Proposes a model to handle mobile health data with irregular measurements.

problem Handling heterogeneous, multi-resolution data in mobile health.
method Individualized dynamic latent factor model for irregular multi-resolution time series data.
result Superior performance compared to existing methods in simulation and smartwatch data applications.

Paper operationalizes individual fairness using side-information and a unified representation.

problem Difficulty in eliciting a human specification of a similarity metric for individual fairness.
method Proposes a Pairwise Fair Representation (PFR) model that learns from fairness graph and side-information.
result Unified PFR model effectively operationalizes individual fairness without human specification.

ContiVAE estimates individual dose-response curves from unobserved confounders using observational data.

problem Estimating causal effects of continuous treatments considering unobserved confounders.
method Variational auto-encoder with a Tilted Gaussian prior distribution modeling hidden confounders as latent variables.
result ContiVAE outperforms existing methods by up to 62% in predicting individual dose-response curves.

FAST-DAD distills complex ensemble models into faster, more accurate individual models.

problem Deploying complex AutoML ensemble predictors on tabular data is slow, large, and opaque.
method Data augmentation strategy based on Gibbs sampling from a self-attention pseudolikelihood estimator.
result FAST-DAD distillation produces significantly better individual models than standard training.

Social media reduces individual investors' disposition effect through negative information.

problem The disposition effect in individual investors selling profitable assets too early and holding onto losing assets for too long.
method Analysis of post data and trading data from Xueqiu.com.
result Social media information significantly reduces the disposition effect.

Paper tackles estimating individual treatment effects from observational data.

problem Estimating the difference between outcomes with and without treatment from single observation.
method Formulated as inference from hidden variables, uses a model of four causal populations, proposes ECM algorithm.
result ECM algorithm provides better performance compared to baseline methods on synthetic and real-world data.

New model clusters cells and individuals, revealing genetic influences on cell types.

problem Clustering nested data with group-level and observation-level variables.
method Nested Atoms Model (NAM), Bayesian nonparametric approach.
result Identifies clusters of genetically similar individuals with homogeneous cell-type profiles.

Develops verifiers to check if machine learning models treat similar individuals equally.

problem Ensuring fairness in machine learning models by checking if similar individuals are treated differently.
method Constructs verifiers for proving individual fairness of machine learning models, considering relaxations of the problem.
result Developed verifiers for linear and kernelized polynomial/radial basis function classifiers.

Proposes a new model for estimating individual treatment effects.

problem Estimating individual treatment effects from observational data is challenging.
method Integrates diffusion modeling and conformal inference with propensity score and covariate approximation.
result Establishes rigorous theoretical guarantees and demonstrates competitive performance.

Modeling individual cardiovascular responses from wearable sensor data.

problem Capturing and understanding cardiovascular responses to physical activity and sleep changes.
method Attentional convolutional neural network to learn signatures from minute-level sensor data.
result Generated signatures generalize and outperform baseline models in predicting cardiovascular variables.

New method aligns brain data across individuals for better brain decoding.

problem Inter-individual variability in brain response patterns limits decoder generalization.
method SpectralOT method that embeds cortical geometry into Laplace-Beltrami eigenmodes.
result SpectralOT strikes balance between aligning functional features and preserving anatomical structure.

Framework infers coordination strategies from movement data.

problem Inferring individual movement strategies from group data.
method Formalizes Coordination Strategy Inference Problem; provides methodology to infer strategies.
result Framework accurately infers strategies in simulated and real-world datasets.

Study integrates diverse data sources to predict mental health conditions.

problem Predict individuals' mental health conditions using a heterogeneous network approach.
method Leverage a heterogeneous information network (HIN) to model social interaction, health data, and survey data. Apply recommender system (RS) and node classification (NC) paradigms to predict mental health states.
result RS and NC methods outperform traditional logistic regression models in predicting mental health conditions.

TCFimt forecasts causal effects of multiple interventions from individual data.

problem Estimating causal effects of temporal multi-interventions from individual data.
method TCFimt uses adversarial tasks in seq2seq framework to alleviate bias and contrastive learning to decouple effects.
result TCFimt outperforms state-of-the-art methods in predicting future outcomes and choosing optimal treatments.

The article presents methods to select models from behavioral learning data, with applications to contextual bandits.

problem Model selection for behavioral learning data, especially in non-stationary environments.
method Two model selection methods: a general hold-out procedure and an AIC-type criterion, adapted for non-stationary dependent data.
result Theoretical error bounds for these methods are close to those of the standard i.i.d. case.

New model recommends stocks considering individual preferences and diversification.

problem Inaccurate stock price predictions and ignoring investment theories.
method Portfolio Temporal Graph Network Recommender (PfoTGNRec) incorporating diversification-enhancing sampling.
result PfoTGNRec outperforms state-of-the-art models in real-world data.

Unified theory for semiparametric data fusion with individual-level data.

problem Handling data fusion problems, especially in settings with diverse data sources and designs.
method Extending a comprehensive theory to handle conditional and marginal distribution alignments, providing universal results for influence functions and efficient influence functions.
result Paves the way for machine-learning debiased, semiparametric efficient estimation.

Two simple methods learn fair metrics from data to improve fairness in ML tasks.

problem Lack of widely accepted fair metrics for many ML tasks hinders individual fairness adoption.
method Presented two simple ways to learn fair metrics from various data types.
result Fair training with learned metrics improves fairness on three ML tasks.

Proposes a deep learning method for modeling dynamic individual-level latent trajectories with changing parameters.

problem Modeling longitudinal data with changing individual-level dynamics parameters.
method Combines deep learning for dimensionality reduction and differential equations for dynamic modeling, allowing different parameters for sub-periods.
result Successfully identifies dynamic parameters and predictors of resilience.

Heteroskedasticity biases uplift model rankings, leading to inefficient treatment allocation.

problem Bias in uplift model rankings due to heteroskedasticity.
method Theoretical analysis and simulation on real-world data.
result Heteroskedasticity can cause individuals with high treatment effects to be ranked at the bottom, leading to inefficient treatment allocation.

We consider the problem of fitting a linear model to data held by individuals who are concerned about their privacy. Incentivizing most players to truthfully report their data to the analyst constrains our design to mechanisms that provide a privacy guarantee to the participants; we use differential privacy to model in…

2015-06-10abs ↗pdf ↗

A method for identifying joint and individual subspaces from multi-view data.

problem Unclear conditions for reliably identifying joint and individual subspaces from noisy, high-dimensional measurements.
method Rigorously quantifies conditions based on signal rank, principal angles, and noise levels. Characterizes spectrum perturbations of product of projection matrices.
result Estimates joint and individual subspaces more accurately than existing approaches in simulations and real-world applications.

Estimates individual treatment effects using gradient interpolation and kernel smoothing.

problem Estimating individualized continuous treatment effects in observational data.
method Augment training data with independently sampled treatments and inferred counterfactual outcomes using gradient interpolation and kernel smoothing.
result Our method outperforms state-of-the-art methods on counterfactual estimation error.

Paper introduces a method to learn physics between digital twins using imperfect models.

problem Learning physics from imperfect data and low-fidelity models.
method Bayesian Hierarchical modeling with physics-informed Gaussian processes.
result Models learning between digital twins are less uncertain than independent models but not over-confident.

Technical report predicts eating and food purchasing behaviors of free-living individuals.

problem Predicting eating and food purchasing behaviors of free-living individuals.
method Applied multiple machine learning algorithms (Logistic Regression, RBF-SVM, Random Forest, Gradient Boosting) to minute-level features from sensors and environmental context.
result Gradient Boosting model had the highest mean accuracy score (0.7289) for predicting eating events before 0 to 4 minutes.