Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

96192288384 · Jun 202019922001200920172026
48 results for observational sciences

This paper improves prediction accuracy for multi-input classification tasks using p-value aggregation.

problem Generating accurate predictive sets with guaranteed coverage for multi-input classification tasks.
method Integrates p-values from each observation to reduce the size of the predicted label set while maintaining class-conditional coverage.
result The method reduces the size of the predicted label set while preserving the required coverage guarantee.

SLdisco uses supervised learning to discover causal models from observational data.

problem Estimating causal effects from observational data with limited samples and sparse models.
method Supervised machine learning to map observational data to causal equivalence classes.
result SLdisco is more conservative, less sensitive to sample size, and provides better model inference.

Develops scalable methods to assess sensitivity and uncertainty in continuous treatment effects.

problem Estimating effects of continuous-valued interventions from observational data, especially when ignorability and positivity assumptions are violated.
method Continuous treatment-effect marginal sensitivity model (CMSM), scalable algorithm, uncertainty-aware deep models.
result Derives bounds that agree with observed data and a defined level of hidden confounding.

The paper uses a graph autoencoder to learn unbiased plant-pollinator interaction embeddings.

problem Sampling bias in citizen science data affects ecological network analysis.
method Bipartite graph variational autoencoder with HSIC for fairness.
result The method mitigates sampling bias and provides unbiased embeddings.

Matrix completion is a classical problem in data science wherein one attempts to reconstruct a low-rank matrix while only observing some subset of the entries. Previous authors have phrased this problem as a nuclear norm minimization problem. Almost all previous work assumes no explicit structure of the matrix and uses…

2019-04-17abs ↗pdf ↗

We present a novel method for obtaining high-quality, domain-targeted multiple choice questions from crowd workers. Generating these questions can be difficult without trading away originality, relevance or diversity in the answer options. Our method addresses these problems by leveraging a large corpus of domain-speci…

2017-07-19abs ↗pdf ↗

Causal relationships in time series with latent variables are discovered using LPCMCI.

problem Discovering causal relationships in complex, time-series data with hidden variables.
method Evaluated LPCMCI algorithm for finding generators compatible with multi-dimensional, autocorrelated time series with latent variables.
result LPCMCI performs better than random guessing but is not optimal.

New model estimates species population trends from citizen science data.

problem Interannual confounding in citizen science data.
method Double Machine Learning framework to estimate population change and propensity scores for confounding adjustment.
result Spatially detailed trend estimates from citizen science data with low error rates.

Automated detection of new, interesting, unusual, or anomalous images within large data sets has great value for applications from surveillance (e.g., airport security) to science (observations that don't fit a given theory can lead to new discoveries). Many image data analysis systems are turning to convolutional neur…

2018-06-21abs ↗pdf ↗

In this paper, we show how simple logistic growth that was studied intensively during the last 200 years in many domains of science could be extended in a rather simple way and with these extensions is capable to produce a collection of behaviors widely observed in an enormous number of real-life systems in Economics, …

2008-02-24abs ↗pdf ↗

LUQ learns QoI from dynamical systems for consistent observation inversion.

problem Quantifying uncertainties on model inputs corresponding to observable QoI in dynamical systems.
method LUQ framework for SIPs, including data filtering, dynamics learning, observation classification, and feature extraction.
result LUQ provides tractable solutions to SIPs for dynamical systems, enabling uncertainty quantification.

Deep Reinforcement Learning (DRL) has emerged as a powerful control technique in robotic science. In contrast to control theory, DRL is more robust in the thorough exploration of the environment. This capability of DRL generates more human-like behaviour and intelligence when applied to the robots. To explore this capa…

2019-10-16abs ↗pdf ↗

S-DIDML integrates structural DID with ML for causal inference in high-dimensional data.

problem Causal inference in high-dimensional observational panel data with confounding variables.
method Structural identification with high-dimensional estimation, Neyman orthogonality, cross-fitting, causal forests, semi-parametric models.
result Precision in identifying policy-sensitive groups and optimizing resource allocation.

Blockchain technology, and more specifically Bitcoin (one of its foremost applications), have been receiving increasing attention in the scientific community. The first publications with Bitcoin as a topic, can be traced back to 2012. In spite of this short time span, the production magnitude (1162 papers) makes it nec…

2019-06-21abs ↗pdf ↗

StepMix estimates mixture models with covariates for social science applications.

problem Estimating latent classes with covariates in social science models.
method Pseudo-likelihood estimation using one-, two-, and three-step approaches.
result Unified framework for expectation-maximization subroutines.

Symmetric observations don't necessarily imply symmetric causal explanations.

problem Inferring causal models from observed correlations is challenging and computationally intensive.
method An explicit example using a tripartite probability distribution over binary events.
result Symmetries in observations cannot be used to reduce the hypothesis space of causal models.

metabeta uses neural networks to speed up Bayesian mixed-effects regression.

problem Bayesian mixed-effects regression is computationally expensive.
method metabeta is a neural network model that pre-trains to estimate posterior distributions.
result metabeta achieves comparable performance to MCMC at a fraction of the time.

Machine learning methods have been remarkably successful for a wide range of application areas in the extraction of essential information from data. An exciting and relatively recent development is the uptake of machine learning in the natural sciences, where the major goal is to obtain novel scientific insights and di…

2019-05-21abs ↗pdf ↗

Estimates price sensitivity from transaction data using a novel odds ratio method.

problem Estimate price sensitivity from transaction-level data with partially observed treatment assignments.
method Recursive partitioning procedure with adversarial imputation for robust estimation.
result Validated on synthetic data and applied to three case studies, demonstrating heterogeneity in treatment effects.

The paper proposes using network science to improve portfolio optimization by reducing noise in covariance estimation.

problem Noise in covariance estimation leads to suboptimal portfolio performance.
method The paper introduces SR-IFN, a network-based method to filter out noise from empirical covariance, enhancing portfolio optimization.
result The SR-IFN network improves portfolio performance by selecting peripheral, diversified assets and inversely weighting them based on centrality.

Foundation models alter medical data science workflow, challenging veridical data science principles.

problem Foundation models disrupt traditional data science practices in medicine.
method Critically examined the medical foundation model lifecycle and its deviation from veridical data science principles.
result Foundation models challenge veridical data science principles of predictability, computability, and stability.

A society or country with income equally distributed among its people is truly a fiction! The phenomena of socioeconomic inequalities have been plaguing mankind from times immemorial. We are interested in gaining an insight about the co-evolution of the countries in the inequality space, from a data science perspective…

2017-12-31abs ↗pdf ↗

Gaussian processes model sparse data in astrophysics and chemistry.

problem Scarcity of data in high-energy astrophysics and synthetic chemistry.
method Gaussian processes for uncertainty-aware predictions and inferences.
result GPs enable predictions and model latent emission from black holes and molecules.

Novel method for Bayesian model comparison using deep learning.

problem Comparing complex models in science with intractable likelihood functions.
method Simulation-based, purely deep learning approach that amortizes model fitting costs.
result Achieves excellent results in accuracy, calibration, and efficiency.

We find polynomial-time solutions to the word problem for free-by-cyclic groups, the word problem for automorphism groups of free groups, and the membership problem for the handlebody subgroup of the mapping class group. All of these results follow from observing that automorphisms of the free group strongly resemble s…

2006-08-23abs ↗pdf ↗

Defines data science as a natural ecosystem with challenges and missions.

problem Challenges and missions in data science due to 5D complexities and data life cycle phases.
method Systemic and data-centric view of data science as a fusion of data universe and its challenges, formalizing a general-purpose architecture.
result Essential data science as a natural ecosystem integrating specific disciplines and high-impact applications.