Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

6.3%12.5%18.8%25.0% · Apr 199319922001200920182026
48 results for discovery significance

The paper integrates statistical significance and discriminative power in pattern discovery.

problem Discovering actionable patterns that meet rigorous statistical significance and discriminative power criteria.
method Integrates statistical significance and discriminative power criteria into state-of-the-art algorithms.
result Improves discriminative power and statistical significance of discovered patterns without quality deterioration.

FSR efficiently discovers significant patterns with few resampled datasets.

problem Mining significant patterns in transactional data, especially subgroups.
method FSR uses resampling to bound the supremum deviation of quality statistics, providing rigorous guarantees on false discoveries.
result FSR effectively discovers significant subgroups with a small number of resampled datasets.

Q-SAVI model improves drug discovery accuracy with prior knowledge of chemical space.

problem Challenges in drug discovery due to covariate shift and limited labeled data.
method Probabilistic model with domain-informed prior distributions over functions.
result Q-SAVI outperforms state-of-the-art techniques in predictive accuracy and calibration.

Estimates target GGM using auxiliary studies with false discovery rate control.

problem Estimating high-dimensional GGMs from related studies.
method Transfer learning with Trans-CLIME and debiased Trans-CLIME estimators.
result Debiased Trans-CLIME estimator provides element-wise asymptotic normality and false discovery rate control.

Amortized Causal Discovery learns to infer causal graphs from time-series data, improving performance.

problem Inference of causal graphs from time-series data is inefficient due to fitting new models for each sample.
method Proposes Amortized Causal Discovery, a variational model that leverages shared dynamics across samples with different causal graphs.
result Significant improvements in causal discovery performance demonstrated experimentally.

Paper discovers governing equations from data using differential invariants.

problem Discovering partial differential equations from data is challenging.
method The paper proposes a pipeline based on differential invariants to reduce the search space and adhere to symmetry.
result DI-SINDy method outperforms other symmetry-informed methods in PDE discovery.

New method combines gradient optimization with constraint-based techniques for causal discovery.

problem Causal discovery from observational data, especially with small sample sizes.
method Differentiable dd-separation scores using percolation theory and soft logic for gradient-based optimization of conditional independence constraints.
result Empirical evaluations show robust performance in low-sample regimes, surpassing traditional methods.

Neural causal discovery methods fail to accurately uncover causal structures due to the faithfulness property.

problem Accuracy in neural causal discovery is limited, especially when distinguishing between existing and non-existing causal relationships.
method Systematic evaluation of neural causal discovery methods, focusing on their performance in finite sample regimes and their ability to recover ground-truth graphs.
result Neural networks lack the precision to reliably recover ground-truth causal graphs, even for small graphs and large sample sizes.

This work tackles causal graph discovery with stochastic interventions to minimize the number of interventions.

problem Discovering the true causal graph from observational data with limited interventions.
method Proposes a stochastic intervention model and studies verification and search problems with approximation algorithms.
result Provides approximation algorithms with competitive ratios for verification and search problems.

Simple bounds show most cross-sectional predictability findings are likely true.

problem Determining the validity of cross-sectional return predictability findings.
method Developed simple and intuitive bounds on the false discovery rate (FDR).
result Bounds show the FDR is small, indicating most findings are likely true.

CardiGraphormer uses SSL and GNNs to improve drug discovery.

problem Challenges in drug discovery due to combinatorial chemical space and limited approved drugs.
method Combines self-supervised learning, Graph Neural Networks, and Cardinality Preserving Attention.
result Enhanced predictive performance and interpretability in drug discovery.

Bayesian methods improve drug discovery experiment design.

problem Optimizing drug screening experiments in high-dimensional data.
method Bayesian inference and optimisation with upper confidence bound algorithms, Thompson sampling, and sparse tree search.
result Sparse tree search techniques outperform other methods in drug toxicity screening.

MEC-IP uses IP to efficiently find MECs in BNs from observational data.

problem Discovering Markov Equivalent Classes (MECs) in Bayesian Networks (BNs) efficiently.
method Clique-focusing strategy and EMSG for MEC discovery via Integer Programming.
result Significant reduction in computational time and improved accuracy.

Paper tackles MIAs vulnerability by controlling FDR, providing guarantees on false discoveries.

problem Vulnerability of deep learning models to membership inference attacks (MIAs).
method Designs a novel membership inference attack method that provides FDR guarantees.
result Demonstrates the effectiveness of the method in various settings.

TimeGraph creates synthetic datasets for robust time-series causal discovery.

problem Lack of reliable synthetic benchmark datasets for robust time-series causal discovery.
method Developed comprehensive synthetic datasets with temporal properties, including trends, seasonality, and noise.
result Demonstrated significant variations in algorithm performance under realistic temporal conditions.

EDL discovers state-covering skills without relying on task rewards.

problem Discovering skills in reinforcement learning without a task-oriented reward function.
method EDL optimizes information-theoretic objective using different machinery to address coverage problem.
result EDL discovers state-covering skills more effectively than existing methods.

SurvNet selects important variables in DNNs with false discovery rate control.

problem Variable selection in deep neural networks (DNNs) for interpretability.
method Backward elimination procedure based on a new variable importance measure.
result SurvNet estimates and controls false discovery rate of selected variables.

This paper compares feature selection methods for biomarker discovery in toxicant-treated fish.

problem Choosing the most suitable method for biomarker discovery in toxicant exposure studies.
method Three feature selection methods: SAM, mRMR, and GeoDE are compared.
result Different methods perform better in different cases, requiring dataset-specific decisions.

SPOT improves differentiable causal discovery by estimating skeleton posterior for latent confounders.

problem Scalable and accurate estimation of causal skeletons in the presence of latent confounders.
method SPOT (Skeleton Posterior-guided OpTimization) framework that estimates skeleton posterior and integrates it with differentiable causal discovery.
result SPOT enhances differentiable causal discovery by reducing the search space and improving accuracy.

The paper tackles online FDR control in hypothesis testing with contextual features.

problem Controlling false discoveries in online hypothesis testing with contextual features.
method Proposes a new class of online testing procedures that learn significance levels sequentially, incorporating contextual information and previous results.
result Proves that the proposed procedures control online FDR under standard assumptions and outperform existing methods in terms of statistical power.

Customer momentum is a positive relationship between a firm's returns and past returns of its customers.

problem Understanding the relationship between a firm's returns and its customers' past returns.
method Examined customer momentum using a long-short equally-weighted decile portfolio and Fama-French factor models.
result Customer momentum generates significant monthly returns and is statistically significant.

CatNet controls FDR in LSTM models using SHAP feature importance and Gaussian mirrors.

problem Controlling False Discovery Rate (FDR) in LSTM models with feature selection.
method CatNet uses SHAP values for feature importance and Gaussian Mirror algorithm for FDR control. It introduces a kernel-based independence measure to handle feature correlations.
result CatNet reduces overfitting and improves model interpretability on simulated and real-world data.

New algorithm discovers causal relationships from observational data efficiently.

problem Inferring direct causal parents from a large set of variables.
method Orthogonal structure search approach, scaling to large graphs, guarantees for nonlinear relationships.
result Significant improvements over existing methods in causal discovery from observational data.

Diamond method controls FDR for trustworthy feature interaction discovery in ML models.

problem Limited interpretability of ML models due to black box nature.
method Diamond method integrates model-X knockoffs framework to control FDR for non-additive interactions.
result Diamond method ensures accurate discovery of feature interactions with FDR control.

Unified algorithm for incorporating various prior knowledge in multiple testing.

problem Improving power and precision in multiple testing procedures with prior knowledge.
method p-filter algorithm that incorporates four types of prior knowledge: null hypotheses, penalties, groups, and independence.
result Unified framework allows for recovery of various known algorithms.

BLADE uses Bayesian methods to discover complex systems from scarce data.

problem Efficiently discovering governing equations of complex dynamical systems from limited data.
method Combines replica-exchange stochastic gradient Langevin Monte Carlo with active learning.
result Reduces measurement requirements by 60% for Lotka-Volterra and 40% for Burgers' equation.

MetaCaDI learns causal graphs and unknown interventions from few data instances.

problem Discovering causal mechanisms in systems with high data costs and unknown interventions.
method MetaCaDI is a Bayesian meta-learning framework that optimizes for rapid adaptation to new intervention targets.
result MetaCaDI significantly outperforms state-of-the-art methods in causal graph recovery and intervention target prediction.

New method for identifying causal relationships in financial time series data.

problem Identifying causal relationships in nonstationary financial time series data.
method Refined constraint-based causal discovery algorithm (CD-NOTS) for nonstationary time series data.
result CD-NOTS effectively identifies causal connections in financial applications.

Improved jet tagging reduces systematic uncertainties and enhances signal purity.

problem Boosted resonance decay signals from jets are difficult to distinguish from background.
method Adversarial neural networks to decorrelate jet substructure tagger.
result Adversarial trained tagger outperforms conventional methods in discovery significance.