Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

4487131174 · Jun 202019922001200920182026
48 results for online discovery

Private online FDR control for adaptive testing under differential privacy.

problem Controlling false discoveries in adaptive multiple hypothesis testing with privacy constraints.
method Private online algorithms based on non-private results, ensuring privacy and statistical performance.
result Strong guarantees for privacy and statistical performance in FDR and power.

New method controls false discoveries in online testing with deadlines.

problem Controlling false discoveries in online hypothesis testing with decision deadlines.
method Benjamini-Hochberg-type procedure over a moving window of hypotheses with adaptive threshold parameters.
result Controls false discovery rate at every stage and adaptively chosen stopping times.

New rules control false discoveries in online anomaly detection for time series data.

problem Controlling false discoveries in anomaly detection for time series data.
method Novel online false discovery rate control (FDRC) rules for time series anomaly detection.
result Ensures high power in detecting anomalies even when the alternative is rare and test statistics are serially dependent.

The paper develops online methods to control familywise error rate in growing hypothesis testing sequences.

problem Controlling familywise error rate in a growing sequence of hypotheses over time.
method Unified algorithmic concepts for offline and online FWER control, including new adaptive online algorithms.
result Substantial gains in power demonstrated and formally proved in a Gaussian sequence model.

New findings control FDR for online testing methods under positive dependence.

problem Maintaining FDR control for online testing methods under positive dependence.
method Developed new methods to control FDR for online testing procedures under positive dependence.
result SAFFRON and LORD control FDR under positive dependence, not just conditional superuniformity.

The paper tackles online FDR control in hypothesis testing with contextual features.

problem Controlling false discoveries in online hypothesis testing with contextual features.
method Proposes a new class of online testing procedures that learn significance levels sequentially, incorporating contextual information and previous results.
result Proves that the proposed procedures control online FDR under standard assumptions and outperform existing methods in terms of statistical power.

New method controls false discoveries in real-time data streams.

problem Online testing of hypotheses with strict error constraints and no future data.
method Structure-adaptive sequential testing (SAST) with alpha-investment algorithm.
result Substantial power gain over existing online testing rules.

Multiple hypothesis testing is a core problem in statistical inference and arises in almost every scientific field. Given a set of null hypotheses H(n)=(H1,,Hn)\mathcal{H}(n) = (H_1,\dotsc, H_n), Benjamini and Hochberg introduced the false discovery rate (FDR), which is the expected proportion of false positives among rejected nu…

2016-03-29abs ↗pdf ↗

AIHT improves online high-dimensional quantile regression by separating support discovery and refinement.

problem Online high-dimensional quantile regression with structural sparsity.
method Adaptive Iterative Hard Thresholding (AIHT) alternates stochastic updates with adaptive hard-thresholding steps.
result AIHT achieves logarithmic regret for the sliding-window objective in high-dimensional settings.

Online method selects candidates from data streams, ensuring irreversible decisions.

problem Conformal selection's incompatibility with irreversible decisions in online scenarios.
method Online Conformal Selection with Accept-to-Reject Changes (OCS-ARC) incorporating online Benjamini-Hochberg procedure.
result OCS-ARC controls FDR at or below nominal level, improving selection power.

New algorithms control FDX while achieving more power in online multiple testing.

problem Problems with previous online multiple testing methods, including high FDX and low power.
method Developed new dynamic algorithms that adjust testing levels based on accumulated wealth.
result SupLORD algorithm achieves higher power and FDR control in synthetic experiments.

Robust Bayesian changepoint detection with ββ-divergences reduces false discovery rates.

problem Detecting changepoints in non-stationary streaming data with high accuracy.
method Doubly robust Bayesian Online Changepoint Detection (BOCD) using ββ-divergences.
result False discovery rates of changepoints reduced from over 90% to 0%.

C-PP-COAD detects anomalies with limited real data, reducing dependency on real calibration data.

problem Limited real calibration data for online anomaly detection.
method Context-aware prediction-powered conformal online anomaly detection (C-PP-COAD).
result Significantly reduces dependency on real calibration data without compromising FDR control.

The aim of process discovery, originating from the area of process mining, is to discover a process model based on business process execution data. A majority of process discovery techniques relies on an event log as an input. An event log is a static source of historical data capturing the execution of a business proc…

2017-04-25abs ↗pdf ↗

MO2 learns useful behaviours from past experience for new tasks.

problem Discovering useful behaviours from past experience and transferring them to new tasks.
method Model-Based Offline Options (MO2) framework supporting sample-efficient bottleneck option discovery over continuous state-action spaces.
result MO2 outperforms recent option learning methods on complex long-horizon continuous control tasks.

Forest Fire Clustering discovers cell types from single-cell data.

problem Discovering cell types from large-scale single-cell sequencing data.
method Iterative label propagation and parallelized Monte Carlo simulation.
result Forest Fire Clustering outperforms state-of-the-art methods on diverse benchmarks.

A new method for multiple testing reduces false discoveries while maximizing power.

problem Maximizing statistical power while controlling false discoveries in multiple testing scenarios.
method Adaptive sampling approach inspired by multi-armed bandits to minimize sample size.
result The method achieves sample complexity close to information theoretic lower bounds and outperforms uniform sampling.

A new metric, Weighted Regret, unifies FDR and power evaluation in online multiple testing.

problem The asymmetric costs of false positives and false negatives in automated pipelines.
method Introducing Weighted Regret and Decoupled-OMT (DOMT) to unify FDR and power evaluation.
result DOMT achieves an order-optimal sublinear mitigation of threshold depletion in bursty environments.

Model learns to discover and disambiguate entities and relations in text streams.

problem Learning to follow and resolve mentions in a continuous text stream.
method End-to-end trainable memory network for online, one-shot learning.
result Improves disambiguation and discovery skills with minimal supervision.

New method learns decisions from collective preferences without individual covariates.

problem Making decisions online without individual covariates.
method Collaborative filtering, matrix completion bandit, ε-greedy policy, online gradient descent, inverse propensity weighting.
result Method outperforms benchmarks and reveals new discoveries.

DiffATD efficiently discovers targets in partially observable environments using diffusion dynamics.

problem Efficiently discovering targets in partially observable environments with limited sampling.
method DiffATD uses diffusion dynamics to maintain a belief distribution over unobserved states, balancing exploration and exploitation.
result DiffATD outperforms baselines and supervised methods in diverse domains.

New methods identify concepts in trained embeddings reliably without human labels.

problem Identifying interpretable concepts in trained embedding spaces without human labels.
method Explicitly connecting concept discovery to PCA and ICA, proposing novel approaches for dependent concepts.
result Proven methods outperform competitors on a variety of experiments, achieving up to 29% better alignment with ground truth.

Paper proposes NAC for efficient network discovery in incomplete networks.

problem Efficiently discover vertices with specific attributes in incomplete networks.
method Formulates network discovery as a reinforcement learning problem, uses deep reinforcement learning with task-specific network embeddings.
result Offline planning leads to significantly improved performance compared to online discovery algorithms.

GOCPD detects change points by maximizing the probability of two independent models.

problem Large false discovery rates in online change point detection methods.
method GOCPD uses ternary search to find change points by maximizing the probability of two independent models.
result GOCPD accelerates CPD with logarithmic complexity for single change point detection.

Paper proposes using word embeddings to detect trolls in social media debates.

problem Preventing online harassment through rapid detection of offensive posts.
method Word embedding models for identifying fast-changing topics and negative content.
result GloVe model helps in discovering new keywords for trolling detection.

Automates organizing diverse web data into a hierarchical topic model.

problem Manual classification of all scientific and popular scientific knowledge is impractical.
method Proposes an algorithm to aggregate multiple collections into a single hierarchical topic model.
result Demonstrates a web service for topical exploratory search.

Causal relationships in time series with latent variables are discovered using LPCMCI.

problem Discovering causal relationships in complex, time-series data with hidden variables.
method Evaluated LPCMCI algorithm for finding generators compatible with multi-dimensional, autocorrelated time series with latent variables.
result LPCMCI performs better than random guessing but is not optimal.

This paper discovers classification models from sequential data without prior knowledge.

problem Lack of prior knowledge in defining kernels for online classification.
method Adapts GP-based time-series structure discovery with SMC to learn new features from sequential data.
result Improves classification accuracy by 10% on real-world data.

OLCS-Ranker improves peptide identification accuracy and speed on hard datasets.

problem Efficiently identifying peptides from MS/MS data, especially on hard datasets with many false positives.
method Cost-sensitive online learning model and iterative online learning algorithm.
result OLCS-Ranker outperforms existing methods in accuracy and speed on large datasets.

Paper proposes a recursive GPSSM for efficient online learning.

problem Efficient online learning for dynamical models with limited prior information.
method Recursive Gaussian Process State-Space Model with adaptive capabilities for domains and hyperparameters.
result Superior accuracy, computational efficiency, and adaptability compared to state-of-the-art methods.

ADDIS improves power in online FDR control for conservative nulls.

problem Lack of power in adaptive FDR control algorithms for conservative nulls.
method ADDIS: adaptive discarding algorithm for online FDR control.
result ADDIS achieves best of both worlds: high power for conservative nulls and no loss for uniformly distributed nulls.

New algorithm detects changes in high-dimensional data with mean and variance.

problem Challenges in detecting changes in high-dimensional data with mean and variance.
method Complete graph-based approach to detect changes of mean and variance from low to high-dimensional online data.
result The proposed method outperforms existing methods in terms of detection power.

Bayesian reflex models AI learning like the autonomic nervous system.

problem Online learning in dynamic AI environments.
method Bayesian online algorithms with belief maintenance, sequential updating, and uncertainty-driven action balancing.
result Unified framework for adaptive AI learning.

Novel semi-supervised method for online structure learning in noisy data streams.

problem Discovering complex relations in noisy data streams with limited labelled data.
method Combines graph-cut minimization and first-order logic for online, single-pass label completion.
result Improves accuracy of structure learning system by completing missing labels.