Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

13.0%26.1%39.1%52.1% · Jun 202019922001200920182026
48 results for data analysis pipelines

Study proposes a statistical testing framework for evaluating clustering pipelines.

problem Quantifying the statistical reliability of clustering results from data analysis pipelines.
method Selective inference-based statistical testing framework for clustering pipelines.
result The proposed test controls the type I error rate and is effective in validating clustering results.

Paper proposes a statistical test for feature selection pipelines using selective inference.

problem Assessing the significance of feature selection pipelines in data analysis.
method Selective inference technique applied to feature selection pipelines composed of various algorithms.
result The proposed statistical test controls false positive feature selection probabilities.

Method predicts computational reproducibility of large population studies data analysis pipelines.

problem Difficulty in evaluating reproducibility of large population studies due to computational and storage requirements.
method Formulated as collaborative filtering process with constraints on training set construction.
result One sampling method, 'Random File Numbers (Uniform)', predicts reproducibility with good accuracy.

A rigorous ML pipeline for binary classification in biomedical studies, focusing on pancreatic cancer.

problem Handling bias in ML models for complex biomedical data.
method Customizable ML analysis pipeline with 9 algorithms, hyperparameter optimization, and thorough evaluation.
result Comparison of ML algorithms to ExSTraCS, highlighting interpretability and bias handling.

TODS automates time series outlier detection with customizable pipelines.

problem Automated detection of outliers in time series data.
method Modular system with 70 primitives for data processing, time series analysis, and detection algorithms. GUI and data-driven searcher for pipeline design.
result Automated discovery and construction of effective outlier detection pipelines.

Method quantifies error contributions in image classification pipelines.

problem Understanding and optimizing complex machine learning pipelines.
method Quantifies error contributions from computational steps, algorithms, and hyperparameters.
result Random search outperforms Bayesian optimization in quantifying error contributions.

Analyzes 6M Python notebooks and 2M enterprise DS pipelines to guide investments in data science.

problem Challenges in following the rapidly evolving landscape of data science technologies and applications.
method Downloaded and analyzed over 6M Python notebooks and 2M enterprise DS pipelines, performing statistical and comparative analyses.
result Identifies actionable conclusions for system builders and technology bets for practitioners based on current trends.

Method quantifies error from pipeline components in image classification.

problem Understanding error sources in complex machine learning pipelines.
method Computes error contribution and propagation from computational steps, algorithms, and hyperparameters.
result Random search method accurately quantifies error contribution and propagation.

Study reveals similarities in knowledge flows between pharmaceutical and AI industries.

problem Understanding the dynamics of drug pipelines in global pharmaceutical industry.
method Multilayer network analysis of drug pipeline, global supply chain, and ownership data.
result Proven similarities in knowledge flows between pharmaceutical and AI industries.

The study of neurocognitive tasks requiring accurate localisation of activity often rely on functional Magnetic Resonance Imaging, a widely adopted technique that makes use of a pipeline of data processing modules, each involving a variety of parameters. These parameters are frequently set according to the local goal o…

2016-10-13abs ↗pdf ↗

Pipeline integrates cross-sectional and longitudinal multi-omics data for IBD research.

problem Integrating diverse data types from the same individuals for disease understanding.
method Statistical and deep learning methods for variable selection, feature extraction, and joint integration.
result Identified microbial pathways, metabolites, and genes discriminating IBD status.

MarmoNet automates analysis of marmoset brain axonal projections.

problem Automatically detect and segment axonal tracer signals in noisy, cluttered images.
method Uses machine learning, specifically CNNs and image registration, to process and map axonal projections.
result Automated pipeline extracts and maps axonal projections robustly.

Functional Magnetic Resonance Imaging (fMRI) relies on multi-step data processing pipelines to accurately determine brain activity; among them, the crucial step of spatial smoothing. These pipelines are commonly suboptimal, given the local optimisation strategy they use, treating each step in isolation. With the advent…

2017-10-02abs ↗pdf ↗

DeepLine automates ML pipeline generation using reinforcement learning.

problem Automatic generation of end-to-end ML pipelines combining multiple algorithms.
method Deep Reinforcement Learning with hierarchical actions filtering.
result DeepLine outperforms state-of-the-art approaches in accuracy and computational cost.

Sabrina integrates financial data and domain knowledge for better visualization.

problem Scattered financial data across various sources makes it hard for analysts to understand the economy.
method Sabrina uses a pipeline to fuse firm-specific and macroeconomic data, visualizing it in a unified interface.
result Sabrina aids financial analysts in their analysis process, as shown in a user study.

PEHRT harmonizes EHR data for translational research.

problem Barriers in using EHR data for translational research.
method Common pipeline including open-source code, visualization tools, and detailed documentation.
result PEHRT harmonizes EHR data to standardized ontologies and generates robust embeddings.

Study finds whitepaper narratives do not predict market factor structure.

problem Predicting market behavior from cryptocurrency whitepaper claims.
method Zero-shot NLP classification combined with CP tensor decomposition of market data.
result Weak alignment between whitepaper claims and market statistics and latent factors.

A new method combines synthetic data analysis and DP generation to produce accurate uncertainty estimates.

problem Invalid inferences from DP synthetic data analysis.
method Combining synthetic data analysis techniques from MI and NA Bayesian modeling with a novel noise-aware synthetic data generation algorithm.
result Accurate confidence intervals from DP synthetic data are produced, wider with tighter privacy.

This work classifies monodromy in vineyards using singularity theory.

problem Understanding and predicting monodromy in vineyards for topological data analysis.
method Using a connection with singularity theory, the study classifies monodromy in vineyards of 1-manifolds in R^2.
result Monodromy in vineyards occurs only if they contain a specific singularity of the distance function.

Automates supervised learning pipeline design with matrix and tensor factorization.

problem Designing effective supervised learning pipelines with many choices.
method Uses matrix and tensor factorization to model pipeline search space and develops greedy experiment design protocols.
result Demonstrates the effectiveness of the approach on real-world classification problems.

Study recommends PLM choices for minimizing calibration error in NLP tasks.

problem Minimizing calibration error in PLM-based NLP predictions.
method Compared various options for PLM encoding, size, uncertainty quantifier, and fine-tuning loss.
result Recommendations for a well-calibrated PLM-based prediction pipeline.

Automatically computes reference ranges for UK Biobank cardiac data.

problem Improving healthcare by discovering patterns in large-scale population data.
method Fully automatic pipeline for 3D cardiac MR image analysis.
result Statistically significant agreement between manual and automatic indexes.

Stable density-based clustering via multiparameter persistence.

problem Density-based clustering stability to data perturbations.
method Degree-Rips construction, correspondence-interleaving distance, multiparameter stability analysis.
result Persistable pipeline yields stable, consistent density-based clustering.

Cilia are hairlike structures protruding from nearly every cell in the body. Diseases known as ciliopathies, where cilia function is disrupted, can result in a wide spectrum of disorders. However, most techniques for assessing ciliary motion rely on manual identification and tracking of cilia; this process is laborious…

2018-03-20abs ↗pdf ↗

Adaptive Bayesian model optimizes machine learning pipelines.

problem Automated model selection and hyperparameter tuning for datasets.
method Combines adaptive Bayesian regression and neural network basis function with acquisition function.
result Identifies high-performance pipelines efficiently and outperforms baseline methods.

The paper uses deep learning to speed up spatial and visual connectivity analysis.

problem Slow calculation of spatial and visual connectivity metrics.
method Investigates machine learning models and a pipeline for training them on spatial and visual connectivity analysis.
result Deep learning models significantly speed up the analysis process.

Optimizes deep learning pipelines with novel algorithms for smooth and non-smooth functions.

problem Optimizing deep learning pipelines for smooth and non-smooth functions.
method Provided matching lower and upper bounds for smooth convex and non-convex functions, and developed PPRS for non-smooth convex functions.
result PPRS achieves near-linear speed-up and convergence time for non-smooth non-convex problems.

A hybrid RL-Bayesian search configures machine learning pipelines efficiently.

problem Optimizing hyper-parameters in a hierarchical conditional space.
method Combines Reinforcement Learning and Bayesian Optimization.
result Outperforms state-of-the-art methods in pipeline optimization.