Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

23456890 · Jun 202019922001200920182026
48 results for automated discovery

Bayesian symbolic regression automates model discovery from data.

problem Learning closed-form mathematical models from data using heuristic methods.
method Probabilistic approach to symbolic regression, connecting to information theory and statistical physics.
result Probabilistic approach provides model plausibility and performance guarantees.

TSML tackles anomaly detection and pattern discovery in industrial time series data.

problem Extracting and exploiting information from large industrial data to reduce downtimes and manufacturing errors.
method TSML uses a pipeline of lightweight filters to process industrial time series data in parallel.
result TSML effectively detects anomalies and discovers patterns in industrial time series data.

Software automates metabolomics data analysis for reproducible results.

problem Automating reproducible metabolomics data analysis.
method Object-oriented software engineering, Java, XML database, GUI, version control system.
result MeKDDaM-SAGA successfully guides metabolomics applications.

Automated discovery of adaptive attacks improves adversarial defense evaluation.

problem Challenges in reliably evaluating adversarial defenses.
method Formalizes adaptive attacks as reusable building blocks in a search space for automatic discovery.
result Our tool discovers significantly stronger attacks than AutoAttack, improving adversarial defense evaluation.

ChemCrow enhances LLMs for chemistry tasks, automating complex chemical processes.

problem Limited access to computational chemistry tools for large-language models.
method Integrating 18 expert-designed chemistry tools into an LLM (ChemCrow).
result ChemCrow autonomously plans and executes chemical syntheses and discoveries.

Machine learning identifies physical laws from data, but discrepancies remain.

problem Discovering universal physical laws from data alone is challenging.
method Sparse Identification of Nonlinear Dynamics (SINDy) method to identify governing equations.
result Measurement noise and secondary physical mechanisms obscure the underlying law of gravitation.

System automates discovery and classification of training videos for career progression.

problem Difficulties in planning and navigating career paths due to changing job requirements and emerging sectors.
method Extracted educational videos, built a machine learning classifier, and optimized probability thresholds.
result Significant improvements in model performance by incorporating video attributes.

Study improves causal model discovery by relaxing assumptions for latent variables.

problem Discovering causal models with latent variables under weaker assumptions.
method Uses Answer Set Programming to discover semi-Markovian causal models with weakened Faithfulness assumption.
result Weakened Faithfulness assumption preserves power and speeds up discovery for causal models with latent variables.

Automated digital twin discovery from biological data improves drug discovery and personalized medicine.

problem Developing reliable digital twins from noisy, incomplete biological data.
method Symbolic and sparse regression, Bayesian frameworks, deep learning, and large language models.
result Sparse regression generally outperforms symbolic regression, especially with Bayesian frameworks.

AutoSciDACT detects scientific anomalies in noisy data.

problem Detecting anomalies in large, noisy scientific datasets.
method Contrastive pre-training for low-dimensional data representations, two-sample test using NPLM.
result Strong sensitivity to small anomalies across various scientific domains.

Automates discovering PDEs from data in dynamical systems.

problem Identifying PDEs from data in dynamical systems is challenging.
method ARGOS-RAL framework using sparse regression with recurrent adaptive lasso.
result ARGOS-RAL effectively identifies PDEs from noisy and non-uniformly distributed data.

GPR enhances materials discovery by automating parameter space exploration.

problem Automating exploration of large, high-dimensional parameter spaces in materials science.
method Gaussian process regression with inhomogeneous measurement noise and anisotropic kernels.
result Importance and benefits of tuning GPR for materials science experiments.

Bayesian methods improve drug discovery experiment design.

problem Optimizing drug screening experiments in high-dimensional data.
method Bayesian inference and optimisation with upper confidence bound algorithms, Thompson sampling, and sparse tree search.
result Sparse tree search techniques outperform other methods in drug toxicity screening.

Process mining is a research field focused on the analysis of event data with the aim of extracting insights in processes. Applying process mining techniques on data from smart home environments has the potential to provide valuable insights in (un)healthy habits and to contribute to ambient assisted living solutions. …

2016-09-12abs ↗pdf ↗

This paper analyzes and compares different Automated Market Maker mechanisms.

problem Impermanent loss in Constant Function Market Makers.
method Mean-Variance analysis of liquidity providers' profit and loss, comparison of different mechanisms.
result Optimized oracle-based mechanisms outperform Constant Function Market Makers.

Automated discovery of diverse self-organized patterns in complex systems.

problem Automated identification of interesting spatially localized patterns in self-organizing systems.
method Intrinsically motivated machine learning algorithms (POP-IMGEPs) combined with deep auto-encoders and CPPN primitives.
result Efficiency and effectiveness of the proposed method in discovering diverse patterns compared to baselines.

Automates organizing diverse web data into a hierarchical topic model.

problem Manual classification of all scientific and popular scientific knowledge is impractical.
method Proposes an algorithm to aggregate multiple collections into a single hierarchical topic model.
result Demonstrates a web service for topical exploratory search.

Machine discovers PDEs from spatiotemporal data without prior knowledge.

problem Discovering PDEs from complex spatiotemporal data without prior knowledge.
method Sparse Spatiotemporal System Discovery (extS3extd ext{S}^3 ext{d}) using Sparse Bayesian Learning.
result Automatically discovers ten types of PDEs from simulation data.

Machine learning predicts quantum advantage in noisy quantum walks.

problem Finding optimal graph types and coherence requirements for quantum advantage.
method Convolutional neural network trained on simulated examples of quantum walks on cycle graphs.
result Machine learning can predict quantum advantage for a wide range of decoherence parameters.

Automated GMA using accelerometers detects abnormal infant movements with human-level accuracy.

problem Undetected perinatal stroke leads to lifelong disability.
method Wearable accelerometers and Discriminative Pattern Discovery (DPD) for automated GMA.
result Automated method correctly identifies abnormal movements with human-level accuracy.

This thesis relaxes assumptions for causal discovery, making methods applicable to more complex systems.

problem Learning causal structures from observational data with latent variables.
method Alternative definition of k-Triangle Faithfulness for non-Gaussian distributions and uniform consistency proof.
result Uniform consistency of causal discovery algorithm under modified faithfulness assumption.

This study compares price discovery in ETH and BTC markets between centralized and decentralized exchanges.

problem Understanding price discovery dynamics in cryptocurrency markets.
method Comparative analysis of centralized and decentralized exchanges, using econometric tools.
result Centralized exchanges lead in ETH price discovery, while futures markets lead in BTC.

Automates kernel discovery for longitudinal data analysis.

problem Handling irregularly sampled, sparse longitudinal data with multilevel correlation.
method Combines deep neural networks and non-parametric kernel methods to discover complex multilevel correlation structure.
result Significantly outperforms state-of-the-art methods on benchmark data sets.

PySR method automates discovering equations from data in chaotic dynamics and epidemics.

problem Discovering equations from complex data in dynamical systems.
method Symbolic regression methods, focusing on PySR.
result PySR method efficiently infers equations from chaotic dynamics and epidemic models, matching original forms.

Generative AI improves stock selection by synthesizing features from diverse data sources.

problem Automating feature discovery in stock market data.
method Used large language models with retrieval-augmented generation and structured prompting to synthesize features from various data sources.
result AI-generated features consistently outperform baselines, with Sharpe improvements ranging from 14% to 91%.

Causal discovery algorithms can help generate legal arguments.

problem Leveraging causal discovery algorithms in legal decision-making.
method Prepared a legal dataset, annotated with 17 legal concepts, applied causal discovery algorithms, and quantified degrees of belief.
result Some causal relationships help generate viable legal arguments.

Enhances KANs for accuracy and interpretability with multi-exit architecture.

problem Unclear optimal depth for KANs and difficulty in optimization and interpretation.
method Introduces multi-exit KANs with each layer having its own prediction branch.
result Multi-exit KANs outperform single-exit versions on various datasets.

Unified platform for statistical and machine learning in bioinformatics.

problem Workflow inefficiencies in using multiple tools for data analysis.
method Automated hyperparameter optimization, feature importance analysis, statistical tests.
result Accelerates biological discovery workflows with methodological soundness.

ServeNet classifies web services without manual feature engineering.

problem Manual feature engineering limits the performance of conventional machine learning methods in web service classification.
method ServeNet uses a deep neural network to automatically abstract service names and descriptions into high-level features.
result ServeNet achieves higher accuracy and robustness in web service classification compared to other machine learning methods.

Derives token price process for AMM tokens, finds leverage effect and pricing discrepancies.

problem Derives token price process for AMM tokens.
method Derives CEV process for token price, derives closed-form option prices, introduces liquidity-adjusted Greeks.
result Token price process is CEV, with leverage effect and pricing discrepancies.