Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

117235352469 · Jun 202019922001200920182026
48 results for applied ML literature

A new ML-based filter improves data assimilation for nonlinear systems.

problem Improving data assimilation for nonlinear systems using ensemble methods.
method Developed a machine learning-based conditional mean filter (ML-EnCMF) integrating ANN and linear functions.
result ML-EnCMF outperforms EnKF and likelihood-based EnCMF in nonlinear systems.

This guide simplifies applying differential privacy to machine learning models.

problem Limited practical guidance for achieving good privacy-utility-computations in ML models.
method Comprehensive self-contained guide covering theory and practical implementation.
result Achieves best possible DP ML model with rigorous privacy guarantees.

Sage platform protects ML models trained on sensitive data from leakage.

problem Protecting sensitive data in machine learning models exposed to untrusted domains.
method Develops block composition for privacy accounting and privacy-adaptive training to manage privacy budget and utility tradeoff.
result Enables continuous training of models on sensitive data streams while maintaining global DP guarantees.

99% of papers use real-world data, but only 3% provide formal comparisons.

problem Lack of complete argumentative chains in demonstrating algorithmic effectiveness in machine learning papers.
method Systematic review of NeurIPS papers from 2017, assessing completeness of argumentative steps.
result Only 3% of papers provide formal comparisons, indicating incomplete argumentative chains.

Integrates ML with operations knowledge to improve distributional forecasts in healthcare.

problem Challenges of ML in operational settings, especially lack of distributional information and integration of operations literature.
method Introduces Boosted Generalized Normal Distribution (bbGND) using gradient boosting with tree learners.
result Improves wait and service time forecasting by 6% and 9% compared to ML benchmarks.

This study examines how data types affect ML algorithms' performance in Bitcoin price prediction.

problem Improving the accuracy of Bitcoin price forecasts for financial gain.
method Constructed continuous and trend data from Bitcoin's historical data, applied various ML algorithms, and compared their performance using accuracy and AUC.
result Data type significantly impacts ML algorithms' performance in Bitcoin price prediction.

PD-ML-Lite uses lightweight cryptography for private distributed machine learning.

problem Privacy issues in learning from distributed data.
method Applying lightweight cryptographic protocols to build learning algorithms.
result Achieves the same accuracy as non-private methods while maintaining privacy.

A rigorous ML pipeline for binary classification in biomedical studies, focusing on pancreatic cancer.

problem Handling bias in ML models for complex biomedical data.
method Customizable ML analysis pipeline with 9 algorithms, hyperparameter optimization, and thorough evaluation.
result Comparison of ML algorithms to ExSTraCS, highlighting interpretability and bias handling.

This paper provides a ML framework for diabetes prediction and care management.

problem Diabetes prediction and care management challenges in real-world healthcare.
method Illustrates a Machine Learning framework for T2DM prediction and risk stratification.
result ML models align with physician's disease management steps.

Machine learning improves portfolio optimization by reducing estimation risk.

problem Suboptimal portfolio choices due to estimation risk in traditional methods.
method Using machine learning to estimate optimal portfolio weights from asset returns.
result Machine learning significantly reduces estimation risk compared to traditional methods.

Survey on uncertainty in ML and DL, covering sources, quantification, and decision-making.

problem Understanding and quantifying uncertainty in ML and DL for risk-sensitive applications.
method Structured review of literature, categorizing uncertainty, assessing uncertainty quantification techniques.
result Broadened scope of uncertainty discussion and updated DL uncertainty quantification methods.

Adversarial attacks can fool ML energy theft detection models.

problem Vulnerability of ML-based energy theft detection models to adversarial attacks.
method Design of an adversarial measurement generation algorithm.
result ML models can be significantly fooled by adversarial attacks, reducing their detection accuracy.

This work explores using deep NNs to learn quantum systems from probability distributions.

problem Learning quantum systems from limited probability distribution data.
method Using deep neural networks to reconstruct quantum Hamiltonian from probability distributions.
result Deep neural networks can learn quantum Hamiltonians from probability distributions.

Develops Active Fourier Auditor to estimate ML model properties without reconstructing them.

problem Verifying and auditing properties of Machine Learning models in real-world applications.
method A new framework that quantifies ML model properties using Fourier coefficients, without reconstructing the model.
result Active Fourier Auditor (AFA) is more accurate and sample-efficient than baselines for estimating robustness, individual fairness, and group fairness.

This work investigates the impact of staleness in distributed ML systems and offers insights into convergence.

problem The effects of staleness on the convergence of distributed machine learning algorithms are inconclusive and challenging to monitor.
method Extensive experiments with various ML models and algorithms under delayed updates.
result The empirical findings reveal the diverse effects of staleness on ML algorithm convergence and match the best-known convergence rate.

The study examines how experimental design choices affect machine learning model performance.

problem Lack of guidelines on choosing experimental designs and machine learning models.
method 12 experimental designs, 7 families of predictive models, 7 test functions, 8 noise settings.
result Guidelines for practical applications of DOE and ML are provided.

This paper surveys fairness notions in ML and recommends the most suitable one for real-world scenarios.

problem Ensuring ML systems do not discriminate against specific individuals or sub-populations.
method Identifying fairness-related characteristics of real-world scenarios and analyzing the behavior of fairness notions.
result A decision diagram to recommend the most suitable fairness notion for specific setups.

Machine learning confound removal biases results, leading to misleading predictions.

problem Common confound removal methods in machine learning lead to misleading predictions.
method Featurewise removal of confound variance by linear regression before applying ML.
result This common deconfounding approach can leak information, amplifying null or moderate effects.

Develops fair feature importance scores for tree-based models to interpret fairness.

problem Ensuring fairness in machine learning models, especially tree-based ones.
method Inspired by decision trees, proposes a novel fair feature importance score based on mean decrease in group bias.
result Valid interpretations of fairness for tree-based ensembles and surrogates of other ML systems.

Paper compares ML methods for credit scoring, highlighting feature selection and scaling impacts.

problem Determining default risk in credit scoring models.
method Eight ML methods (SVM, Naive Bayes, DT, RF, XGBoost, KNN, MLP, LR) with feature selection and scaling.
result Feature selection and scaling improve model performance in credit scoring.

Paper proposes a new ML framework to enhance creativity in music generation.

problem Current generative models struggle to produce music outside their training dataset.
method Develops a new ML objective to address creativity limitations.
result Proposed framework could improve generative models' creativity.

Researchers validate ML scenario generators by checking dependencies and detecting memorization effects.

problem Validation of machine learning-based scenario generators differs from classical methods due to data-driven dependencies.
method Two novel validation aspects: checking dependencies and detecting memorization effects. Novel memorization ratio introduced.
result Validation methods successfully detect dependencies and memorization effects in ML-based scenario generators.