Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

3977931,1901,586 · Jun 202019922001200920182026
48 results for machine learning difficulty

The paper analyzes how curriculum learning improves machine learning performance.

problem Lack of theoretical analysis for curriculum learning in machine learning.
method Formulated an ideal difficulty score and analyzed its contribution in convex problems.
result The expected convergence rate decreases with the ideal difficulty score.

Research uses CPS to estimate uncertainty in ML radio metric models.

problem Estimating uncertainty in machine learning models for radio metrics and path loss.
method Conformal Prediction (CP) in Conformal Predictive Systems (CPS) with diverse difficulty estimators.
result CPS models maintain high coverage and reliability across different cities.

Study predicts firm defaults using machine learning on Italian credit data.

problem Predicting firm defaults to inform bank lending policies.
method Used large granular credit data from Italian Central Credit Register, combined with public balance sheet data, and applied ensemble techniques and random forest models.
result Ensemble techniques and random forest provide the best results for predicting firm defaults.

In this paper, we propose AutoCompete, a highly automated machine learning framework for tackling machine learning competitions. This framework has been learned by us, validated and improved over a period of more than two years by participating in online machine learning competitions. It aims at minimizing human interf…

2015-07-08abs ↗pdf ↗

This paper applies secure multi-party computation to K-means clustering to protect private data.

problem Privacy-preserving K-means clustering for distributed private data.
method Secure multi-party computation (MPC) techniques to protect private data during K-means clustering.
result Privacy-preserving K-means clustering is feasible and effective for both horizontal and vertical data distribution.

This work embeds annotations into a multidimensional space to measure classification difficulty.

problem Uncertainty in machine learning models during annotation phase.
method Develops a Bayesian Dirichlet-Multinomial framework to embed annotations and uses stochastic Expectation Maximization with MCMC.
result Embeddings reflect semantic similarities of original classes, aiding in measuring classification difficulty.

New β3β^3-IRT model improves test performance and assesses classifier quality.

problem Improving test performance and assessing classifier quality.
method Proposes β3β^3-IRT model to model continuous responses and assess latent abilities.
result The β3β^3-IRT model outperforms standard models on various datasets and provides a new metric for classifier evaluation.

Paper provides conditions for reliable use of pre-trained embeddings in econometrics.

problem Uncertainty in using pre-trained embeddings for econometric tasks.
method Derives sufficient conditions and convergence rates for machine learning models with pre-trained embeddings.
result Establishes theoretical foundations for reliable use of pre-trained embeddings in econometrics.

This study compares machine learning methods for high-cardinality categorical variables.

problem Machine learning struggles with high-cardinality categorical variables.
method Empirical comparison of tree-boosting, deep neural networks, and linear mixed effects models.
result Tree-boosting with random effects outperforms deep neural networks with random effects.

Study examines challenges and applications of machine learning in finance.

problem Challenges in applying machine learning to financial research due to market idiosyncrasies and methodological differences.
method Discussion of adjustments needed to conventional machine learning methodology to account for financial market peculiarities.
result Machine learning can be unified with financial research as a robust complement to econometric methods.

The paper builds interpretable models for property markets using machine learning.

problem Noise in real market data and differences from ideal data.
method Combining classical linear regression with kriging for land parcels, and RuleFit method for flats.
result Effective models can be built for property markets while maintaining interpretability.

Survey of machine learning methods for Windows malware classification.

problem Difficulties in malware classification through data collection, labeling, feature creation, and selection.
method Review of current methods and challenges in malware classification.
result Discussion of constraints and unaddressed problems for machine learning in cybersecurity.

Machine learning enhances wireless network authentication for diverse devices.

problem Complex dynamic wireless environments challenge conventional authentication methods.
method Intelligent authentication using machine learning for diverse physical layer attributes.
result Machine learning-based authentication provides cost-effective, reliable, and situation-aware security.

This paper explores how machine learning can improve life insurance risk assessment.

problem Limited use of machine learning in life insurance due to statistical models' efficiency.
method Review and extension of traditional actuarial methodologies with machine learning techniques.
result Developed Python library for life insurance data, improving risk modeling.

Method for explaining machine learning survival models using counterfactuals.

problem Tackles the challenge of explaining survival models in machine learning.
method Introduces a condition based on the difference of mean times to event for counterfactual explanation. Reduces the problem to a convex optimization problem for Cox models and applies Particle Swarm Optimization for other models.
result Demonstrates the effectiveness of the proposed method through numerical experiments.

This work sets theoretical limits on meta-learning performance.

problem Understanding the difficulty of adapting machine learning models to real-world data distributions.
method Information-theoretic lower bounds on convergence rates for meta-learning algorithms.
result Theoretical bounds on parameter estimation error for hierarchical Bayesian models of meta-learning.

Novel ML approach optimizes large portfolios without covariance matrix issues.

problem Static and dynamic portfolio optimization for many assets.
method Machine learning for constrained optimization, avoiding covariance matrix computation.
result Significant excess returns in U.S. and China equity markets.

Deep networks prioritize easier examples over harder ones, leading to faster training.

problem Understanding how deep networks prioritize examples of varying difficulty.
method Investigated the effect of linear vs non-linear learning modes on example difficulty.
result Non-linear dynamics tend to sequentialize the learning of examples of increasing difficulty.

Improved Monte Carlo simulations using RBMs for phase transitions.

problem Slow mixing times in Monte Carlo simulations for complex systems.
method Fit unnormalized probability to a restricted Boltzmann machine and use its feature detection ability for efficient updates.
result Improved acceptance ratio and autocorrelation time near phase transition points.

Unified framework for interpreting complex regression models with many predictors.

problem Interpreting nonparametric regression models with many predictors.
method Derivative-based approach for existing tools like partial-dependence plots.
result New technique called accumulated total derivative effects plot for complex models.

AEFS selects features from high-dimensional data using autoencoders.

problem Feature selection for high-dimensional data in computer vision and machine learning.
method Combines autoencoder regression and group lasso for unsupervised feature selection.
result AEFS selects more important features than traditional methods, including linear and nonlinear information.

Machine learning bypasses Kohn-Sham equations for faster DFT calculations.

problem Solving the Kohn-Sham equations for electronic structure problems.
method Directly learning density-potential and energy-density maps for test systems and molecules.
result Improved accuracy and lower computational cost demonstrated for molecular geometries.

Artificial neural networks are simple and efficient machine learning tools. Defined originally in the traditional setting of simple vector data, neural network models have evolved to address more and more difficulties of complex real world problems, ranging from time evolving data to sophisticated data structures such …

2012-10-24abs ↗pdf ↗

Proposes an IRT-based ensemble method to improve machine learning accuracy.

problem Improving the accuracy of machine learning models, especially for hard-to-classify instances.
method Introduces Item Response Theory (IRT) to evaluate sample difficulty and classifier ability, creating three models with different assumptions.
result The proposed IRT ensemble model outperforms other methods on 19 datasets.

Tutorials on optimization methods for machine learning problems.

problem Solving supervised machine learning problems using optimization methods.
method Discusses various optimization problems and algorithms for machine learning, including logistic regression and deep neural networks.
result Explains the challenges and approaches for training deep neural networks.

Subpopulation attacks poison data to misclassify naturally distributed points.

problem Improving accuracy of machine learning predictions through adversarial data modification.
method Introducing a novel subpopulation attack framework, using influence functions and gradient optimization.
result Subpopulation attacks are effective and stealthy, making them difficult to defend against.

We propose SPARFA-Trace, a new machine learning-based framework for time-varying learning and content analytics for education applications. We develop a novel message passing-based, blind, approximate Kalman filter for sparse factor analysis (SPARFA), that jointly (i) traces learner concept knowledge over time, (ii) an…

2013-12-19abs ↗pdf ↗

Study uses LCA to identify ARDS sub-phenotypes improving predictive models.

problem Complex and heterogeneous nature of ARDS makes early recognition difficult.
method Applied latent class analysis to identify sub-groups, then built predictive models.
result Significantly improved prediction performance for two sub-phenotypes of ARDS.