DDAT framework improves machine learning by dynamically adjusting difficulty.
problem Improving time-continuous emotion prediction models.
method Dynamic Difficulty Awareness Training (DDAT) framework.
result DDAT framework outperforms existing methods in emotion prediction.
New measure quantifies task difficulty for machine learning models.
problem Quantifying the inherent difficulty of machine learning tasks.
method Inductive bias complexity measure.
result Tasks requiring generalization over many dimensions are more difficult.
We create synthetic Morse code datasets for machine learning.
problem Creating challenging datasets for neural networks.
method Algorithm to generate synthetic Morse code datasets of varying difficulty.
result Network performance is affected by noise and feature set expansion.
The paper analyzes how curriculum learning improves machine learning performance.
problem Lack of theoretical analysis for curriculum learning in machine learning.
method Formulated an ideal difficulty score and analyzed its contribution in convex problems.
result The expected convergence rate decreases with the ideal difficulty score.
New theory shows perfect supervised learning is possible with enough data and complexity.
problem Theoretical limits of supervised learning tasks.
method Introducing bandlimiting into machine learning theory.
result Practical machine learning tasks are asymptotically solvable with sufficient data and model complexity.
Gradient boosting with randomized trees reduces discontinuities and complexity.
problem Discontinuities in regression functions due to sparse training data.
method Gradient boosting machine with partially randomized decision trees.
result Improves robustness and computational efficiency of gradient boosting.
Research uses CPS to estimate uncertainty in ML radio metric models.
problem Estimating uncertainty in machine learning models for radio metrics and path loss.
method Conformal Prediction (CP) in Conformal Predictive Systems (CPS) with diverse difficulty estimators.
result CPS models maintain high coverage and reliability across different cities.
A new method estimates uncertainty without explicit prediction models.
problem Costly data acquisition in machine learning.
method Distance-weighted Class Impurity method for uncertainty estimation.
result Distance-weighted Class Impurity effectively estimates uncertainty without prediction models.
Study predicts firm defaults using machine learning on Italian credit data.
problem Predicting firm defaults to inform bank lending policies.
method Used large granular credit data from Italian Central Credit Register, combined with public balance sheet data, and applied ensemble techniques and random forest models.
result Ensemble techniques and random forest provide the best results for predicting firm defaults.
Study on neural networks to identify redundancy issues in safe machine learning.
problem Identifying redundancy in neural network architectures for safe machine learning.
method Experiments with MNIST database using neural network classifiers.
result Underlines difficulties in using neural network classifiers for safe systems.
In this paper, we propose AutoCompete, a highly automated machine learning framework for tackling machine learning competitions. This framework has been learned by us, validated and improved over a period of more than two years by participating in online machine learning competitions. It aims at minimizing human interf…
This paper applies secure multi-party computation to K-means clustering to protect private data.
problem Privacy-preserving K-means clustering for distributed private data.
method Secure multi-party computation (MPC) techniques to protect private data during K-means clustering.
result Privacy-preserving K-means clustering is feasible and effective for both horizontal and vertical data distribution.
Improved PAC-Bayesian bounds by considering example difficulty.
problem Improving generalization bounds in machine learning.
method Introducing a modified excess risk that leverages example difficulty to reduce variance and tighten PAC-Bayesian bounds.
result Tighter PAC-Bayesian generalization bounds for machine learning models.
Machine learning risks in finance pricing and hedging
problem Understanding and managing risks in financial models
method Analyzing machine learning applications in finance, focusing on pricing and hedging of financial options
result Identifies various sources of risk and potential mitigation strategies
ProteinNet provides a standardized data set for protein structure prediction.
problem Lack of standardized data sets for protein structure prediction.
method Created high-quality sequence alignments, multiple data splits, and validation sets.
result Facilitates fair assessment of machine learning models for protein structure.
This article reviews datasets for COVID-19 detection using ML.
problem Lack of accessible datasets for COVID-19 detection research.
method Analyzed 96 papers on COVID-19 detection from January 2020 to June 2020.
result Identified and summarized datasets used in COVID-19 detection studies.
This work embeds annotations into a multidimensional space to measure classification difficulty.
problem Uncertainty in machine learning models during annotation phase.
method Develops a Bayesian Dirichlet-Multinomial framework to embed annotations and uses stochastic Expectation Maximization with MCMC.
result Embeddings reflect semantic similarities of original classes, aiding in measuring classification difficulty.
New β3-IRT model improves test performance and assesses classifier quality.
problem Improving test performance and assessing classifier quality.
method Proposes β3-IRT model to model continuous responses and assess latent abilities. result The β3-IRT model outperforms standard models on various datasets and provides a new metric for classifier evaluation. Paper provides conditions for reliable use of pre-trained embeddings in econometrics.
problem Uncertainty in using pre-trained embeddings for econometric tasks.
method Derives sufficient conditions and convergence rates for machine learning models with pre-trained embeddings.
result Establishes theoretical foundations for reliable use of pre-trained embeddings in econometrics.
This study compares machine learning methods for high-cardinality categorical variables.
problem Machine learning struggles with high-cardinality categorical variables.
method Empirical comparison of tree-boosting, deep neural networks, and linear mixed effects models.
result Tree-boosting with random effects outperforms deep neural networks with random effects.
New blockchain metrics improve cryptocurrency trading and prediction.
problem Improving trading and prediction in the volatile cryptocurrency market.
method Developed blockchain metrics based on public data from Bitcoin mining nodes.
result Blockchain metrics provide statistical advantage in trading Bitcoin assets.
Discussing AI's difficulty and physics' simplicity, suggesting AI benefits from physics principles.
problem AI difficulty compared to physics simplicity.
method Drawing on physical intuition and theoretical physics to improve AI.
result AI and physics principles are strongly coupled through sparsity.
Study examines challenges and applications of machine learning in finance.
problem Challenges in applying machine learning to financial research due to market idiosyncrasies and methodological differences.
method Discussion of adjustments needed to conventional machine learning methodology to account for financial market peculiarities.
result Machine learning can be unified with financial research as a robust complement to econometric methods.
End-to-end framework learns from imperfect annotations directly.
problem Training machine learning models on imperfect human annotations.
method End-to-end framework merging aggregation with model training and modeling annotator competencies.
result Accuracy gains of up to 25% over state-of-the-art annotation aggregation methods.
The paper builds interpretable models for property markets using machine learning.
problem Noise in real market data and differences from ideal data.
method Combining classical linear regression with kriging for land parcels, and RuleFit method for flats.
result Effective models can be built for property markets while maintaining interpretability.
Paper tackles fairness in insurance machine learning models using active learning.
problem Reducing labeling effort and promoting fairness in insurance machine learning.
method Introduces a fair active learning method to sample informative and fair instances.
result Achieves a balance between model performance and fairness in insurance datasets.
Survey of machine learning methods for Windows malware classification.
problem Difficulties in malware classification through data collection, labeling, feature creation, and selection.
method Review of current methods and challenges in malware classification.
result Discussion of constraints and unaddressed problems for machine learning in cybersecurity.
Machine learning enhances wireless network authentication for diverse devices.
problem Complex dynamic wireless environments challenge conventional authentication methods.
method Intelligent authentication using machine learning for diverse physical layer attributes.
result Machine learning-based authentication provides cost-effective, reliable, and situation-aware security.
This paper explores how machine learning can improve life insurance risk assessment.
problem Limited use of machine learning in life insurance due to statistical models' efficiency.
method Review and extension of traditional actuarial methodologies with machine learning techniques.
result Developed Python library for life insurance data, improving risk modeling.
Method for explaining machine learning survival models using counterfactuals.
problem Tackles the challenge of explaining survival models in machine learning.
method Introduces a condition based on the difference of mean times to event for counterfactual explanation. Reduces the problem to a convex optimization problem for Cox models and applies Particle Swarm Optimization for other models.
result Demonstrates the effectiveness of the proposed method through numerical experiments.
This work sets theoretical limits on meta-learning performance.
problem Understanding the difficulty of adapting machine learning models to real-world data distributions.
method Information-theoretic lower bounds on convergence rates for meta-learning algorithms.
result Theoretical bounds on parameter estimation error for hierarchical Bayesian models of meta-learning.
Novel ML approach optimizes large portfolios without covariance matrix issues.
problem Static and dynamic portfolio optimization for many assets.
method Machine learning for constrained optimization, avoiding covariance matrix computation.
result Significant excess returns in U.S. and China equity markets.
ML improves flood forecasting by leveraging local data.
problem Human calibration, limited data, and computational difficulty.
method Transfer learning and ML for high-dimensional scenarios.
result ML systems achieve timely and accurate flood prediction.
Deep networks prioritize easier examples over harder ones, leading to faster training.
problem Understanding how deep networks prioritize examples of varying difficulty.
method Investigated the effect of linear vs non-linear learning modes on example difficulty.
result Non-linear dynamics tend to sequentialize the learning of examples of increasing difficulty.
Improved Monte Carlo simulations using RBMs for phase transitions.
problem Slow mixing times in Monte Carlo simulations for complex systems.
method Fit unnormalized probability to a restricted Boltzmann machine and use its feature detection ability for efficient updates.
result Improved acceptance ratio and autocorrelation time near phase transition points.
Unified framework for interpreting complex regression models with many predictors.
problem Interpreting nonparametric regression models with many predictors.
method Derivative-based approach for existing tools like partial-dependence plots.
result New technique called accumulated total derivative effects plot for complex models.
This paper critiques machine learning fairness from a consequentialist perspective.
problem The ethical foundations of fairness in machine learning.
method Consequentialist analysis of existing fairness definitions and machine learning perspectives.
result Consequentialism highlights ethical tradeoffs in fairness definitions and decision making.
AEFS selects features from high-dimensional data using autoencoders.
problem Feature selection for high-dimensional data in computer vision and machine learning.
method Combines autoencoder regression and group lasso for unsupervised feature selection.
result AEFS selects more important features than traditional methods, including linear and nonlinear information.
Machine learning bypasses Kohn-Sham equations for faster DFT calculations.
problem Solving the Kohn-Sham equations for electronic structure problems.
method Directly learning density-potential and energy-density maps for test systems and molecules.
result Improved accuracy and lower computational cost demonstrated for molecular geometries.
Artificial neural networks are simple and efficient machine learning tools. Defined originally in the traditional setting of simple vector data, neural network models have evolved to address more and more difficulties of complex real world problems, ranging from time evolving data to sophisticated data structures such …
Proposes an IRT-based ensemble method to improve machine learning accuracy.
problem Improving the accuracy of machine learning models, especially for hard-to-classify instances.
method Introduces Item Response Theory (IRT) to evaluate sample difficulty and classifier ability, creating three models with different assumptions.
result The proposed IRT ensemble model outperforms other methods on 19 datasets.
InfoBridge uses diffusion bridges to estimate mutual information accurately.
problem Estimating mutual information between random variables.
method Formulated mutual information estimation as a domain transfer problem using diffusion bridge models.
result Demonstrated unbiased estimator for various data types.
Machine learning improves ice flow tracking in satellite images.
problem Improving accuracy of ice flow tracking in multi-spectral satellite images.
method Adversarial learning method to predict future ice flow.
result Adversarial learning improves ice flow tracking accuracy.
Tutorials on optimization methods for machine learning problems.
problem Solving supervised machine learning problems using optimization methods.
method Discusses various optimization problems and algorithms for machine learning, including logistic regression and deep neural networks.
result Explains the challenges and approaches for training deep neural networks.
Machine learning method characterizes network interference in A/B tests.
problem Compromised A/B test reliability due to network interference.
method Causal network motifs and machine learning models.
result Outperforms conventional methods in characterizing network interference.
Subpopulation attacks poison data to misclassify naturally distributed points.
problem Improving accuracy of machine learning predictions through adversarial data modification.
method Introducing a novel subpopulation attack framework, using influence functions and gradient optimization.
result Subpopulation attacks are effective and stealthy, making them difficult to defend against.
We propose SPARFA-Trace, a new machine learning-based framework for time-varying learning and content analytics for education applications. We develop a novel message passing-based, blind, approximate Kalman filter for sparse factor analysis (SPARFA), that jointly (i) traces learner concept knowledge over time, (ii) an…
Study uses LCA to identify ARDS sub-phenotypes improving predictive models.
problem Complex and heterogeneous nature of ARDS makes early recognition difficult.
method Applied latent class analysis to identify sub-groups, then built predictive models.
result Significantly improved prediction performance for two sub-phenotypes of ARDS.