Iterative subtraction method outperforms other feature ranking techniques in high-energy physics.
problem Determining the most important features for classification in high-energy physics experiments.
method Comparison of feature ranking methods including Iterative Addition, Iterative Removal, and BDT Selection Frequency.
result Iterative Removal method is the most efficient for feature ranking in classification tasks.
Modified BDT model includes zero rate jumps for crisis risk.
problem Quantify risk of future crises in bond prices and derivatives.
method Modified BDT tree model with zero rate jumps, modified algorithms.
result Calibrated tree models show different option prices and volatilities.
AI model identifies chemical agents in MCI with high accuracy.
problem Identifying chemical agents in mass casualty incidents.
method Reverse engineered signs/symptoms, trained using ANN, BDT, and WISER.
result WISER outperformed ANN and BDT in identifying chemical agents.
Study improves forex forecasting accuracy using machine learning models.
problem Improving accuracy in predicting foreign exchange rates.
method Employed LSTM neural networks and Gradient Boosting Classifier for forecasting.
result Achieved 99.449% accuracy in forecasting USD/BDT exchange rates.
GDT improves reinforcement learning by matching future state information efficiently.
problem Efficient learning of multi-task policies from trajectory data.
method Generalized Decision Transformer (GDT) for offline hindsight information matching.
result GDT enables effective offline multi-task state-marginal matching and imitation learning.
Machine learning tools are commonly used in modern high energy physics (HEP) experiments. Different models, such as boosted decision trees (BDT) and artificial neural networks (ANN), are widely used in analyses and even in the software triggers. In most cases, these are classification models used to select the "signal"…
Study compares ZBDT model to BDT for financial derivatives valuation.
problem Valuation of financial derivatives under catastrophic events.
method Introduced Zero Black-Derman-Toy (ZBDT) model with jumps to zero interest rate.
result ZBDT model better matches financial slowdown risk.
The study analyzes when Bayesian averaging over decision trees is reliable.
problem When do Bayesian model averaging weights over decision trees provide reliable information?
method Closed-form solution for Bayesian decision trees with Catalan-exponential priors.
result Established a complete non-asymptotic theory of rational commitment thresholds.
New method solves stochastic control problems with delays using deep learning.
problem Stochastic control problems with delayed control in drift and diffusion.
method Characterization via Riccati PDEs and deep learning scheme.
result Illustrates effect of delay on Markowitz portfolio allocation problem.
Boosted decision trees improved for particle identification in high-energy physics.
problem Overfitting in boosted decision trees hampers their performance in particle identification.
method Meta-learning techniques of boosting and bagging to mitigate overfitting.
result The proposed algorithm achieves performance close to that of deep neural networks on a benchmark data set.
Since the machine learning techniques are improving rapidly, it has been shown that the image recognition techniques in deep neural networks can be used to detect jet substructure. And it turns out that deep neural networks can match or outperform traditional approach of expert features. However, there are disadvantage…
Dual-stage sEMG classification improves gesture recognition accuracy.
problem Improving accuracy in hand gesture recognition from sEMG signals.
method Dual-stage classification approach: first stage groups similar activities, second stage classifies within groups.
result Dual-stage classification yields significantly higher accuracy than single-stage approach.
A novel method for classification with rejection using ensemble of cost-sensitive classifiers.
problem Avoid risky misclassification in error-critical applications.
method Learning an ensemble of cost-sensitive classifiers.
result Improved classification accuracy and flexibility in loss selection.
The number of possible methods of generalizing binary classification to multi-class classification increases exponentially with the number of class labels. Often, the best method of doing so will be highly problem dependent. Here we present classification software in which the partitioning of multi-class classification…
Sequence classification is an important data mining task in many real world applications. Over the past few decades, many sequence classification methods have been proposed from different aspects. In particular, the pattern-based method is one of the most important and widely studied sequence classification methods in …
Few-shot image classification is improved by correcting CNNs' texture bias.
problem Few-shot image classification performance is hindered by CNNs' texture bias.
method Corrected CNNs' texture bias using a simpler method than state-of-the-art approaches.
result State-of-the-art performance on miniImageNet task achieved.
New NHCAs improve multi-category classification efficiency.
problem Efficient multi-category classification for real-world problems.
method Twin SVM (TWSVM), Generalized eigenvalue proximal SVM (GEPSVM), Regularized GEPSVM (RegGEPSVM), and Improved GEPSVM (IGEPSVM) with OAA, BT, and TDS approaches.
result TDS-TWSVM outperforms other methods in classification accuracy.
Paper compares XGB and BPNN for music style classification.
problem Efficient music style classification using different methods.
method Feature extraction for timbral texture, rhythmic content, and pitch content; comparative evaluation of XGB and BPNN.
result XGB outperforms BPNN for small datasets in music classification.
Classification outperforms regression in portfolio construction, yielding higher Sharpe ratios.
problem Determining which machine learning approach (classification vs. regression) is more effective for portfolio construction.
method Used stacking ensemble of gradient boosted tree, random forest, and neural network models.
result Classification yields higher Sharpe ratios and economically significant alphas compared to regression.
C-HMCNN(h) improves HMC classification by leveraging class hierarchy.
problem Hierarchical multi-label classification with class hierarchy constraints.
method Exploits class hierarchy to produce coherent predictions for multi-label classification.
result C-HMCNN(h) outperforms state-of-the-art models in HMC classification.
Advances few-shot classification by treating it as supervised learning and proposing new training techniques.
problem Formulating the ability of humans to learn from limited data in machine learning.
method Formulated few-shot classification as a supervised learning problem and introduced multi-episode and cross-way training techniques.
result Proposed training strategies accelerate the training process without accuracy loss.
Improves NILM with multi-label SRC, outperforming state-of-the-art.
problem Non-intrusive load monitoring (NILM) for energy disaggregation.
method Modified multi-label sparse representation based classification (SRC).
result Significant improvement over state-of-the-art techniques with minimal training data.
New approach improves classification guarantees by focusing on direction rather than regression risk.
problem Improving classification guarantees in binary classification problems.
method Establishing a geometric distinction between classification and regression, leveraging scale invariance.
result Improved guarantees for classification risk compared to regression risk.
We study realizations of Lie algebras by vector fields. A correspondence between classification of transitive local realizations and classification of subalgebras is generalized to the case of regular local realizations. A reasonable classification problem for general realizations is rigorously formulated and an algori…
Study selective classification with halfspaces, achieving error bounds under Gaussian distributions.
problem Modeling relationships in subsets of data defined by selection rules.
method Sparse linear classifiers for subsets defined by halfspaces, focusing on Gaussian feature distributions.
result First PAC-learning algorithm for homogeneous halfspace selectors with error guarantee $\bigO*{\sqrt{\mathrm{opt}}}$.
A new network-based high-level data classification method using betweenness centrality.
problem Traditional data classification techniques focus on physical features, while high-level classification considers semantic meaning.
method Proposes a network-based high-level classification technique using betweenness centrality.
result Competent classification performance in nine real datasets compared to traditional models.
This thesis evaluates text-based vs audio-based classification of mental health interviews.
problem Classifying psychiatric illness using text-based methods.
method Design and evaluate a text classification network on mental health interviews, using belabBERT.
result Text-based classification is a strong alternative to audio-based methods.
Study on error probability for classification of heavy-tailed renewal processes.
problem Error probability in classification of heavy-tailed renewal processes.
method Asymptotic expressions for Bhattacharyya bound on misclassification error probabilities.
result Obtained asymptotic expressions for misclassification error probabilities.
This review explores resampling techniques for imbalanced binary classification.
problem Imbalanced classes lead to poor prediction results in classification.
method Classical, cost-sensitive, and Neyman-Pearson paradigms with resampling techniques and classification methods.
result Complex dynamics among resampling techniques, base methods, metrics, and imbalance ratios.
Directly compute classification by learning features with class scores.
problem Classification efficiency and accuracy on various datasets.
method PCA for feature encoding, supervised learning model with encoder-decoder structure.
result Effective classification performance on multiple datasets.
Conventional techniques for supervised classification constrain the classification rules considered and use surrogate losses for classification 0-1 loss. Favored families of classification rules are those that enjoy parametric representations suitable for surrogate loss minimization, and low complexity properties suita…
This paper tackles tweet classification by identifying purpose and position.
problem Difficulties in determining user intention and attitude in short, informal tweets.
method Transformed tweet classification into a multi-label problem and applied a multi-label classification method with post-processing.
result The method effectively classifies tweet purpose and position, outperforming individual classification methods.
Text classification on drug SMILES strings yields competitive drug type classification results.
problem Classifying drug types using conventional text classification methods.
method Treated drug SMILES as sentences and applied basic NLP methods for classification.
result Competitive drug type classification results achieved.
BERT improves fine-grained sentiment classification.
problem Fine-grained sentiment classification of text.
method Used BERT for fine-grained sentiment classification.
result BERT outperforms other models for fine-grained sentiment classification.
Paper proposes fully Bayesian approach for RVM classification, improving accuracy especially in imbalanced data.
problem Difficulty in conducting RVM classification due to lack of closed-form solution for weight parameter posterior.
method Proposes Generic Bayesian and Fully Bayesian approaches with hierarchical hyperprior structure.
result Improves classification performance, especially in imbalanced data.
CRCEN neural network tackles imbalanced classification.
problem Challenges in training conventional classifiers on imbalanced datasets.
method CRCEN neural network with a novel weighted cross entropy loss function.
result CRCEN outperforms baseline models on benchmark datasets.
This work bounds classification error in machine learning for low Bayes error conditions.
problem Understanding the error mismatch between Bayes error and model-based classification error.
method Applying classification error bounds to study the relationship with Kullback-Leibler divergence and proposing a linear approximation for low Bayes error conditions.
result A linear approximation of the classification error bound for low Bayes error conditions is proposed.
Proposes an angle-based framework for multicategory cost-sensitive classification.
problem Cost-sensitive multicategory classification challenges.
method Angle-based cost-sensitive classification framework without sum-to-zero constraint.
result Proposed boosting algorithms yield competitive classification performances.
Optimal fuzzy classification aggregation functions are weighted means.
problem Characterizing optimal fuzzy classification aggregation functions.
method Proving optimality of weighted arithmetic means for fuzzy classification.
result Optimal fuzzy classification aggregation functions are weighted means.
Set classification problems arise when classification tasks are based on sets of observations as opposed to individual observations. In set classification, a classification rule is trained with N sets of observations, where each set is labeled with class information, and the prediction of a class label is performed a…
Interactive tool helps choose and understand classification metrics.
problem Common metrics for binary classification have limitations.
method Graphical application to visualize and explore evaluation metrics.
result Promotes careful attention to interpretation of metrics.
Complete classification of symmetric spaces actions.
problem Classifying symmetric spaces actions.
method Isometric cohomogeneity-one actions classification.
result Classification completed for all symmetric spaces of noncompact type.
Paper reduces neural network complexity for image classification.
problem High computational complexity in deep neural networks.
method Proposes a two-step classification process: coarse-grain and fine-grain.
result Achieves similar accuracy with less computational complexity.
MODWST improves classification tasks with wavelet scattering.
problem Signal classification challenges.
method Combines MODWT and WST for feature extraction.
result MODWST outperforms CNNs in limited data scenarios.
This paper reviews metrics for evaluating multi-class classification models.
problem Evaluating and comparing multi-class classification models.
method Review and analysis of metrics.
result Promising multi-class metrics are highlighted and their usages are demonstrated.
Data in real-world application often exhibit skewed class distribution which poses an intense challenge for machine learning. Conventional classification algorithms are not effective in the case of imbalanced data distribution, and may fail when the data distribution is highly imbalanced. To address this issue, we prop…
Debiasing techniques can worsen gender bias in text classification, but a tweak improves both.
problem Debiasing techniques can inadvertently increase gender bias in text classification.
method Investigated traditional debiasing techniques and found they worsen bias. Suggested a minor adjustment.
result A minor adjustment to debiasing techniques can reduce gender bias while maintaining high classification accuracy.
The paper classifies bundles over complex projective plane.
problem Classifying S3-bundles over CP2. method Two-step approach: PL-homeomorphism classification via Kreck-Stolz invariants, followed by homotopy equivalence classification using surgery theory.
result Established the homotopy equivalence classification of S3-bundles over CP2.