Enhances multi-tag classification using low-dimensional vector representations and virtual data.
problem Improving the performance of multi-tag classifiers.
method Embedding raw data into a low-dimensional feature space, then generating virtual data from linear operations on these vectors, to train multi-tag classifiers.
result Significant improvement in F1 scores (up to 224%) compared to training directly with raw data.
This paper provides new insight into maximizing F1 scores in the context of binary classification and also in the context of multilabel classification. The harmonic mean of precision and recall, F1 score is widely used to measure the success of a binary classifier when one class is rare. Micro average, macro average, a…
A new stable similarity measure for time series using persistent homology.
problem Constructing a robust measure of time series similarity.
method Persistent homology for stability, bi-conditional periodicity score for similarity.
result Stability of the bi-conditional periodicity score under perturbations and dimension reduction.
Study improves radio show segmentation using audio embeddings.
problem Automated segmentation of radio shows.
method Created audio embeddings from multi-class classification tasks on different datasets, evaluated performance against text-only baseline.
result Audio embeddings from non-speech sound event classification significantly outperformed text-only baseline by 32.3% in F1-measure.
A deep learning algorithm for ECG segmentation.
problem ECG signal segmentation for various sampling rates and monitors.
method UNet-like neural network for adaptive and generalized ECG segmentation.
result F1-measures for ECG segmentation are at least 97.8%, 99.5%, and 99.9%.
Proposes MCC-F1 curve for better binary classification evaluation.
problem Misleading performance evaluations with ROC and PR curves for imbalanced data.
method Introduces MCC-F1 curve combining MCC and F1 score.
result MCC-F1 curve provides clearer classifier differentiation.
Communities in social networks or graphs are sets of well-connected, overlapping vertices. The effectiveness of a community detection algorithm is determined by accuracy in finding the ground-truth communities and ability to scale with the size of the data. In this work, we provide three contributions. First, we show t…
Improved classifier for PU data using logistic regression.
problem Analysis of Positive Unlabeled data under SCAR assumption.
method Fitting misspecified logistic regression model to PU data.
result The classifier performs on par or better than competitors on real data sets.
The study evaluates AI model performance measures for medical use.
problem Selecting appropriate performance measures for AI models in medical practice.
method Assessed 32 performance measures across five domains for binary outcomes.
result 17 measures are both proper and reflect decision-analytic performance.
Simplifies F-measure for better interpretability.
problem Lack of intuitive interpretation of F-measure.
method Introduces F* (F-star) transformation.
result F* provides an immediate practical interpretation.
Neural network for water treatment anomaly detection with GA architecture optimization.
problem Detect anomalies in water treatment systems.
method Genetic algorithms for NN architecture optimization, NAB metric, F1-metric drawbacks analysis, techniques to improve AD quality.
result Improved anomaly detection quality through genetic algorithms and techniques.
Two different formulas for macro F1 lead to significant differences in classification evaluation.
problem Evaluation discrepancies in binary, multi-class, and multi-label classification problems.
method Comparison of two formulas for macro F1 metric.
result The two formulas can result in up to a 0.5 difference and different classifier rankings.
SentiCite analyzes citations for sentiment and nature, improving on existing methods.
problem Identifying quality scientific work amidst many citations.
method Sentiment analysis of citations with motivation detection.
result SentiCite outperforms state-of-the-art methods with a F1-measure of 0.71.
Improved deep learning for cardiac disease detection from limited ECG data.
problem Generalizing deep learning models to unseen classes and variations in cardiac disease detection.
method Utilizing unsupervised learning to construct latent spaces for better generalization.
result Significant improvements in F1-scores compared to state-of-the-art deep learning solutions.
AWARE-FX uses AI to audit foreign-exchange risk disclosures in corporate reports.
problem Weakly structured foreign-exchange risk disclosures in corporate reports.
method Combines lexicon, logic, encoders, and aggregation methods to convert text into traceable measures.
result FinBERT outperforms in most comparisons, improving F1 scores by up to 0.077.
Improved ECG classification using multi-task learning.
problem Low frequency of rare diagnoses in ECGs.
method Developed a multi-task CNN to classify multiple diagnoses from 12-lead ECGs.
result Adding common classes improves performance on rarer classes.
Paper fine-tunes a language model to predict long-term stock buy signals.
problem Predicting long-term stock price movements with narrative text.
method Fine-tuning a small language model on 10-K reports for buy/sell decisions.
result Buy signals generated from 10-K text are most precise at 6 and 9 months, providing 4.8-9% improvement over random selection.
In this paper, we study local solutions F=(F1,..,Fn) of a general functional equation of the form F1(U1(x,y))+....+Fn(Un(x,y))=0. A such equation will be called an ``abelian functional equation'' (Afe). We will restrict ourselves to the case when the inner functions Ui's are real rational functions. First we prove that…
Two randomized algorithms improve hypergraph learning accuracy and efficiency.
problem Efficiently learning and tagging images in hypergraphs.
method Block randomized SVD and conjugate gradient method.
result Both methods achieve high accuracy and reduce computational requirements.
Data-driven approaches outperform random choices in multi-label classification with Naive Bayes.
problem Improving multi-label classification performance with Naive Bayes classifiers.
method Comparison of data-driven, a priori, and random approaches on 12 benchmark datasets.
result Data-driven methods significantly outperform random baselines on F1 scores and Subset Accuracy.
OT domain adaptation improves aphasia detection across languages.
problem Detecting aphasia in low-resource languages with limited data.
method Utilized OT domain adaptation to map linguistic features across multiple languages.
result OT domain adaptation significantly improved F1 scores for French and Mandarin aphasia detection.
Improved performance in shape identification tasks using elastic metrics in t-SNE and UMAP.
problem Improper metrics in dimensionality reduction techniques lead to poor performance in machine learning applications.
method Incorporating elastic metrics into t-SNE and UMAP for functional data.
result Improved F1 scores on shape identification tasks for three benchmark datasets.
Deep learning models outperform traditional methods in automated chief complaint classification for syndromic surveillance.
problem Improving accuracy and speed of automated classification of emergency department records for outbreak detection.
method Implemented two LSTM and GRU models compared to MNB and SVM classifiers trained on 3.6 million de-identified records.
result RNN models outperform bag-of-words classifiers, especially for chief complaints.
Proposes a new undersampling method for imbalanced data classification.
problem Challenges of oversampling and undersampling in imbalanced data.
method Bilevel optimization framework for identifying optimal subset of majority training data.
result Improves F1 scores by up to 10% compared to state-of-the-art methods.
Improved performance in classifying domestic activities.
problem Classifying domestic activities effectively.
method Ensemble learning system based on CNN and LSTM.
result Significant improvement in F1-score (92.19% vs baseline 84.49%).
Paper studies Landsberg curvature of a specific Finsler metric.
problem Analyzing Landsberg curvature of a particular Finsler metric.
method Derived Landsberg curvature and mean Landsberg curvature of the twisted product Finsler metric.
result Necessary and sufficient conditions for the metric to be Landsberg or weakly Landsberg are established.
System optimizes retrieval for personal assistants using reinforcement learning.
problem Needing a neural retrieval-based Q&A system for user memory.
method Direct optimization of F1-score using reinforcement learning.
result Improved retrieval performance on test sets.
Given two maps f1 and f2 from the sphere Sm to an n-manifold N, when are they loose, i.e. when can they be deformed away from one another? We study the geometry of their (generic) coincidence locus and its Nielsen decomposition. On the one hand the resulting bordism class of coincidence data and the corresponding Niels…
Paper proposes new features to improve Bitcoin address classification.
problem Identifying criminal activities in Bitcoin transactions.
method New features based on transaction history summarized by block height.
result Improved performance of Bitcoin address classification significantly.
Modeling human language learning with multi-checkpoint machine translation.
problem Improving machine translation quality for language education.
method Ensemble of multi-checkpoints from a single model, sampling n-best sequences.
result Achieved 37.57 macro F1 score, outperforming baseline.
New scoring rules improve probabilistic classification model evaluation.
problem Traditional scoring rules misalign with the preference for correct classifications.
method Introduces Penalized Brier Score (PBS) and Penalized Logarithmic Loss (PLL) to modify proper scoring rules.
result PBS and PLL better identify optimal checkpoints and early stopping points, leading to superior F1 scores.
Paper semantifies bioassay text using neural networks.
problem Semantifying unstructured bioassay text descriptions.
method Neural-network-based approach to automatically semantify bioassay text.
result Neural-based semantification significantly outperforms a frequency-based baseline (72% F1 vs 47% F1).
This paper explores using KDE for balanced sampling in imbalanced datasets.
problem Imbalanced class distribution in data science.
method Kernel density estimation (KDE) for resampling the minority class.
result KDE-based resampling outperforms other techniques in F1-score and G-mean.
The paper benchmarks data stream classifiers for human activity recognition on connected devices.
problem Challenges in human activity recognition on connected devices, particularly high memory consumption and low F1 scores.
method Evaluation of five stream classification algorithms on real and synthetic datasets, measuring both performance and resource consumption.
result HT and MF classifiers show superior performance and resilience to concept drift compared to other algorithms.
PCA-Triage optimizes sensor data sampling for IoT networks.
problem Excessive sensor data in IoT networks exceeds available bandwidth.
method PCA-Triage uses streaming incremental PCA loadings to adaptively triage sensor data.
result PCA-Triage achieves high inference performance with minimal bandwidth usage.
Deep learning models accurately recognize and estimate physical activity types and energy expenditure from wrist accelerometer data.
problem Rigorous evaluation of wrist-worn accelerometers for assessing physical activity across the lifespan.
method Built deep learning networks to extract spatial and temporal representations from time-series data, recognizing physical activity types and estimating energy expenditure.
result Deep learning models achieved high performance: F1 scores of 0.82, 0.81, and 95 for sedentary, locomotor, and lifestyle activities, respectively; root mean square error of 1.1 for EE estimation.
DeepBeat uses deep learning to assess signal quality and detect arrhythmia in wearable devices.
problem Detecting atrial fibrillation from wearable devices with noise.
method Multi-task deep learning approach using convolutional denoising autoencoders.
result Significantly improved AF detection accuracy compared to traditional methods.
LoRAS improves model performance on imbalanced datasets by better oversampling the minority class.
problem Imbalanced datasets lead to poor model performance, especially for the majority class.
method Localized Random Affine Shadowsampling (LoRAS) to oversample minority class data.
result LoRAS generates better ML models in terms of F1-Score and Balanced accuracy compared to SMOTE and its extensions.
Proposes BA method for unbiased time series anomaly detection evaluation.
problem Anomalies in time series data are rare, making F1-score unreliable.
method Introduces Balanced Point Adjustment (BA) to address F1-score bias.
result BA provides fairer evaluation of time series anomaly detectors.
Paper classifies Brazilian music genres using song lyrics with BLSTM network.
problem Classifying Brazilian music genres from lyrics.
method Used BLSTM network combined with SVM, Random Forest, and word embeddings.
result BLSTM outperforms other models with an F1-score of 0.48.
GWN improves multimodal data fusion accuracy for chronic pain patients.
problem Dynamic and unspecified uncertainties in multimodal data fusion.
method Inspired by Global Workspace Theory, GWN is a neural network architecture that dynamically attends to multiple modalities.
result GWN achieved higher F1 scores (0.92 and 0.75) for multimodal discrimination and classification tasks.
Study uses multimodal machine learning to predict ICD-10 codes.
problem Improving accuracy and interpretability of ICD-10 code predictions.
method Developed separate models for text and tabular data, integrated using an ensemble method.
result Best-performing model achieved micro-F1 of 0.7633 and micro-AUC of 0.9541.
Gestalt combines two models to improve SQuAD2.0 performance.
problem Improving the accuracy of answering questions in context paragraphs.
method A stacking ensemble of ALBERT and RoBERTa models, combined with a CNN-based meta-model.
result Best ensemble achieved 87.117 EM and 90.306 F1 scores, improving baseline by 0.55% and 0.61% respectively.
Deep ROC analysis improves model selection and interpretation in medical and AI applications.
problem Inadequate performance measures for binary classifiers.
method Deep ROC analysis, translating AUC and partial AUC into balanced average accuracy and post-test measures.
result Deep ROC analysis provides balanced average accuracy, average sensitivity, and average specificity.
New model detects crying in real-world settings with improved accuracy.
problem Generalization of cry detection models to real-world environments.
method Evaluated machine learning approaches on a novel dataset of real-world infant crying.
result Improved F1 score of 0.613 for crying event recognition in real-world settings.
Model creates summaries of patient notes to save time and reduce errors.
problem Improper summarization of patient notes leads to inefficiencies and errors.
method Developed an LSTM model to sequentially label topics in history of present illness notes.
result Achieved an F1 score of 0.876, indicating the model's effectiveness.
TrackNet tracks tiny balls in sports videos with high precision.
problem Accurately tracking tiny and fast-moving objects like tennis balls in sports videos.
method Developed a heatmap-based deep learning network, TrackNet, trained on public and additional labeled videos.
result Achieved high precision (99.7%) in ball tracking on public domain videos.
DISPR uses diffusion models to predict 3D cell shapes from 2D images.
problem Predicting 3D cell shapes from 2D microscopy images.
method Diffusion model trained to predict 3D shapes from 2D microscopy images as a prior.
result Adding DISPR predictions to minority cell classes improves classification accuracy.