NFTs with diverse rare attributes sell at higher prices.
problem Understanding how rarity affects NFT market dynamics.
method Analyzed 3.7M NFT transactions across 410 collections.
result Rarer NFTs sell for higher prices and are less risky.
Study evaluates financial misstatement detection methods, highlighting evaluation process impact.
problem Detecting financial reports with high misstatement risk.
method Proposes a new, realistic evaluation framework focusing on misstatement rarity, time dimension, and detection latency.
result Evaluation process significantly impacts system performance, revealing model and feature type effectiveness.
Paper proposes a risk index combining frequency and severity of abnormal driving patterns.
problem Assessing driver risk based on telematics data.
method Combines frequency of abnormal driving patterns with severity quantified through tail rarity.
result Developed a risk index that enables reliable discrimination and ranking of drivers.
CutMix enhances feature learning in neural networks, improving test accuracy.
problem Understanding and improving feature learning in neural networks using patch-level augmentation.
method Three distinct methods: vanilla training, Cutout training, and CutMix training were studied.
result CutMix training yields the highest test accuracy and learns all features and noise vectors evenly.
Method predicts rarity of image features to support research integrity investigations.
problem Difficulty in determining if image reuse is by chance or intentional.
method Statistical estimation of ORB features' chance occurrence across PubMed Open Access Subset dataset.
result The method produces decreasingly smaller p-values for more complex imagery, supporting null hypothesis.
LSVM with EBT reduces quasar detection errors by 10x.
problem Rare quasar detection in astronomy with high cost of misclassification.
method Linear Support Vector Machine (LSVM) with Ensemble Bagged Trees (EBT) and Learning from Mistakes.
result 10x reduction in False Negative Rate for quasar detection.
This paper explores constantly curved holomorphic 2-spheres in complex Grassmannian and confirms their rarity.
problem Classifying constantly curved holomorphic 2-spheres of degree 6 in the complex Grassmannian G(2,5). method Invoking the moduli space structure of sextic curves in Fano 3-folds and using PSL2-transvectant and engaged unitary analyses. result The moduli space of constantly curved sextic curves in G(2,5) is semialgebraic of dimension 2, with only one nonhomogeneous member. DBDT uses deep boosting decision trees for fraud detection.
problem Fraud detection in imbalanced data.
method Gradient boosting with neural networks (SDT), AUC maximization.
result DBDT significantly improves fraud detection performance.
Deep learning predicts NFT prices with high accuracy.
problem Dynamic valuation of non-fungible tokens (NFTs).
method Trained deep learning model on Ethereum blockchain data.
result Highly accurate price predictions of NFTs.
OMASGAN generates anomalous samples on distribution boundary to improve anomaly detection.
problem Missed anomalies and low AD performance due to OoD samples in generative models.
method OMASGAN generates anomalous samples on the estimated distribution boundary using a GAN approach, refining AD models.
result OMASGAN improves AD performance by at least 0.24 and 0.07 points on average on MNIST and CIFAR-10 datasets.
New framework estimates treatment effects in extreme data.
problem Hindered by unavailability of counterfactual outcomes and rarity of extreme data.
method Proposes a new framework based on extreme value theory.
result Quantifies treatment effects using tail decay rates of potential outcomes.
AI boosts study of rare weather extremes with lower costs.
problem Difficulty in studying rare weather events due to limited data and models.
method Coupling AI forecasts with physics models using rare-event algorithms.
result Efficiently characterizes very rare events like once-per-millennium heatwaves.
This paper evaluates data enrichment techniques for rare event detection in manufacturing.
problem Rare events in manufacturing lead to unplanned downtime and high energy consumption.
method Time series data augmentation, sampling, and imputation techniques combined with supervised machine learning.
result Data enrichment enhances rare failure event detection and prediction by up to 48%.
Improved forecasting of suicide attempts using LSGPs for patients with little data.
problem Challenges in predicting suicide attempts due to their rarity and patient heterogeneity.
method Introduced Latent Similarity Gaussian Processes (LSGPs) to capture patient heterogeneity.
result LSGPs outperform baseline models, even without kernel-design, and offer new insights into patient similarity.
We introduce the notion of connection thickness of spheres in a Cayley graph, related to dead-ends and their retreat depth. It was well-known that connection thickness is bounded for finitely presented one-ended groups. We compute that for natural generating sets of lamplighter groups on a line or on a tree, connection…
Regions of high-dimensional input spaces that are underrepresented in training datasets reduce machine-learnt classifier performance, and may lead to corner cases and unwanted bias for classifiers used in decision making systems. When these regions belong to otherwise well-represented classes, their presence and negati…
Human falls rarely occur; however, detecting falls is very important from the health and safety perspective. Due to the rarity of falls, it is difficult to employ supervised classification techniques to detect them. Moreover, in these highly skewed situations it is also difficult to extract domain specific features to …
We develop a new method to estimate failure probabilities in complex systems.
problem Estimating failure probabilities in safety-critical autonomous systems is challenging due to the rarity of failures and large state spaces.
method We propose an adaptive importance sampling algorithm that minimizes forward Kullback-Leibler divergence and uses Markov score ascent methods.
result Our method provides more accurate failure probability estimates than existing techniques.
Better methods to detect insider threats need new anticipatory analytics to capture risky behavior prior to losing data. In search of the best overall classifier, this work empirically scores 88 machine learning algorithms in 16 major families. We extract risk features from the large CERT dataset, which blends real net…
Develops RES metrics for stable rare-event forecasting evaluation.
problem Challenges in evaluating forecasts of rare events.
method Rare-event-stable (RES) metrics designed to maintain stable thresholds under extreme rarity.
result RES metrics maintain stable thresholds, consistent model rankings, and near-complete prevalence invariance.
Noisy labeled data is more a norm than a rarity for crowd sourced contents. It is effective to distill noise and infer correct labels through aggregation results from crowd workers. To ensure the time relevance and overcome slow responses of workers, online label aggregation is increasingly requested, calling for solut…
New index improves anomaly detection in correlated time series data.
problem Challenges in evaluating cluster quality for anomaly detection.
method Introduced Synchronized Anomaly Agreement Index (SAAI) to assess cluster quality.
result Maximizing SAAI improves anomaly detection accuracy by 0.23 compared to SSC and by 0.32 compared to X-Means.
This paper explores object detection in the small data regime, where only a limited number of annotated bounding boxes are available due to data rarity and annotation expense. This is a common challenge today with machine learning being applied to many new tasks where obtaining training data is more challenging, e.g. i…
New algorithm identifies best arm in rare event scenarios.
problem Identifying the best arm with tiny reward probability.
method Approximated Compound Poisson process for faster algorithms.
result Improved computational efficiency with minor sample complexity increase.
Reinforcement learning improves insurance claims reserving by learning from all claim trajectories.
problem Traditional reserving models learn only from settled claims, missing valuable data from ongoing claims.
method Formulated as a Markov decision process, uses reinforcement learning to update OCL estimates sequentially.
result Soft Actor-Critic implementation achieves competitive claim-level accuracy and strong aggregate performance.
The paper explains how importance sampling can be used for optimization of rare events.
problem Minimizing tail risks in stochastic optimization formulations.
method Importance sampling for reducing sample requirements in estimating rare events.
result Effective importance sampling techniques for optimization of rare events.
EX-DRL improves extreme quantile prediction for financial risk management.
problem Inaccurate estimation of extreme quantiles in loss distributions.
method EX-DRL uses Generalized Pareto Distribution (GPD) to model the tail of the loss distribution and Quantile Regression (QR) to improve extreme quantile prediction.
result EX-DRL provides more precise estimates of extreme quantiles, improving risk metrics reliability.
Method detects anomalies on attributed graphs with few labeled instances.
problem Detecting anomalies on connected instances (attributed graphs) with limited labeled data.
method Embed nodes in latent space using GCNs, training to distinguish normal and anomalous nodes.
result Method outperforms existing methods on real-world attributed graph datasets.
PRESTO improves rare event prediction by shrinking towards proportional odds model.
problem Difficult to predict rare events due to class imbalance.
method PRESTO relaxes proportional odds model by estimating separate weights for transitions between categories, imposing L1 penalty to shrink towards proportional odds.
result PRESTO consistently estimates decision boundary weights under sparsity assumption, improving rare probability estimation.
Synthetic augmentation improves financial machine learning performance in variance-dominant regimes.
problem Data scarcity in financial machine learning.
method Formalized synthetic augmentation, introduced size-matched null augmentation, and developed a non-parametric block permutation test.
result Synthetic augmentation is beneficial only in variance-dominant regimes, such as persistent volatility forecasting.
Noisy labeled data is more a norm than a rarity for self-generated content that is continuously published on the web and social media. Due to privacy concerns and governmental regulations, such a data stream can only be stored and used for learning purposes in a limited duration. To overcome the noise in this on-line s…
Prognostics or Remaining Useful Life (RUL) Estimation from multi-sensor time series data is useful to enable condition-based maintenance and ensure high operational availability of equipment. We propose a novel deep learning based approach for Prognostics with Uncertainty Quantification that is useful in scenarios wher…
Consider a two-class clustering problem where we observe Xi=ℓiμ+Zi, Zi∼iidN(0,Ip), 1≤i≤n. The feature vector μ∈Rp is unknown but is presumably sparse. The class labels ℓi∈{−1,1} are also unknown and the main interest is to estimate them. We are interested …
New model helps identify suspect footwear from crime scene prints.
problem Identifying a suspect's footwear from crime scene prints among thousands of similar shoes.
method Developed a hierarchical Bayesian model with spatially varying coefficients.
result Improved accuracy and reliability in forensic shoe print analysis.
Approach for modeling EHR data with rare features, improving prediction and interpretation.
problem Challenges in modeling rare binary features in EHR data.
method Tree-guided feature selection and logic aggregation for large-scale regression.
result Improved prediction and model interpretation of suicide risk in EHR data.
CovidCare uses EMR data to predict patient outcomes in emerging epidemics.
problem Intelligent prognosis for patients with emerging infectious diseases during rapid epidemics.
method Transfer learning and knowledge distillation from existing EMR data.
result CovidCare outperforms baseline methods in predicting patient length of stay.
Paper examines adversarial attacks on weather forecasting models, focusing on TC trajectory prediction.
problem Adversarial attacks can mislead downstream TC trajectory predictions in DLWF models.
method Proposes Cyc-Attack, a method using a surrogate model and skewness-aware loss function to generate adversarial TC paths.
result Cyc-Attack achieves higher true positive rates and lower false alarm rates compared to conventional methods.
We consider the problem of evaluating the quality of startup companies. This can be quite challenging due to the rarity of successful startup companies and the complexity of factors which impact such success. In this work we collect data on tens of thousands of startup companies, their performance, the backgrounds of t…
DFRot improves LLMs by reducing outlier and massive activation effects.
problem Reducing outlier and massive activation effects in rotated LLMs.
method Weighted loss function and orthogonal Procrustes transforms for rotation matrix refinement.
result DFRot achieves dual free (Outlier-Free and Massive Activation-Free) with significant improvements in perplexity.
Study flaws in generative model evaluation metrics, especially for diffusion models.
problem Flaws in existing metrics for evaluating generative models, particularly for diffusion models.
method Systematic study of generative models, human perception experiments, and analysis of feature extractors.
result State-of-the-art perceptual realism of diffusion models is not reflected in commonly reported metrics.
FL improves insurance claims loss prediction without sharing data.
problem Limited data volume and variety due to privacy concerns.
method Federated Learning (FL) to update a global model using local data insights.
result Improved claims loss forecasting compared to individual models.
FROB model improves robustness and reliable confidence for few-shot OoD detection.
problem Challenges in few-shot classification and OoD detection due to limited samples and adversarial attacks.
method FROB model combines support boundary generation and few-shot Outlier Exposure (OE) for improved robustness and reliable confidence.
result FROB achieves generalization to unseen OoD and maintains robustness independent of few-shot number.