This paper assesses GAN diversity using classification-based covariate shift.
problem Evaluating the diversity of GAN distributions is challenging.
method Developed tools to assess GAN diversity using a classification-based perspective.
result Popular GANs have significant problems reproducing training dataset properties.
A novel online feature selection method using DPP for diversity.
problem Online feature selection for diverse feature sets.
method DPP-based framework with three stages: sampling, local criteria, and global criteria.
result Demonstrated better compactness and comparable/outsuperior performance.
New framework shows diverse training data improves subgroup and overall performance.
problem Lack of understanding how diverse training data affects subgroup and overall performance.
method Casts data collection as part of the learning process, analyzes dataset compositions, and guides dataset design.
result Diverse representation in training data improves subgroup and overall performance.
Synthesizing high resolution photorealistic images has been a long-standing challenge in machine learning. In this paper we introduce new methods for the improved training of generative adversarial networks (GANs) for image synthesis. We construct a variant of GANs employing label conditioning that results in 128x128 r…
Study finds open data sets favor Western locales, impacting classifier performance.
problem Impact of biased open data sets on classifier performance in the developing world.
method Analysis of two large, publicly available image data sets and classifiers trained on them.
result Open data sets exhibit a bias towards Western locales, affecting classifier performance.
New analysis reveals diverse problem-solving behaviors in machine learning models.
problem Understanding and evaluating the diverse problem-solving behaviors of machine learning models.
method Spectral Relevance Analysis to characterize and validate machine learning models.
result Standard performance metrics fail to distinguish diverse problem-solving behaviors.
Model assesses credit risk using behavioral data from Experian and Bank of Italy.
problem Improving credit risk assessment in financial institutions.
method Statistical and machine learning techniques applied to behavioral data from Experian and Bank of Italy.
result Demonstrates transferability of the model from private to central data.
Deep learning improves mammography assessment with high accuracy.
problem Challenges in mammography assessment due to noise, resolution, and lack of ground truths.
method Proposes a classification approach using multi-scale deep tissue classifiers.
result Highest AUC of 0.9 achieved in classifying suspicious tissue patches.
InvestorBench benchmarks LLM agents in financial tasks.
problem Lack of a comprehensive benchmark for LLM-based financial agents.
method Developed a benchmark with diverse financial tasks and datasets.
result Evaluated LLM agents' performance across various financial products and market environments.
Study shows diverse data sources improve cryptocurrency forecasting models.
problem Improving cryptocurrency market forecasting accuracy.
method Integrating various data types, including on-chain metrics, traditional indices, and macroeconomic indicators.
result Data source diversity significantly enhances forecasting model performance.
Proposes a method to generate diverse translations by conditioning on target domain.
problem NMT models lack diversity in translations, even with search algorithms.
method Condition the decoder on a latent variable representing target domain, generated by a target encoder.
result Generated diverse translations without affecting performance or training time.
AlphaEval evaluates alpha mining models efficiently and comprehensively.
problem Lack of systematic evaluation for alpha mining models.
method Unified, parallelizable evaluation framework assessing predictive power, stability, robustness, financial logic, and diversity.
result AlphaEval achieves evaluation consistency comparable to comprehensive backtesting, providing more comprehensive insights and higher efficiency.
PerSense assesses personality traits from text for commonsense reasoning.
problem Estimating human personality traits from text for mental health analysis.
method Aggregated Probability Density Functions (PDF) and Machine Learning (ML) models.
result PerSense algorithms achieve comparable results to ground truth data, with high accuracy for personality assessment and commonsense prediction.
VTAB benchmarks diverse visual tasks to assess representation learning effectiveness.
problem Lack of a unified evaluation for general visual representations.
method Developed VTAB, a benchmark for diverse visual tasks, and evaluated many representation learning algorithms.
result VTAB revealed insights into the effectiveness of various representation learning methods.
HRF enhances tree diversity in random forests to improve performance.
problem Selection bias and lack of diversity in random forests.
method Introducing heterogeneity during tree construction by assigning lower weights to features used in previous trees.
result HRF outperforms other ensemble methods in accuracy across 52 datasets.
Meta-Dataset benchmarks few-shot learning models with diverse datasets.
problem Lack of diverse and realistic datasets for evaluating few-shot learning models.
method Meta-Dataset: a new benchmark with diverse datasets and realistic tasks.
result Meta-Dataset uncovers important research challenges in few-shot learning.
PQMass assesses generative model quality using chi-squared tests.
problem Assessing the quality of generative models without density assumptions.
method Divides sample space into regions, applies chi-squared tests to p-values.
result Effectively assesses generative model quality, novelty, and diversity.
Proposes a spectral method to assess and combine multiple data visualizations.
problem Evaluating and combining the strengths of different data visualization algorithms.
method Spectral method for assessing and combining multiple visualizations.
result Proposes a visualization eigenscore to quantify relative performance and a consensus visualization.
EGR refines and assesses protein complex structures.
problem Improving the accuracy of protein complex 3D structures for drug discovery.
method E(3)-equivariant graph neural network (GNN) for multi-task refinement and assessment.
result EGR achieves state-of-the-art performance in refining and assessing protein complexes.
Paper introduces risk assessment for contextual bandits without experiments.
problem Evaluate policies using logged data in context bandits.
method Lipschitz risk functionals and Off-Policy Risk Assessment (OPRA) framework.
result OPRA provides finite sample guarantees for various risk estimates.
The paper introduces a method to assess the reliability of model explanations.
problem Assessing the quality and reliability of model explanations.
method An Ordinal Consensus Approach using diverse bootstrapped surrogate explainers.
result Uncertainty estimates offer actionable insights beyond standard surrogate explainers.
New method assesses individual training points' privacy risk without retraining.
problem Privacy vulnerability of individual training points in membership inference attacks.
method Derives a closed-form decomposition of individual black-box MIA vulnerability, extending to deep networks.
result Proposes a surrogate score operating on last-layer representations that requires only a single trained model.
Unified model predicts stock and systemic risks from diverse financial data.
problem Isolating financial tasks leads to missed cross-scale dependencies.
method Shared Transformer backbone with modular task heads for cross-modal attention and multi-task optimization.
result Uni-FinLLM significantly outperforms baselines in stock forecasting, credit-risk assessment, and systemic-risk detection.
Two proteins are homologous if they have a common evolutionary origin, and the binary classification problem is to identify proteins in a candidate set that are homologous to a particular native protein. The feature (explanatory) variables available for classification are various measures of similarity of proteins. The…
Enhances Transformers for better risk assessment in finance.
problem Transformer models lack sensitivity to extreme financial losses.
method Integrates Loss-at-Risk function with Value at Risk (VaR) and Conditional Value at Risk (CVaR).
result Improves risk prediction and management in financial datasets.
A new method clusters heterogeneous subgroups for accurate causal learning.
problem Diverse causal relationships across different time spans, regions, or strategies.
method Nonlinear Causal Kernel Clustering
result Reduction in prediction error through enhanced causal learning.
Study finds optimal board gender diversity for emissions performance.
problem Association between board gender diversity and emissions performance.
method Panel regressions, machine learning, explainable AI.
result Optimal board gender diversity for emissions performance is approximately 35%.
Study assesses consistency and reproducibility of LLMs in finance and accounting tasks.
problem Consistency and reproducibility of LLM outputs in finance and accounting research.
method Extensive experimentation with 50 independent runs across 5 tasks using 3 OpenAI models.
result Task-specific patterns of consistency and reproducibility, with binary classification and sentiment analysis achieving near-perfect reproducibility.
CAT framework improves AI medical screening fairness and reliability.
problem Imbalanced data, varying performance across cohorts, and patient-level inconsistencies in traditional metrics.
method CAT framework introduces patient-level assessment, entropy-based distribution weighting, and cohort-weighted sensitivity and specificity.
result Enhanced predictive reliability, fairness, and interpretability of AI-driven medical screening models.
The Wisdom of Crowds is a phenomenon described in social science that suggests four criteria applicable to groups of people. It is claimed that, if these criteria are satisfied, then the aggregate decisions made by a group will often be better than those of its individual members. Inspired by this concept, we present a…
New global adversarial attacks improve DNN robustness assessment.
problem Vulnerability of deep neural networks to adversarial attacks.
method Proposed global adversarial example pairs and attack methods.
result DNNs hardened with local adversarial training are vulnerable to global attacks.
TinyXRA assesses financial risks from 10-K reports using a lightweight transformer model.
problem Comprehensive risk assessment from financial reports, distinguishing between upside and downside risk.
method Lightweight transformer model with dynamic attention, incorporating skewness, kurtosis, and Sortino ratio.
result State-of-the-art predictive accuracy and transparent risk assessments.
GANs mode collapse solved with Bures distance.
problem GANs mode collapse or mode dropping.
method Use Bures distance to match real and fake batch diversity in feature space.
result Diversity matching reduces mode collapse and improves sample quality.
DeepEvolution improves testing of deep neural networks by generating diverse test cases.
problem Lack of diversity in generated test cases from random fuzzing or transformations.
method Search-based approach using metaheuristics to ensure diversity in test cases.
result DeepEvolution significantly increases neuronal coverage and detects latent defects.
LSTM model predicts climate impacts on floods and droughts.
problem Predicting climate impacts on individual watersheds is challenging.
method Large-scale LSTM training on extensive data sets.
result LSTM model outperforms state-of-the-art models in predicting extreme flows.
Proposes a modified Morgan-Pitman test for evaluating variances in machine learning models.
problem Limited ability to account for sampling variability in model selection.
method Enhances the classic Morgan-Pitman test for robustness in non-linear models with heavy-tailed distributions or outliers.
result Demonstrates the test's effectiveness and practical utility in model evaluation and selection.
This paper uses LLMs to improve equity stock ratings by ingesting diverse financial and news data.
problem Challenges in traditional stock rating methods, including data overload, inconsistencies, and delayed reactions.
method Application of LLMs to generate multi-horizon stock ratings using various datasets.
result LLMs enhance the accuracy and consistency of stock ratings, outperforming traditional methods in forward returns.
A new model calibrates survival predictions for better risk assessment.
problem Calibrated time-to-event predictions are crucial but underexplored.
method Survival function estimator using neural network draws, without adversarial learning.
result The model outperforms existing approaches in calibration and distribution concentration.
Multivariate regular variation plays a role assessing tail risk in diverse applications such as finance, telecommunications, insurance and environmental science. The classical theory, being based on an asymptotic model, sometimes leads to inaccurate and useless estimates of probabilities of joint tail regions. This pro…
Extends Shifts dataset for MS lesion segmentation and marine vessel power estimation.
problem Distributional shift in training and deployment data for ML models.
method Develops new datasets for high-risk industrial applications.
result Demonstrates robustness and uncertainty estimation in new industrial tasks.
The paper introduces CoCoCat bonds for multi-region natural catastrophes, accounting for complex dependencies.
problem Valuation of multi-region contingent convertible bonds under complex dependencies.
method Developed a model accounting for inter-regional dependencies using change-of-measure techniques.
result Significant impact of inter-regional dependencies on CoCoCat bond pricing.
A new metric assesses generative models for molecules in drug design.
problem Difficulty in evaluating generative models for molecules.
method Fréchet ChemNet Distance (FCD) using a deep neural network trained to predict drug activities.
result FCD detects diversity and similarity of generated molecules to real ones.
Stein variational neural network ensembles improve diversity and uncertainty estimation.
problem Lack of proper Bayesian justification and diversity guarantees in deep neural network ensembles.
method Particle-based inference methods, specifically Stein variational gradient descent (SVGD), operating in weight space, function space, and hybrid settings.
result SVGD methods improve diversity and uncertainty estimation, approaching the true Bayesian posterior more closely.
New framework assesses LLM security risks in BFSI.
problem Lack of domain-specific security evaluation for LLMs in BFSI.
method Risk-aware evaluation framework combining taxonomy, automated red-teaming, and ensemble judging.
result Higher decoding stochasticity and adaptive interaction lead to more severe disclosures.
This review assesses statistical and machine learning methods for coral bleaching.
problem Coral bleaching due to rising sea temperatures and environmental factors.
method Statistical and machine learning models for predicting and analyzing coral bleaching.
result Statistical and machine learning methods are crucial for effective reef management.
Novel method decorrelates batches of triplets for active metric learning.
problem Correlation among triplets degrades active learning performance.
method Proposes a novel method to decorrelate batches of triplets, balancing informativeness and diversity.
result Method outperforms state-of-the-art in active metric learning.
Paper introduces a framework for managing cyber risk with insurance and cybersecurity models.
problem Pervasive challenges in managing cyber risk, especially for capital allocation.
method Combines insurance frequency-severity models with cybersecurity cascade models for comprehensive cyber risk assessment. Facilitates informed capital allocation through a two-pillar framework.
result Demonstrates the necessity of comprehensive cost-benefit analysis for budget-constrained companies.
Deep RL algorithms generally generalize better than specialized schemes.
problem Deep RL agents fail to generalize beyond their training environments.
method Presented a benchmark and experimental protocol to systematically assess generalization in deep RL.
result Vanilla deep RL algorithms outperform specialized generalization schemes.