TimeInf estimates data contribution in time series data, improving model performance and anomaly detection.
problem Estimating data contribution in time series datasets with temporal dependencies.
method Model-agnostic data contribution estimation method using influence scores.
result TimeInf effectively detects time series anomalies and outperforms existing methods.
The paper introduces Absolute Shapley Value to handle negative contributions in machine learning model training.
problem Negative marginal contributions in machine learning model training.
method Investigates three philosophies: Original Shapley Value, Zero Shapley Value, and Absolute Shapley Value.
result Absolute Shapley Value significantly outperforms other definitions in evaluating data importance.
Anomaly detection using dimensionality reduction has been an essential technique for monitoring multidimensional data. Although deep learning-based methods have been well studied for their remarkable detection performance, their interpretability is still a problem. In this paper, we propose a novel algorithm for estima…
FedCM measures contributions in real-time for federated learning.
problem Fairly allocating contributions in federated learning systems.
method FedCM calculates impact based on current and previous rounds with attention aggregation.
result FedCM is more sensitive to data quality and quantity in real-time.
NDDV estimates data point value from a single stochastic trajectory.
problem Estimating marginal contributions of data points over stochastic training paths.
method Introduces Neural Dynamic Data Valuation (NDDV) using stochastic state and adjoint equations.
result NDDV provides a one-run, trajectory-conditioned estimator of data point value.
In this paper, we introduce Anomaly Contribution Explainer or ACE, a tool to explain security anomaly detection models in terms of the model features through a regression framework, and its variant, ACE-KL, which highlights the important anomaly contributors. ACE and ACE-KL provide insights in diagnosing which attribut…
The paper studies quantile contributions and their relationship with order statistics in heavy-tailed distributions.
problem Challenges of classical statistical models in heavy-tailed distributions.
method Theoretical study of quantile contribution statistic and its relationship with order statistics. Derivation of closed-form expression for joint CDF of order statistics and quantile contributions.
result Established asymptotic normality of quantile contributions and characterized their limiting distribution.
Federated Machine Learning (FML) creates an ecosystem for multiple parties to collaborate on building models while protecting data privacy for the participants. A measure of the contribution for each party in FML enables fair credits allocation. In this paper we develop simple but powerful techniques to fairly calculat…
A study on optimizing data augmentation weights for improved test-time predictions.
problem Improving robustness of predictions during testing with data augmentation methods.
method A weighted Test-Time Augmentation (TTA) approach based on variational Bayesian framework to optimize weights.
result Optimizing weights suppresses unwanted data augmentations and improves prediction performance.
In Random Forests, proximity distances are a metric representation of data into decision space. By observing how changes in input map to the movement of instances in this space we are able to determine the independent contribution of each feature to the decision-making process. For binary feature vectors, this process …
A neural network and evolutionary algorithm framework designs nonlinear optical molecules.
problem Designing efficient nonlinear optical materials.
method Multi-stage Bayesian neural network (msBNN) and corrected Lewis-mode group contribution method (cLGC) combined with evolutionary algorithm (EA).
result Accurately and efficiently designs molecules with different optical properties using a small data set.
The paper proposes a new method for product recommendation that considers revenue contributions and user similarity.
problem High dimensionality and sparsity in user-item data, especially in terms of revenue contributions.
method The approach encodes revenue contributions in the user-item matrix and computes customer similarity using suitable distance measures.
result The method segments users based on revenue-based similarity and supports recommendations aligned with profitability objectives.
Data science relies on pipelines that are organized in the form of interdependent computational steps. Each step consists of various candidate algorithms that maybe used for performing a particular function. Each algorithm consists of several hyperparameters. Algorithms and hyperparameters must be optimized as a whole …
Determining risk contributions of unit exposures to portfolio-wide economic capital is an important task in financial risk management. Computing risk contributions involves difficulties caused by rare-event simulations. In this study, we address the problem of estimating risk contributions when the total risk is measur…
Automated anomaly detection is essential for managing information and communications technology (ICT) systems to maintain reliable services with minimum burden on operators. For detecting varying and continually emerging anomalies as differences from normal states, learning normal relationships inherent among cross-dom…
Non-experts have long made important contributions to machine learning (ML) by contributing training data, and recent work has shown that non-experts can also help with feature engineering by suggesting novel predictive features. However, non-experts have only contributed features to prediction tasks already posed by e…
Study on risk contributions of portfolios using lambda quantile risk measures.
problem No known allocation rule for non-positively homogeneous risk measures.
method Defined lambda quantiles on portfolio compositions, derived derivatives, and introduced generalized Euler contributions.
result Explicit formulae for the derivatives of lambda quantiles, showing their homogeneity properties.
New model shows most users on Q&A site Stack Overflow ignore badges.
problem Understanding user behavior in Q&A websites with badges.
method Probabilistic model applied to Stack Overflow data.
result Majority of users remain apathetic to badges but still contribute.
Unified method for local GBDT feature contributions.
problem Need for local model interpretation for GBDT.
method Unified computation mechanism for instance-level feature contributions.
result Unified computation mechanism for GBDT feature contributions.
New measure EC assesses node contributions in nonlinear, time-varying systems.
problem Existing node contribution measures assume linear, time-invariant dynamics, failing for complex, real-world systems.
method Defined 'emergent contribution (EC)' as a dynamical leverage measure from Jacobians of differentiable models.
result EC diverges from average controllability under persistent regime switching and sign reversal, identifying limits of local linearization.
New method for sparse data using L1-NMF with improved sparsity control.
problem Sparse data with false zeros and heavy-tailed noise.
method Component-wise L1-NMF with weighted penalization and coordinate descent.
result Effective in handling sparse data with false zeros.
FGSV defends against shell company attacks in group data valuation.
problem Shell company attacks on group-level data valuation.
method Developed a provably fast and accurate approximation algorithm for FGSV.
result Empirical results show significant improvement in computational efficiency and accuracy.
Study ablated data augmentation techniques and their mathematical equivalence to penalties.
problem Lack of mathematical understanding of differences between ablated data augmentation techniques.
method Formal model of mean ablated data augmentation and inverted dropout for linear regression; empirical validation for deep networks.
result Ablated data augmentation and inverted dropout are mathematically equivalent to penalties in optimization.
P3LS preserves privacy while integrating data across companies.
problem Privacy concerns in cross-organizational data exchange and integration.
method Privacy-preserving federated learning technique using SVD-based PLS and random masks.
result Improves prediction performance on process-related indicators.
We introduce a new factor model for log volatilities that performs dimensionality reduction and considers contributions globally through the market, and locally through cluster structure and their interactions. We do not assume a-priori the number of clusters in the data, instead using the Directed Bubble Hierarchical …
This paper introduces Sigma, a domain-specific computational representation for collaboration in large-scale for the field of economics. A computational representation is not a programming language or a software platform. A computational representation is a domain-specific representation system based on three specific …
Paper proposes incentives for federated learning to ensure truthful contributions.
problem Ensuring truthful contributions from decentralized users in federated learning.
method Introduces a scoring rule based framework to incentivize truthful reporting of local hypotheses at a Bayesian Nash Equilibrium.
result Proposed solution verified using MNIST and CIFAR-10 datasets, showing decreasing scores for low-quality hypotheses.
Improves privacy amplification by shuffling for differential privacy.
problem Enhancing privacy guarantees in systems with anonymous data contributions.
method Theoretical and numerical analysis of Rényi differential privacy parameters and privacy amplification by shuffling.
result First asymptotically optimal analysis of Rényi differential privacy parameters for shuffled outputs.
This paper presents a new fuzzy k-means algorithm for the clustering of high-dimensional data in various subspaces. Since high-dimensional data, some features might be irrelevant and relevant but may have different significance in the clustering process. For better clustering, it is crucial to incorporate the contribut…
Federated learning is a recent advance in privacy protection. In this context, a trusted curator aggregates parameters optimized in decentralized fashion by multiple clients. The resulting model is then distributed back to all clients, ultimately converging to a joint representative model without explicitly having to s…
This paper learns multi-modal embeddings from text, audio, and video views/modes of data in order to improve upon down-stream sentiment classification. The experimental framework also allows investigation of the relative contributions of the individual views in the final multi-modal embedding. Individual features deriv…
In this paper we consider a multivariate model-based approach to measure the dynamic evolution of tail risk interdependence among US banks, financial services and insurance sectors. To deeply investigate the risk contribution of insurers we consider separately life and non-life companies. To achieve this goal we apply …
Gaussian processes model sparse data in astrophysics and chemistry.
problem Scarcity of data in high-energy astrophysics and synthetic chemistry.
method Gaussian processes for uncertainty-aware predictions and inferences.
result GPs enable predictions and model latent emission from black holes and molecules.
Generative model identifies temporal count data components with regime-dependent contributions.
problem Modeling temporal count data with regime-dependent dynamics.
method Generative framework combining regime-adaptive dynamics with Poisson log-normal emissions.
result Established identifiability of the model and revealed co-variation patterns and regime shifts.
The paper establishes a connection between different risk measures and their risk contributions.
problem Understanding the relationship between conditional coherent and deviation risk measures.
method Axiomatic framework and continuous-time risk contribution analysis.
result Risk contributions of time-consistent risk measures are also time-consistent.
SHAP clustering explains model predictions by grouping similar feature contributions.
problem Lack of explainability in black-box models.
method SHAP values for feature contributions, supervised clustering of SHAP values.
result Insight into pathways leading to similar predictions.
Paper introduces contribution measures for systemic risk in crypto markets.
problem Evaluating systemic risk and quantifying risk interactions in cryptocurrency markets.
method Develops various contribution ratio measures based on MCoVaR, MCoES, and MMME.
result Establishes sufficient conditions for comparing contribution measures between sets of random vectors.
In-Run Data Shapley offers efficient data attribution for large-scale models.
problem Existing data attribution methods are computationally intensive and cannot target specific models.
method In-Run Data Shapley, which efficiently attributes data contributions to a specific model without re-training.
result In-Run Data Shapley achieves significant efficiency, enabling data attribution for pretraining models.
Deep learning improves survival analysis for complex data types.
problem Limited application of DL in survival analysis for complex data.
method Comprehensive review of DL methods for time-to-event analysis.
result Methods often ignore complex settings like multiple risks and censoring.
The paper proposes a method to detect and filter noisy or mislabeled data using pointwise mutual information.
problem Detecting and filtering noisy or mislabeled data in deep learning models.
method A mutual information-based framework quantifying statistical dependencies between inputs and labels.
result The method effectively filters low-quality samples, improving classification accuracy by up to 15%.
Efficiently estimates SAGE values using causal structure learning.
problem Computational infeasibility of exact SAGE calculations.
method Uses causal structure learning to identify conditional independencies and accelerate SAGE approximation.
result Empirically demonstrates efficient and accurate estimation of SAGE values.
Study finds risk management significantly improves pension scheme efficiency in Kenya.
problem Improving efficiency of pension schemes in Kenya.
method Panel data analysis of 128 pension schemes from 2015-2021.
result Risk management significantly mediates the relationship between corporate governance and pension scheme efficiency.
Recommending items to users is a challenging task due to the large amount of missing information. In many cases, the data solely consist of ratings or tags voluntarily contributed by each user on a very limited subset of the available items, so that most of the data of potential interest is actually missing. Current ap…
Algorithm generates private continuous-time data for sensitive domains.
problem Private generation of continuous-time data for sensitive domains.
method Mean-field Langevin dynamics and noisy particle gradient descent.
result Strong privacy guarantees for one-time data contributions.
Paper outlines a new mathematical language for experiments.
problem Formalizing the scientific process for automation.
method Formulates the scientific process in precise mathematical language.
result Novel contributions in data processing, bias variance, and deficiency.
This paper speeds up mean curvature computation for high-dimensional data.
problem Efficiently computing mean curvature in high-dimensional datasets.
method Two contributions: algebraic identity and truncated SVD approximation.
result Mean curvature computation reduced from O(m4) to O(k2m+kmp2). The paper proposes a dynamic risk measure approach for evaluating defined-contribution pension funds.
problem Periodic evaluation of defined-contribution pension funds to manage risk and improve projections.
method Dynamic risk measure criterion, model-free reinforcement learning, Lee-Carter mortality model.
result Periodic evaluations lead to more risk-averse strategies, while mortality improvements encourage risk-seeking behaviors.
This paper extends risk parity to continuous-time, solving risk budgeting problems.
problem Achieving robust risk across different assets in continuous-time.
method Characterizing risk contributions and solving risk budgeting problems using continuous-time terminal variance.
result Risk contributions and risk budgets can be represented as predictable processes in continuous-time.