Researchers predict NBA player salaries using machine learning, avoiding overfitting.
problem Predicting NBA player salaries based on performance statistics.
method Selected important determinants, used Random Forest machine learning, avoided overfitting.
result Very satisfactory salary predictions identified for important factors.
Develops objective metrics to evaluate NFL offensive linemen performance.
problem Objective evaluation of NFL offensive linemen performance is lacking.
method Uses statistical analysis of performance metrics to objectively evaluate offensive linemen.
result Identifies overvalued and undervalued offensive linemen.
Develops a framework to estimate NBA player salary ROI.
problem Measuring the relative return of player salaries in NBA.
method Five-part framework: GCP measure, SGV calculation, cash flow series, ROI calculation.
result Illustrates framework with 2022-2023 NBA data, showing top and bottom performers.
New model forecasts pension fund contributors and retirees.
problem Forecasting and estimating future populations and financial flows.
method 3-dimensional Markov chain probabilistic model.
result Increased reliability and deeper demographic analysis.
Paper uses DRL to evaluate NHL players based on game context.
problem Limited ability to model player performance in game context.
method Applied Deep Reinforcement Learning to learn Q-function from NHL play-by-play data.
result Game Impact Metric (GIM) correlates with player success and future salary.
The excessive compensation packages of CEOs of U.S. corporations in recent years have brought to the foreground the issue of fairness in economics. The conventional wisdom is that the free market for labor, which determines the pay packages, cares only about efficiency and not fairness. We present an alternative theory…
The paper offers algorithms for managing freelancers and in-house workers in online labor markets.
problem Managing freelancers and in-house workers in online labor markets efficiently.
method Developed algorithms for team formation with outsourcing in an online setting.
result Efficient online algorithms for minimizing costs in hiring and outsourcing.
This study examines how VaR constraints impact corporate managers' decisions and firm value.
problem The impact of Value-at-Risk (VaR) constraints on corporate managers' decisions and firm value.
method The study uses the concavification technique and quantile formulation to derive explicit solutions for optimal effort, terminal firm value, and project choice.
result A VaR requirement generally improves downside protection and reduces bankruptcy probability, but can increase it when the VaR floor is high.
Detects potential local adversarial examples to prevent fraud in credit insurance decisions.
problem Manipulating variables to gain unfair advantages in credit decisions.
method Provides critical features to a human expert to detect and control potential fraud.
result Demonstrates a method to identify and mitigate local adversarial examples.
Fair market valuations ignore future worker profits in employee-owned firms.
problem Ignoring future worker profits in fair market valuations for employee-owned firms.
method Analyzing property rights and residual claimants in employee-owned firms.
result Fair market valuations are inappropriate for employee-owned firms.
New study finds better employee pay leads to better stock performance.
problem Understanding the relationship between employee remuneration and stock performance.
method Developed new asset pricing factors using firm financial characteristics.
result Companies with higher employee remuneration tend to have better stock performance.
fairadapt uses causal inference to mitigate algorithmic bias in data pre-processing.
problem Mitigating algorithmic bias in machine learning predictions.
method Causal graphical model and observed data to address counterfactual questions.
result The method can help eliminate discrimination and justify fair decisions.
A new framework learns to select samples for active learning.
problem Lack of labeled data for training active learning strategies.
method Learning to Sample (LTS) framework with a sampling and boosting model.
result Significantly outperforms baselines in limited label budgets.
New framework values football players based on in-game interactions.
problem Valuing football players based on in-game performance.
method Combining financial models and network theory using a passing matrix.
result Dynamic and individualized player valuation framework.
Study optimal asset allocation for DC plans with inflation and mortality risks.
problem Maximizing expected utility from terminal wealth in a pension plan with inflation and mortality risks.
method Closed-form solutions using a sufficient maximum principle approach for a problem with partial information.
result Closed-form solutions for asset allocation problem.
New method converts LVAs into linear projections for better understanding of complex models.
problem Limited interpretability of nonlinear machine learning models.
method Animated linear projections and radial tours.
result Improved understanding of variable importance in complex models.
Predicts STEM attrition using transcript data.
problem High rates of STEM students leaving fields in higher education.
method Machine learning on student transcript data.
result Attrition from STEM fields can be accurately predicted.
We analyze the income distribution of employees for 9 consecutive years (2001-2009) using a complete social security database for an economically important district of Romania. The database contains detailed information on more than half million taxpayers, including their monthly salaries from all employers where they …
FACE generates actionable counterfactuals that are feasible and coherent with data.
problem Counterfactual explanations can be unachievable and offensive.
method FACE proposes a new approach to generate counterfactuals that are coherent with the data and feasible.
result FACE generates counterfactuals that are coherent with the data and feasible.
Optimal buying and selling times for homes in fluctuating interest rates.
problem Maximizing profit from buying and selling homes in a market with variable interest rates.
method Nested optimal stopping problem solved using a nonnegative concave majorant approach.
result Investor's optimal buying and selling strategies derived for CIR interest rates.
A computational theory reduces agent evaluation errors and speeds up processes.
problem Efficient evaluation of mini agents at reduced cost.
method Developed a computational theory and a meta-learner to handle heterogeneous agents.
result Reduced evaluation errors by 24.1% to 99.0% across various scenarios.
Develops a method to evaluate OPE robustness to hyperparameters and policies.
problem Difficulty in selecting and tuning OPE estimators due to limited experimental evaluations.
method Introduces IEOE (Interpretable Evaluation for Offline Evaluation) to assess robustness.
result Demonstrates improved evaluation of OPE estimators' reliability.
This paper evaluates and validates cluster results using external and internal evaluation methods.
problem Evaluating and validating the quality of clustering results.
method External evaluation using Homogeneity, Correctness, and V-measure scores; internal evaluation using Silhouette Index and Sum of Square Errors.
result Validation of the number of clusters using dendrogram and statistical frequency distribution.
Optimizes crowdsourced preference-based subjective evaluation with online learning.
problem Large-scale evaluation of generative media using crowdsourcing due to combinatorial explosion.
method Automatic optimization of pair combination selections and evaluation volumes with online learning.
result Optimizes evaluation by reducing pair combinations and allocating optimal evaluation volumes.
Enhances BLEU for better SMT evaluation.
problem Improving BLEU for more human-like evaluation of machine translations.
method Adapts BLEU to consider synonyms, word order, and style variations in human translations.
result Improves SMT evaluation metrics and correlates with human translation quality.
Study evaluates human vs. machine review generation, finds human assessments correlate better with lexical overlaps.
problem Evaluating natural language generation models for online reviews is challenging and inconsistent.
method Compared human evaluators with various automated evaluation methods, including discriminative and word overlap metrics.
result Human evaluators do not correlate well with discriminative evaluators, but correlate better with lexical overlaps.
The paper critiques current time series classification evaluation methods.
problem Current performance evaluation methods in time series classification are criticized.
method No specific new method proposed, but a critical analysis of existing methods.
result Suggests a need for discussion and reflection on TSC performance evaluation.
Proposes Nash averaging to improve evaluation in machine learning.
problem Overwhelming choices in evaluation suites and attacks have diluted the basic model.
method Detailed analysis of evaluation scenarios leads to Nash averaging, which adapts to data redundancies.
result Nash averaging encourages maximally inclusive evaluation, reducing bias.
Study finds AUC is most consistent across different prevalence in binary classification.
problem Consistency of model evaluation metrics across varying prevalence in binary classification.
method Analysis of 156 data scenarios with 18 metrics, 5 models, and a random guess model.
result AUC has the smallest variance in evaluating individual models and ranking of models.
This paper identifies and discusses bias in offline recommendation algorithm evaluations.
problem Bias in offline evaluation of recommendation algorithms.
method Describes and discusses the relevance of weighted offline evaluation.
result The need for weighted offline evaluation to reduce bias.
Unified evaluation for both quality and diversity in NLP.
problem Measuring both quality and diversity in NLP models.
method Proposes HUSE, a metric combining human and statistical evaluation.
result HUSE detects both quality and diversity defects in NLP models.
Cramming method evaluates learned policies from contextual bandits efficiently.
problem Evaluating final learned policies from contextual bandit algorithms.
method On-policy evaluation using a single pass of data, ensuring consistency and asymptotic normality.
result Cramming method reduces evaluation standard error by approximately 40% compared to off-policy methods.
Cer-Eval saves LLM evaluation costs while maintaining accuracy.
problem Challenges in evaluating large language models due to large dataset requirements.
method Adapts to different evaluation objectives, uses test sample complexity, and develops a partition-based algorithm.
result Cer-Eval can save 20-40% test points with comparable accuracy and 95% confidence guarantee.
The study addresses biases in evaluating molecular optimization methods and proposes methods to reduce these biases.
problem Biases in in silico evaluation of molecular optimization methods.
method Discussion and empirical investigation of bias reduction methods for predictor misspecification and sample reuse.
result Empirical investigation of bias reduction methods for predictor misspecification and sample reuse.
Paper addresses off-policy evaluation and learning with covariate shift.
problem Evaluating and training a new policy using historical data with a covariate shift.
method Derives efficiency bounds and proposes doubly robust estimators for OPE and OPL under covariate shift.
result Proposes estimators for off-policy evaluation and learning under covariate shift.
Proposes clustering as a new evaluation method for clinical knowledge embedding.
problem Traditional Link Prediction evaluation protocol loses information and harms model accuracy.
method Proposes Clustering Evaluation Protocol as an alternative.
result Experimental results show the proposed protocol can potentially replace Link Prediction.
Study off-policy evaluation in partially observable environments, reducing bias and errors.
problem Bias and large errors in off-policy evaluation for partially observable environments.
method Defined and solved off-policy evaluation for POMDPs, introduced Decoupled POMDP model.
result Demonstrated and compared off-policy evaluation methods, showing benefits of new approach.
AXE evaluates explanations to avoid misleading Rashomon set model selection.
problem Evaluating explanations for Rashomon set models to avoid false selection.
method Proposed AXE method to evaluate explanation quality.
result AXE detects adversarial fairwashing with 100% success rate.
Reduces bias in recommendation system evaluations.
problem Offline evaluation bias in collaborative filtering algorithms.
method Weighted offline evaluation method.
result Reduces bias in recommendation system evaluations.
We consider evaluation methods for payoffs with an inherent financial risk as encountered for instance for portfolios held by pension funds and insurance companies. Pricing such payoffs in a way consistent to market prices typically involves combining actuarial techniques with methods from mathematical finance. We prop…
The paper proposes a method to evaluate ML models for subjective inference, focusing on sentence toxicity.
problem Bias in ML models for subjective inference, especially in real-life applications.
method Proposes a list of specifications to evaluate ML models for subjective inference, illustrated with a sentence toxicity example.
result Demonstrates the importance of considering subjectivity and bias in evaluating ML models.
New method for evaluating LLMs reduces bias in open-ended evaluations.
problem Bias in Elo-based ratings of LLMs due to data redundancies.
method Proposes evaluation as a 3-player game and introduces novel solution concepts.
result Novel method leads to more robust and intuitive ratings.
This paper considers the challenge of evaluating a set of classifiers, as done in shared task evaluations like the KDD Cup or NIST TREC, without expert labels. While expert labels provide the traditional cornerstone for evaluating statistical learners, limited or expensive access to experts represents a practical bottl…
New method improves consistency of reinforcement learning performance evaluations.
problem Inconsistent performance results in reinforcement learning due to flawed evaluation metrics.
method Proposes a new comprehensive evaluation methodology for reinforcement learning algorithms.
result Demonstrates improved reliability of performance measurements for reinforcement learning algorithms.
SVRG reduces gradient evaluations for policy evaluation in reinforcement learning.
problem Policy evaluation in reinforcement learning with high computational costs.
method Two variants of SVRG for policy evaluation that reduce gradient calculations.
result Significant reduction in the number of gradient evaluations while preserving linear convergence speed.
Bayesian approach quantifies uncertainty in LLM evaluations.
problem Statistical uncertainty in evaluating LLM behavior.
method Bayesian evaluation of LLM behavior using probabilistic text generation strategies.
result Bayesian approach provides useful uncertainty quantification about LLM behavior.
Solves challenges in replicating ML/DL model evaluations.
problem Challenges in evaluating and studying ML/DL innovations.
method Proposes MLModelScope for repeatable model evaluation.
result Facilitates rapid adoption of ML/DL innovations.
Paper discusses methods to evaluate defenses against adversarial examples.
problem Difficulty in evaluating adversarial robustness.
method Methodological foundations and best practices for evaluating defenses.
result Suggests new methods to avoid common pitfalls in evaluations.