ABOUT ML aims to improve transparency in ML lifecycle documentation.
problem Lack of standard documentation in machine learning lifecycle.
method Initiative to operationalize ML transparency and standardize documentation.
result Helps address gaps in ML lifecycle documentation.
Deep learning models can discriminate against certain groups, requiring computational methods to ensure fairness.
problem Algorithmic discrimination in deep learning models affecting protected groups.
method Interpretability and mitigation approaches at different stages of deep learning lifecycle.
result Interpretability aids in diagnosing and mitigating algorithmic discrimination in deep learning.
This paper introduces C-DSL to improve data mining outcomes by considering context.
problem Data collection ambiguities, data imbalance, hidden biases, lack of domain info, and data incompleteness.
method Developed Context-Driven Data Science Lifecycle (C-DSL) to address data quality issues.
result Tangible improvements to data mining outcomes were achieved through C-DSL.
Machine learning has evolved into an enabling technology for a wide range of highly successful applications. The potential for this success to continue and accelerate has placed machine learning (ML) at the top of research, economic and political agendas. Such unprecedented interest is fuelled by a vision of ML applica…
Industry lacks tools to secure ML systems, study finds.
problem Insufficient security tools for ML systems in industry.
method Interviews with 28 organizations to identify gaps.
result Need revised Security Development Lifecycle for ML.
A framework combining HSMM and survival analysis for lifecycle-oriented mobility analysis.
problem Understanding individual metro usage dynamics over multi-year horizons.
method A state-based lifecycle modeling framework integrating HSMM and discrete-time survival analysis.
result Identification of interpretable mobility states, transition dynamics, and state-dependent exit and re-entry processes.
The paper optimizes DIA purchase policies using lifecycle models and asset allocation.
problem Determining the optimal allocation to Deferred Income Annuities (DIAs).
method Employed a lifecycle model with utility of consumption and bequest, formalized optimization process, analyzed results, and extended model to include asset allocation.
result Optimal DIA allocation varies based on refundability, asset allocation, and perceived longevity.
Improving software quality through effective organizational learning.
problem Lack of reliable quantification methods for software evolution.
method Leveraging application lifecycle management data to identify and address managerial practices.
result Effective learning from past processes improves software quality indirectly.
A dynamic model of the product lifecycle of (nearly) homogeneous durables in polypoly markets is established. It describes the concurrent evolution of the unit sales and price of durable goods. The theory is based on the idea that the sales dynamics is determined by a meeting process of demanded with supplied product u…
Homeownership boosts wealth and welfare compared to renting, according to new research.
problem The conventional wisdom that renting is better than owning a home.
method Block-bootstrap lifecycle simulation to compare homeownership and renting strategies.
result Homeownership generates more wealth and welfare gains than renting, especially for households with high labor income.
RED-2400 is a public benchmark of trading events from a Solana exchange, labeled by algorithmic rejection.
problem Analyzing algorithmically-rejected trading events for insights into market dynamics.
method Public dataset of 6,660 algorithmically-rejected trading events, linked to post-rejection price and liquidity trajectories.
result First window of a planned series of datasets extending the time horizon and enabling regime-stratified analysis.
A new microeconomic model is presented that aims at a description of the long-term unit sales and price evolution of homogeneous non-durable goods in polypoly markets. It merges the product lifecycle approach with the price dispersion dynamics of homogeneous goods. The model predicts a minimum critical lifetime of non-…
Systematic review of ML models for detecting social media deception.
problem Detecting fake news, spam, and fake accounts on social media.
method 36 studies evaluated using PROBAST tool, identifying biases and limitations.
result Over-reliance on accuracy in imbalanced data settings is a flaw.
Focuses on monitoring and explaining models in real-world applications.
problem Ensuring high quality machine learning services in production environments.
method Statistical techniques for model performance and data monitoring, explanations of predictions.
result Challenges and solutions for implementing monitoring and explanation in production models.
A new indicator measures project risk from activity durations.
problem Managing project risks throughout the lifecycle.
method Activity Risk Index (ARI) based on Schedule Risk Baseline.
result Identifies activities contributing most to project uncertainty.
The process of exploring and exploiting Oil and Gas (O&G) generates a lot of data that can bring more efficiency to the industry. The opportunities for using data mining techniques in the "digital oil-field" remain largely unexplored or uncharted. With the high rate of data expansion, companies are scrambling to develo…
We extend the lifecycle model (LCM) of consumption over a random horizon (a.k.a. the Yaari model) to a world in which (i.) the force of mortality obeys a diffusion process as opposed to being deterministic, and (ii.) a consumer can adapt their consumption strategy to new information about their mortality rate (a.k.a. h…
SaML guides ML models to avoid survey biases.
problem ML models trained on survey data often ignore survey design metadata.
method Nine-step guideline for incorporating survey design metadata in ML lifecycle.
result SaML provides valid population inference from survey data.
The paper audits trading filters, finding a high save-to-miss ratio.
problem Improving the efficiency and accuracy of trading filters in decentralized exchanges.
method A precision audit of filter rules against real trading data, classifying rejection events.
result Conservative save-to-miss ratio of 3.7 : 1, with wider interpretation of 14.8 : 1.
Study identifies NFT whales driving the market with consistent high returns.
problem Lack of financial analysis of NFT trading ecosystem.
method Longitudinal study of 3.8M NFT transactions, classifying traders into whales, dolphins, and minnows.
result Top 0.1% of NFT traders (whales) drive the market with consistent, high returns.
Machine Learning is transitioning from an art and science into a technology available to every developer. In the near future, every application on every platform will incorporate trained models to encode data-based decisions that would be impossible for developers to author. This presents a significant engineering chal…
Three methods detect informed trading on prediction markets, each focusing on different aspects.
problem Detecting informed trading in decentralized prediction markets.
method Composite screen, event-level sign-randomization test, and Information Leakage Score (ILS) framework.
result Different methods detect informed trading on prediction markets, each focusing on different aspects.
Deployment of machine learning (ML) algorithms in production for extended periods of time has uncovered new challenges such as monitoring and management of real-time prediction quality of a model in the absence of labels. However, such tracking is imperative to prevent catastrophic business outcomes resulting from inco…
This study examines non-retail trading on Polymarket, revealing unique behavior patterns and structural limitations.
problem Lack of address-level quote-lifecycle data in Polymarket prediction markets.
method Empirical analysis of 13 million order-filled events using DBSCAN clustering on a six-feature fill-side vector.
result Non-retail behavior is uni-modal, contradicting previous archetypal hypotheses.
STAD adapts models to evolving time-based data shifts.
problem Gradual distribution shifts over time challenge existing test-time adaptation methods.
method Bayesian filtering method that learns time-varying dynamics in hidden features.
result STAD excels in handling small batch sizes and label shift on real-world data.
In a large E-commerce platform, all the participants compete for impressions under the allocation mechanism of the platform. Existing methods mainly focus on the short-term return based on the current observations instead of the long-term return. In this paper, we formally establish the lifecycle model for products, by…
The paper proposes a method to identify fair features in ML data integration.
problem Ensuring fairness in machine learning data integration.
method Causal interventional fairness, conditional independence tests, group testing.
result The proposed algorithm identifies fair features without biasing the dataset.
In life-cycle economics the Samuelson paradigm (Samuelson, 1969) states that the optimal investment is in constant proportions out of lifetime wealth composed of current savings and the present value of future income. It is well known that in the presence of credit constraints this paradigm no longer applies. Instead, …
This paper advances theory on the process of collaboration between entities and its implications on the quality of services, information, and/or products (SIPs) that the collaborating entities provide to each other. It investigates the scenario of outsourced IS projects (such as custom software development) where the e…
ReSkill reconciles RL skill creation with policy optimization.
problem RL policies lack reusable strategies across tasks.
method Integrates skill creation into RL loop with three mechanisms.
result Consistently outperforms existing methods, especially on unseen tasks.
FairPrep aims to improve fairness in machine learning by providing best practices.
problem Lack of best practices in fairness-enhancing interventions.
method Developer-centered design and evaluation framework for fairness-enhancing interventions.
result Hyperparameter tuning and data cleaning methods impact fairness outcomes.
PredictionMarketBench benchmarks trading agents on prediction markets.
problem Evaluating trading agents on prediction markets with realistic conditions.
method Deterministic replay of historical data, execution-realistic simulator, agent interface.
result Fee-aware algorithmic strategies outperform naive agents in volatile episodes.
The paper tackles model failure detection and refitting in real-world systems.
problem Real-world data often fails statistical models due to heterogeneity.
method Develops tools for detecting and identifying model failures and refitting to improve accuracy.
result Empirical and theoretical results show the effectiveness of the proposed methodology.
Isthmus platform simplifies ML/AI integration in healthcare.
problem Challenges in deploying ML models in healthcare due to data quality, regulatory, and security issues.
method Turnkey, cloud-based platform addressing data quality, clinical relevance, and regulatory compliance.
result Reduces time to market for operationalizing ML/AI in healthcare.
Simplifying machine learning (ML) application development, including distributed computation, programming interface, resource management, model selection, etc, has attracted intensive interests recently. These research efforts have significantly improved the efficiency and the degree of automation of developing ML mode…
We solve a lifecycle model in which the consumer's chronological age does not move in lockstep with calendar time. Instead, biological age increases at a stochastic non-linear rate in time like a broken clock that might occasionally move backwards. In other words, biological age could actually decline. Our paper is ins…
New attacks improve privacy audits by analyzing model updates.
problem Improving privacy audits of AI models through sequence analysis.
method Developed SeMI attacks to identify target insertions in model sequences.
result SeMI attacks achieve higher power and tighter privacy audits.
Blockchain helps secure payments between AI agents.
problem Ensuring secure payments between untrusted AI agents.
method Systematized four-stage lifecycle for A2A payments on blockchain.
result Challenges remain in weak intent binding, misuse, and limited accountability.
Study finds consumers are more price-sensitive before livestreams than after.
problem Understanding consumer demand during livestreaming lifecycle.
method Examined consumer demand for live events and recorded versions using data from a livestreaming platform.
result Demand is more price-sensitive before livestreams than after.
This work introduces a bias-variance decomposition for proper scores, improving uncertainty estimation in predictive models.
problem Reliable uncertainty estimation for predictions in safety-critical applications, especially under domain drift.
method Developed a general bias-variance decomposition for proper scores, introducing the Bregman Information as the variance term.
result The decomposition provides novel formulations for different predictive tasks, including classification and model ensembles.
Geometric stability predicts steerability and detects drift in language models.
problem Predicting steerability and detecting drift in language models.
method Supervised and unsupervised geometric stability measures.
result Supervised geometric stability predicts steerability with high accuracy and detects drift earlier.
This study categorizes RWA tokenization challenges and solutions.
problem Navigating the gap between on-chain deterministic code and off-chain probabilistic reality.
method Taxonomy and comparative analysis of RWA protocols, legal and technical standards.
result RWA tokenization requires overcoming legal and technical interoperability issues.
Foundation models alter medical data science workflow, challenging veridical data science principles.
problem Foundation models disrupt traditional data science practices in medicine.
method Critically examined the medical foundation model lifecycle and its deviation from veridical data science principles.
result Foundation models challenge veridical data science principles of predictability, computability, and stability.
Large language models struggle with causal relationships, leading to biases and hallucinations.
problem LLMs struggle with true causal relationships, leading to biases and hallucinations.
method Embed causality into LLMs training process at every stage.
result LLMs need to be trained to understand and apply causal knowledge, not just recite it.
Stablecoins offer efficient settlement but externalize costs and risks.
problem Comparing stablecoins to card networks in retail payments.
method Unified analytical framework (CLEAR) across five dimensions.
result Stablecoins are advantageous in closed-loop and high-friction contexts but structurally disadvantaged as open-loop instruments.
The Australian Government uses the means-test as a way of managing the pension budget. Changes in Age Pension policy impose difficulties in retirement modelling due to policy risk, but any major changes tend to be `grandfathered' meaning that current retirees are exempt from the new changes. In 2015, two important chan…
Unified framework explains retirement and annuitization decisions under age-dependent mortality.
problem Complexity of annuitization decisions due to longevity risk and labor force participation.
method Stochastic control and optimal stopping framework with habit formation and endogenous labor supply.
result Rich sequence of retirement dynamics, including defensive and aggressive labor supply phases.
Improves deep learning robustness by considering task and model.
problem Adversarial attacks on deep learning systems.
method Binary and interval label encoding strategy to redefine classification tasks and design corresponding loss functions.
result Our method enhances robustness without sacrificing accuracy.