New method for fair resource allocation in AI-aware networks with unknown utility functions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
SMOTE-DP enhances synthetic data privacy without sacrificing utility.
Paper establishes utility theory for synthetic data generation.
TVineSynth generates synthetic data to balance privacy and utility.
Paper evaluates synthetic retail data for fidelity, utility, and privacy.
A new model uses neural networks for consistent discrete choice analysis.
The paper tackles optimal policy learning with asymmetric counterfactual utilities in healthcare decisions.
Optimal defenses protect FL models from gradient reconstruction attacks.
BUDS balances privacy and utility by shuffling data, achieving strong privacy with minimal loss.
Advocates focusing on utility functions to avoid unfair outcomes.
The paper proposes using density ratio estimation to evaluate synthetic data quality.
Study shows privacy and utility trade-offs in synthetic data models, impacting fairness and real-world performance.
Modern portfolio theory(MPT) addresses the problem of determining the optimum allocation of investment resources among a set of candidate assets. In the original mean-variance approach of Markowitz, volatility is taken as a proxy for risk, conflating uncertainty with risk. There have been many subsequent attempts to al…
RUMBoost combines RUMs and deep learning for better choice modelling.
New method generates private synthetic data with optimal utility for smooth queries.
The study bounds the utility of empirically optimal portfolios using stock return data.
Neural networks approximate random utility models for choice prediction.
This paper proposes a systematic framework to design a classification model that yields a classifier which optimizes a utility function based on prior knowledge. Specifically, as the data size grows, we prove that the produced classifier asymptotically converges to the optimal classifier, an extended version of the Bay…
Study reveals Data Shapley's inconsistent performance in data selection tasks.
Specifying utility functions is a key step towards applying the discrete choice framework for understanding the behaviour processes that govern user choices. However, identifying the utility function specifications that best model and explain the observed choices can be a very challenging and time-consuming task. This …
Sequential pattern mining is an interesting research area with broad range of applications. Most prior research on sequential pattern mining has considered point-based data where events occur instantaneously. However, in many application domains, events persist over intervals of time of varying lengths. Furthermore, tr…
Differential privacy is a mathematical framework for privacy-preserving data analysis. Changing the hyperparameters of a differentially private algorithm allows one to trade off privacy and utility in a principled way. Quantifying this trade-off in advance is essential to decision-makers tasked with deciding how much p…
Synthetic tabular data synthesis models balance utility and risk.
New ranking system balances fairness and user utility.
Research optimizes a small RES utility's portfolio by dynamically trading in German electricity markets.
User releases data to service provider while balancing privacy and utility.
DiPriMe forests use private medians to create balanced tree splits for privacy-protected data.
FELICIA uses a centralized adversary to improve synthetic medical image generation.
Study uses reinforcement learning to optimize portfolios under recursive utility.
Accurately learning from user data while providing quantifiable privacy guarantees provides an opportunity to build better ML models while maintaining user trust. This paper presents a formal approach to carrying out privacy preserving text perturbation using the notion of dx-privacy designed to achieve geo-indistingui…
In many scenarios, humans prefer a text-based representation of quantitative data over numerical, tabular, or graphical representations. The attractiveness of textual summaries for complex data has inspired research on data-to-text systems. While there are several data-to-text tools for time series, few of them try to …
Investigations have been performed into using clustering methods in data mining time-series data from smart meters. The problem is to identify patterns and trends in energy usage profiles of commercial and industrial customers over 24-hour periods, and group similar profiles. We tested our method on energy usage data p…
Optimizes football play calls using reinforcement learning.
With a rapidly increasing number of devices connected to the internet, big data has been applied to various domains of human life. Nevertheless, it has also opened new venues for breaching users' privacy. Hence it is highly required to develop techniques that enable data owners to privatize their data while keeping it …
NPO method improves LLM unlearning without catastrophic collapse.
The input data features set for many data driven tasks is high-dimensional while the intrinsic dimension of the data is low. Data analysis methods aim to uncover the underlying low dimensional structure imposed by the low dimensional hidden parameters by utilizing distance metrics that consider the set of attributes as…
Study learns linear utility functions from comparisons, showing learnability gaps between passive and active learning.
We consider a framework involving behavioral economics and machine learning. Rationally inattentive Bayesian agents make decisions based on their posterior distribution, utility function and information acquisition cost Renyi divergence which generalizes Shannon mutual information). By observing these decisions, how ca…
Quantifying the importance of each training point to a learning task is a fundamental problem in machine learning and the estimated importance scores have been leveraged to guide a range of data workflows such as data summarization and domain adaption. One simple idea is to use the leave-one-out error of each training …
Recommendation systems are ubiquitous and impact many domains; they have the potential to influence product consumption, individuals' perceptions of the world, and life-altering decisions. These systems are often evaluated or trained with data from users already exposed to algorithmic recommendations; this creates a pe…
The remarkable success of machine learning, especially deep learning, has produced a variety of cloud-based services for mobile users. Such services require an end user to send data to the service provider, which presents a serious challenge to end-user privacy. To address this concern, prior works either add noise to …
It is becoming increasingly clear that users should own and control their data. Utility providers are also becoming more interested in guaranteeing data privacy. As such, users and utility providers should collaborate in data privacy, a paradigm that has not yet been developed in the privacy research community. We intr…
Pruning neural networks adds differential privacy noise, preserving data utility.
Inpatient care is a large share of total health care spending, making analysis of inpatient utilization patterns an important part of understanding what drives health care spending growth. Common features of inpatient utilization measures include zero inflation, over-dispersion, and skewness, all of which complicate st…
Random utility theory models an agent's preferences on alternatives by drawing a real-valued score on each alternative (typically independently) from a parameterized distribution, and then ranking the alternatives according to scores. A special case that has received significant attention is the Plackett-Luce model, fo…
Healthcare data continues to flourish yet a relatively small portion, mostly structured, is being utilized effectively for predicting clinical outcomes. The rich subjective information available in unstructured clinical notes can possibly facilitate higher discrimination but tends to be under-utilized in mortality pred…
ARF synthesizes epidemiological data to match original findings.
Enhances DP linear regression using public data moments.