Heterogeneous SVO leads to diverse policies in sequential social dilemmas.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Research tackles alliance formation in many-player zero-sum games, showing reinforcement learning fails but a contract mechanism can help.
This paper considers stochastic bandits with side observations, a model that accounts for both the exploration/exploitation dilemma and relationships between arms. In this setting, after pulling an arm i, the decision maker also observes the rewards for some other actions related to i. We will see that this model is su…
Controversies around race and machine learning have sparked debate among computer scientists over how to design machine learning systems that guarantee fairness. These debates rarely engage with how racial identity is embedded in our social experience, making for sociological and psychological complexity. This complexi…
Sequential portfolio selection has attracted increasing interests in the machine learning and quantitative finance communities in recent years. As a mathematical framework for reinforcement learning policies, the stochastic multi-armed bandit problem addresses the primary difficulty in sequential decision making under …
Bayesian active learning tackles nuisance parameters, leading to bias and dilemmas.
This paper considers a statistical signal processing problem involving agent based models of financial markets which at a micro-level are driven by socially aware and risk- averse trading agents. These agents trade (buy or sell) stocks by exploiting information about the decisions of previous agents (social learning) v…
This survey outlines methods to ensure fairness in machine learning.
SLIM model tackles graph classification by resolving part-interaction dilemmas.
We propose a unified mechanism for achieving coordination and communication in Multi-Agent Reinforcement Learning (MARL), through rewarding agents for having causal influence over other agents' actions. Causal influence is assessed using counterfactual reasoning. At each timestep, an agent simulates alternate actions t…
New algorithms for batch decision-making with high-dimensional user data.
Thompson sampling is an efficient algorithm for sequential decision making, which exploits the posterior uncertainty to address the exploration-exploitation dilemma. There has been significant recent interest in integrating Bayesian neural networks into Thompson sampling. Most of these methods rely on global variable u…
Paper analyzes convergence of continual learning with adaptive methods.
Many real-world problems like Social Influence Maximization face the dilemma of choosing the best out of options at a given time instant. This setup can be modeled as a combinatorial bandit which chooses out of arms at each time, with an aim to achieve an efficient trade-off between exploration and expl…
NFTs raise concerns like scams, racism, and sexism; centralization vs decentralization debate.
Bayesian approach learns causal concepts from diverse social surveys.
Communities in social networks or graphs are sets of well-connected, overlapping vertices. The effectiveness of a community detection algorithm is determined by accuracy in finding the ground-truth communities and ability to scale with the size of the data. In this work, we provide three contributions. First, we show t…
Adding data can sometimes hurt model performance in multi-source healthcare tasks.
EmTract extracts emotions from financial social media text.
Algorithm learns fair division from noisy feedback in uncertain markets.
New CUSUM method detects changes in Hawkes networks efficiently.
Efficiently selects seed nodes to maximize content influence in unknown social networks.
The Data Clustering (DC) problem is of central importance for the area of Machine Learning (ML), given its usefulness to represent data structural similarities from input spaces. Differently from Supervised Machine Learning (SML), which relies on the theoretical frameworks of the Statistical Learning Theory (SLT) and t…
In this paper we extend the investigation into the transition from sure to probabilistic sniping as introduced in Menkveld and Zoican \cite{mz2017}. In that paper, the authors introduce a stylized version of a competitive game in which high frequency traders (HFTs) interact with each other and liquidity traders. The au…
Programming has been an important skill for researchers and practitioners in computer science and other related areas. To learn basic programing skills, a long-time systematic training is usually required for beginners. According to a recent market report, the computer software market is expected to continue expanding …
Detection of emerging topics are now receiving renewed interest motivated by the rapid growth of social networks. Conventional term-frequency-based approaches may not be appropriate in this context, because the information exchanged are not only texts but also images, URLs, and videos. We focus on the social aspects of…
Algorithm reduces long-term policy regret in ML decision-making.
ASD algorithm maximizes model estimates by adaptively labeling points.
Leo Breiman's Rashomon Effect and Occam Dilemma are re-evaluated in the context of modern machine learning.
The paper examines how slightly biasing towards under-represented groups in sequential selection processes can lead to long-term fairness.
Margin enlargement over training data has been an important strategy since perceptrons in machine learning for the purpose of boosting the robustness of classifiers toward a good generalization ability. Yet Breiman (1999) showed a dilemma that a uniform improvement on margin distribution does NOT necessarily reduces ge…
The study examines causal razors and their logical relations, highlighting a dilemma in causal discovery.
Framework for online resource allocation using social welfare functions.
For many years, achievements and discoveries made by scientists are made aware through research papers published in appropriate journals or conferences. Often, established scientists and especially newbies are caught up in the dilemma of choosing an appropriate conference to get their work through. Every scientific con…
CRB tackles rising rewards in combinatorial online learning.
Whenever customers' choices (e.g. to buy or not a given good) depend on others choices (cases coined 'positive externalities' or 'bandwagon effect' in the economic literature), the demand may be multiply valued: for a same posted price, there is either a small number of buyers, or a large one -- in which case one says …
Many kinds of data can be represented as a network or graph. It is crucial to infer the latent structure underlying such a network and to predict unobserved links in the network. Mixed Membership Stochastic Blockmodel (MMSB) is a promising model for network data. Latent variables and unknown parameters in MMSB have bee…
New method improves OOD detection without sacrificing generalization.
New method evaluates multiple social disparities using machine learning.
The explore{exploit dilemma is one of the central challenges in Reinforcement Learning (RL). Bayesian RL solves the dilemma by providing the agent with information in the form of a prior distribution over environments; however, full Bayesian planning is intractable. Planning with the mean MDP is a common myopic approxi…
Generative models of graph structure have applications in biology and social sciences. The state of the art is GraphRNN, which decomposes the graph generation process into a series of sequential steps. While effective for modest sizes, it loses its permutation invariance for larger graphs. Instead, we present a permuta…
New method uses path signatures for efficient likelihood estimation in time-series data.
SocialInteractionGAN generates realistic human interactions from low-dimensional data.
Study online learning with off-policy feedback in adversarial bandit problems.
The multi-armed bandit (MAB) problem is a classic example of the exploration-exploitation dilemma. It is concerned with maximising the total rewards for a gambler by sequentially pulling an arm from a multi-armed slot machine where each arm is associated with a reward distribution. In static MABs, the reward distributi…
Multinomial logit bandit is a sequential subset selection problem which arises in many applications. In each round, the player selects a -cardinality subset from candidate items, and receives a reward which is governed by a {\it multinomial logit} (MNL) choice model considering both item utility and substitution…
We present PredRNN++, an improved recurrent network for video predictive learning. In pursuit of a greater spatiotemporal modeling capability, our approach increases the transition depth between adjacent states by leveraging a novel recurrent unit, which is named Causal LSTM for re-organizing the spatial and temporal m…
We model anomaly and change in data by embedding the data in an ultrametric space. Taking our initial data as cross-tabulation counts (or other input data formats), Correspondence Analysis allows us to endow the information space with a Euclidean metric. We then model anomaly or change by an induced ultrametric. The in…