Differential privacy reduces bias in adaptive data gathering.
problem Bias in adaptive data gathering, including numeric and complex data.
method Apply differential privacy to reduce bias and correct p-values.
result Near optimal regret bounds for differentially private bandit algorithms.
New method learns collective behavior without communication.
problem Learning optimal individual rules for gathering without communication.
method Collective multi-agent reinforcement learning.
result Gathering behavior can be learned without communication.
BED-LLM uses Bayesian experimental design to improve LLMs' information gathering.
problem Improving LLMs' ability to gather information adaptively.
method Iteratively choosing questions to maximize expected information gain using a probabilistic model.
result BED-LLM achieves substantial performance gains compared to other adaptive design strategies.
FairBED: A Bayesian Experimental Design Approach to Gathering Fairer Data
problem Ensuring fairness in machine learning
method Bayesian Experimental Design
result Improved fairness-accuracy trade-offs
We give a new proof of the representation of implied volatility as a time-average of weighted expectations of local or stochastic volatility. With this proof we clarify the question of existence of 'forward implied variance' in the original derivation of Gatheral, who introduced this representation in his book 'The Vol…
In this article we propose a generalisation of the recent work of Gatheral and Jacquier on explicit arbitrage-free parameterisations of implied volatility surfaces. We also discuss extensively the notion of arbitrage freeness and Roger Lee's moment formula using the recent analysis by Roper. We further exhibit an arbit…
In this short note, we prove by an appropriate change of variables that the SVI implied volatility parameterization presented in Gatheral's book and the large-time asymptotic of the Heston implied volatility agree algebraically, thus confirming a conjecture from Gatheral as well as providing a simpler expression for th…
Investment tool predicts higher returns for Madrid real estate units.
problem Determining which real estate units have higher returns to investment in Madrid.
method Data collection from Idealista.com, descriptive statistics, return index, machine learning algorithms.
result Introduction of machine learning algorithms for rental real estate price prediction.
Neural nets analyze Magic: the Gathering cards for text and images.
problem Classify Magic: the Gathering cards into multiple categories.
method Used Convolutional and Recurrent Neural Networks to analyze card text and images.
result Developed a method to generate card text matching an input image.
Robots gather information resiliently despite failures and attacks.
problem Resilient information gathering in adversarial or failure-prone environments.
method First scalable algorithm for minimal communication, system-wide resiliency, and provable approximation performance.
result Algorithm ensures optimal or near-optimal solutions for any number of failures and attacks.
Modified model prevents volatility from approaching zero.
problem Volatility in the Gatheral model can approach zero, making it statistically indistinguishable.
method Proposed a modified model with Skorokhod reflection to prevent volatility from approaching zero.
result The modified model prevents volatility from approaching zero, preserving the model's flexibility.
We explore training an automatic modality tagger. Modality is the attitude that a speaker might have toward an event or state. One of the main hurdles for training a linguistic tagger is gathering training data. This is particularly problematic for training a tagger for modality because modality triggers are sparse for…
Study benchmarks label noise detection methods, identifying best practices.
problem Label noise in real-world datasets affects model performance and evaluation reliability.
method Decomposed detection methods into label agreement, aggregation, and information gathering components; introduced a unified benchmark task and novel metric.
result In-sample probability aggregation with logit margin label agreement function achieves best results across scenarios.
Two-dimensional case in the theory of dynamical systems admitting the normal shift differs crucially from multidimensional case. Features of two-dimensional case are gathered and studied in this thesis.
Paper simplifies complex AI exploration by predicting future rewards.
problem Training machines to optimally gather complex information.
method Developed a denser reward structure using cross-value to decouple exploration and exploitation.
result Demonstrated successful learning of challenging tasks without shaping or bonuses.
New meta-RL method avoids exploration-exploitation trade-off.
problem Learning to explore and exploit simultaneously in meta-RL.
method Developed new objectives for exploration and exploitation.
result DREAM outperforms existing methods on complex tasks.
Framework detects cyber threats from Twitter tweets.
problem Time-consuming manual extraction of cyber threat intelligence.
method Novelty detection model trained on CVE data.
result F1-score of 0.643 for classifying cyber threat tweets.
Formula adjusts steady-state models for control confounding.
problem Learning steady-state models from operational data can be flawed due to control confounding.
method Derives a formula to adjust for control confounding using structural dynamical causal models.
result Estimates a causal steady-state model from closed-loop operational data.
Proposes a max-utility arm selection strategy for reducing cumulative regret in sequential query recommendations.
problem Reduces cumulative regret in sequential query recommendations for closed loop interactive learning settings.
method Proposes a max-utility arm selection strategy based on the maximum utility of arms.
result Improves cumulative regret substantially compared to baseline algorithms and random selection.
Foundation models struggle with multi-turn exploration but can learn through regular summaries.
problem Foundation models struggle with multi-turn exploration in dynamic environments.
method Implemented a text-based version of the Alchemy environment to test multi-trial learning. Prompting models to summarize their observations at regular intervals enabled them to improve across trials and adapt to changes.
result Foundation models can improve through regular summaries, enabling multi-trial learning and adaptation.
In this survey article we gather classical as well as recent results on minimal geodesics of Riemannian or Finsler metrics, giving special attention to the two-dimensional case. Moreover, we present open problems together with some first ideas as to the solutions.
Assuming geometric Brownian motion as unaffected price process S0, Gatheral & Schied (2011) derived a strategy for optimal order execution that reacts in a sensible manner on market changes but can still be computed in closed form. Here we will investigate the robustness of this strategy with respect to misspecifica…
Abstract relates G2 instantons to Seiberg-Witten monopoles.
problem Understanding the connection between G2 instantons and Seiberg-Witten monopoles.
method Relates Seiberg-Witten monopoles, Fueter sections, and G2 instantons.
result Describes a conjectural relation between G2 instantons and Seiberg-Witten monopoles.
Survey on metrics and assembly maps in positive scalar curvature.
problem Positive Scalar Curvature metrics and their relation to assembly maps.
method Gromov-Lawson index and Baum-Connes assembly map.
result Survey of results connecting metrics and assembly maps.
Gathering the most information by picking the least amount of data is a common task in experimental design or when exploring an unknown environment in reinforcement learning and robotics. A widely used measure for quantifying the information contained in some distribution of interest is its entropy. Greedily minimizing…
Hybrid-FL improves ML model accuracy in non-IID data environments.
problem Performance degradation in FL due to non-IID data.
method Hybrid-FL combines client and server learning, selecting optimal clients and data for aggregation.
result 13.5% higher classification accuracy than previous methods.
Mounting evidences are being gathered suggesting that income and wealth distribution in various countries or societies follow a robust pattern, close to the Gibbs distribution of energy in an ideal gas in equilibrium, but also deviating significantly for high income groups. Application of physics models seem to provide…
It is classical that given any Seifert structure on N, Reidemeister-Schreier's algorithm produces a presentation of all index 2 subgroups of the fundamental group of N, described as the fundamental group of some Seifert manifolds. The new result of this article is concise formulas that gather all possible cases.
Study confirms rough volatility in financial data, independent of microstructure noise.
problem Characterizing volatility in financial markets, especially rough volatility.
method Used range-based volatility estimators to confirm findings from fractional behavior.
result Log-volatility behaves like fractional Brownian motion with an even lower Hurst exponent.
There are two schools of thought regarding market impact modeling. On the one hand, seminal papers by Almgren and Chriss introduced a decomposition between a permanent market impact and a temporary (or instantaneous) market impact. This decomposition is used by most practitioners in execution models. On the other hand,…
We introduce a probabilistic model of labor markets for university graduates, in particular, in Japan. To make a model of the market efficiently, we take into account several hypotheses. Namely, each company fixes the (business year independent) number of opening positions for newcomers. The ability of gathering newcom…
Paper explores multimodal methods for detecting urban micro-events.
problem Detecting urban micro-events in cities with limited geographical coverage.
method Exploring multimodal fusion methods including early, late, hybrid fusion and representation learning.
result Multimodal approach yields higher performance than unimodal alternatives.
Bayesian BIC for multi-trial data improves VAR model order selection.
problem Optimal VAR model order selection for multi-trial event-based data.
method Derive and apply Bayesian Information Criterion (BIC) for multi-trial ensemble data.
result Multi-trial BIC successfully recovers real model order and estimates small model order.
Aims to improve personalized treatment decisions through Bayesian experimental design.
problem Evaluating and improving personalized treatment decisions in contexts like customer service.
method Model-agnostic Bayesian Experimental Design to efficiently gather data and avoid highly sub-optimal treatments.
result Our method achieves superior performance in evaluating and improving treatment decisions compared to traditional approaches.
Increasingly, a huge amount of statistics have been gathered which clearly indicates that income and wealth distributions in various countries or societies follow a robust pattern, close to the Gibbs distribution of energy in an ideal gas in equilibrium. However, it also deviates in the low income and more significantl…
DRLViz interprets deep RL agent memory for better understanding.
problem Understanding complex deep RL agent memory.
method Visual analytics interface to reduce and interpret memory vectors.
result Experts can better understand and investigate agent decisions.
Study optimizes scoring rules for incentivizing agent's information gathering in online settings.
problem Optimizing incentives for agents to acquire information in online settings.
method Designing a sample-efficient algorithm that tailors the UCB algorithm to the strategic agent's model.
result Achieves sublinear T2/3-regret after T iterations, independent of the number of states. Distributed, online data mining systems have emerged as a result of applications requiring analysis of large amounts of correlated and high-dimensional data produced by multiple distributed data sources. We propose a distributed online data classification framework where data is gathered by distributed data sources and…
Abstract notes on robust statistical learning theory.
problem Developing robust estimators for statistical learning.
method Stressing principles of robust estimators construction and analysis.
result Emphasizes main principles of robust estimators construction and analysis.
The BBF, SABR, and rough SABR formulas provide nearly arbitrage-free implied vol approximations.
problem Arbitrage in implied volatility calculations.
method Analytical proofs for BBF, SABR, and rough SABR formulas under specific models.
result These formulas offer asymptotically arbitrage-free approximations of implied volatility.
Study on small-time behavior of rough Bergomi model using large deviations principle.
problem Understanding the small-time behavior of the rough Bergomi model.
method Proved a large deviations principle for a rescaled log stock price process.
result Characterized the small-time behavior of the implied volatility.
This paper compiles formulas involving differential operators and interior products.
problem Scattered identities in differential geometry involving various operators.
method Compilation and extension of formulas using the Schouten-Nijenhuis bracket and interior product.
result New formulas involving the de Rham codifferential and interior product.
Study compares and analyzes various counterfactual estimators.
problem Analyzing and comparing different off-policy estimators.
method Detailed comparison of Empirical Average, Basic Importance Sampling, and Normalized Importance Sampling in different regimes.
result Fused estimators outperform basic ones but can be improved.
Hybrid LSMC-PDE method for Bermudan options under GDMR model.
problem Pricing Bermudan options under the GDMR model.
method Adapted Hybrid LSMC-PDE framework, combining Monte Carlo and PDE methods.
result Hybrid approach yields more accurate and lower error estimates than plain LSMC.
New method calibrates rough stochastic volatility models quickly.
problem Calibrating rough stochastic volatility models is expensive and time-consuming.
method Combines Levenberg-Marquardt with neural networks for fast calibration.
result Neural network approximates implied volatility map efficiently.
We explore how to improve machine translation systems by adding more translation data in situations where we already have substantial resources. The main challenge is how to buck the trend of diminishing returns that is commonly encountered. We present an active learning-style data solicitation algorithm to meet this c…
DPPy offers Python tools for sampling DPPs.
problem Sampling from DPPs is challenging.
method Gathers exact and approximate sampling algorithms for finite and continuous DPPs.
result DPPy provides Python tools for DPP sampling.
Optimization of very expensive black-box functions requires utilization of maximum information gathered by the process of optimization. Model Guided Sampling Optimization (MGSO) forms a more robust alternative to Jones' Gaussian-process-based EGO algorithm. Instead of EGO's maximizing expected improvement, the MGSO use…