Develops a new method for summarizing ranking data using bucket orders.
problem Summarizing ranking data without a vector space structure.
method Introduces a mass transportation metric to quantify bucket representations and selects optimal bucket orders.
result Optimal bucket orders minimize empirical distortion and provide sparse representations of ranking distributions.
Bucketed PCA-NN outperforms DNNs by 96% on MNIST.
problem Benchmarking deep neural networks for supervised classification.
method Applies PCA to individual buckets constructed in two phases, retains neural network architecture, and uses neurons that mirror input signals.
result Bucketed PCA-NN achieves 96% accuracy on MNIST, similar to DNNs.
A new hashing framework learns multiple hash codes for each image to improve hash bucket search efficiency.
problem Existing hashing methods fail to handle complex image retrieval scenarios efficiently.
method Multiple Code Hashing (MCH) framework with deep reinforcement learning.
result Significant improvement in hash bucket search performance compared to single-code methods.
In this work we compare different batch construction methods for mini-batch training of recurrent neural networks. While popular implementations like TensorFlow and MXNet suggest a bucketing approach to improve the parallelization capabilities of the recurrent training process, we propose a simple ordering strategy tha…
New bucketing scheme improves Byzantine robustness for heterogeneous data.
problem Byzantine attacks on federated learning with heterogeneous data.
method Bucketing scheme to adapt robust algorithms to non-iid data.
result Bucketing scheme ensures convergence against Byzantine attacks.
The aim of this paper is to present a dual-term structure model of interest rate derivatives in order to solve the two hardest problems in financial modeling: the exact volatility calibration of the entire swaption matrix, and the calculation of bucket vegas for structured products. The model takes a series of long-ter…
We study the problem of maximizing a monotone submodular function subject to a cardinality constraint k, with the added twist that a number of items τ from the returned set may be removed. We focus on the worst-case setting considered in (Orlin et al., 2016), in which a constant-factor approximation guarantee was g…
New method improves variational bounds on partition function.
problem Computing the partition function of discrete graphical models is intractable.
method Combines gauge transformations with weighted mini-bucket elimination (WMBE).
result WMBE-G strictly improves earlier WMBE approximation for symmetric models.
New method improves probabilistic model inference.
problem Computing partition function is hard and variational methods often fail.
method Sequential summation over variables with mini-bucket elimination and renormalization.
result Robust approximate algorithms show good performance on various models.
Zap predicts user behavior online using diverse techniques.
problem Predicting user behavior on websites.
method Combines sequential data processing techniques with Bloom filters, bucketing, and model calibration.
result Creates website- and task-specific models without website-specific code.
Model predicts Chinese stock market liquidity and customer order behavior.
problem Understanding market liquidity and customer order behavior in the Chinese stock market.
method Dual state-space model using Fourier transform to connect volume-at-price buckets to correlations.
result Customer orders are correlated with market sentiment and stock returns, not with bond returns.
Study examines meso-scale order flows and LOB resilience.
problem Understanding price formation and LOB resilience on meso-scale.
method Empirical analysis of order flows and LOB metrics from tick data.
result Limit order flows and shape are key predictors of price formation and LOB resilience.
SAFLe solves federated learning's trade-off between non-linearity and scalability.
problem Federated Learning's high communication overhead and performance collapse on non-IID data.
method SAFLe introduces a structured head of bucketed features and sparse, grouped embeddings, mathematically equivalent to a high-dimensional linear regression.
result SAFLe achieves a new state-of-the-art in analytic FL, outperforming linear AFL and multi-round DeepAFL.
Paper assesses how pandemic data impacts mortality models.
problem Impact of pandemic data on mortality projections.
method Calibrated Li & Lee model with transformed weekly data.
result Impact quantified, scenarios generated for future mortality.
A new uncertainty principle helps traders better understand market activity.
problem Understanding high-frequency market activity and correlation.
method Integrates market activity, order-flow overlap, and response time into a clock-dependent uncertainty principle.
result Six rules of thumb for traders operating at market-making frequencies.
We present a HJM approach to the projection of multiple yield curves developed to capture the volatility content of historical term structures for risk management purposes. Since we observe the empirical data at daily frequency and only for a finite number of time-to-maturity buckets, we propose a modelling framework w…
New algorithms mitigate impairment effect in stochastic bandits.
problem Impairment effect in stochastic multi-armed bandits.
method Developed two novel algorithms with bucketing to address impairment and temporal constraints.
result Achieved sublinear regret, explicitly capturing impairment cost.
Develops optimal currency hedging strategy for fund managers considering liquidity risk.
problem Choosing optimal foreign exchange (FX) hedge tenors to maximize carry returns within liquidity constraints.
method Time-dispersing total hedge value into future time buckets, maximizing FX carry benefit while adhering to liquidity risk metric (CFaR).
result Hedging strategy operates within liquidity budget, demonstrating practical insights for fund managers.
ForestDSH hashes improve nearest neighbor search in high-dimensional data.
problem High-dimensional classification and nearest neighbor search.
method Distribution-sensitive hashing using a forest of decision trees.
result ForestDSH hashes outperform LSH and state-of-the-art methods in speed and accuracy.
ClusterLOB clusters market events to identify different trading behaviors.
problem Understanding market microstructure and participant behavior in financial markets.
method ClusterLOB uses K-means++ algorithm to cluster market events based on six time-dependent features.
result ClusterLOB identifies three distinct trading behaviors: directional, opportunistic, and market-making participants.
Study of Polymarket's prediction market microstructure using tick-level order book data.
problem Understanding the microstructure of decentralized prediction markets.
method Analysis of a continuous tick-level order book feed and on-chain trade records.
result Trade direction inferred from Polymarket's public order-book feed disagrees with on-chain data in ~59% of cases.
seMCD computes depth functions with statistical guarantees using sequential Monte Carlo.
problem Computing depth functions is computationally challenging, especially in high dimensions.
method Sequential Monte Carlo methodology with theoretical and empirical guarantees.
result The seMCD method provides accurate depth approximations with fewer samples than traditional methods.
Study finds IBS useful for predicting ETF price movements.
problem Predicting short-term price movements in country ETFs.
method Quantitative analysis of historical price data using Mean Reversion.
result IBS can be a useful technical indicator for ETFs.
Algorithm distinguishes light-tailed from non-light-tailed distributions.
problem Characterize the tail of a distribution using hazard rate.
method Careful bucketing scheme based on hazard rate.
result Polynomial number of samples required for success.
During recent years the counterparty risk subject has received a growing attention because of the so called Basel Accord. In particular the Basel III Accord asks the banks to fulfill finer conditions concerning counterparty credit exposures arising from banks' derivatives, securities financing transactions, default and…
New federated learning protocols resist Byzantine failures and offer privacy guarantees.
problem Resisting Byzantine failures in federated learning.
method Proposes robust federated learning protocols with optimal statistical rates and privacy guarantees.
result Achieves nearly optimal statistical rates and tight rate in terms of all parameters for strongly convex losses.
Market events such as order placement and order cancellation are examples of the complex and substantial flow of data that surrounds a modern financial engineer. New mathematical techniques, developed to describe the interactions of complex oscillatory systems (known as the theory of rough paths) provides new tools for…
A fast, non-iterative method for missing value imputation using random trees.
problem Missing value imputation in large and high-dimensional datasets.
method Recursive semi-random hyperplane cuts to assign observations to buckets and calculate weighted averages as imputations.
result Significantly faster than chained equations and scales well to large datasets.
Study explores reinforcement learning in a complex game environment, analyzing rule inference and policy learning.
problem Learning optimal policies in environments with hidden rules.
method Investigated using the Game Of Hidden Rules (GOHR) environment, employing Feature-Centric and Object-Centric state representations with a Transformer-based A2C algorithm.
result Transformer-based A2C models outperform traditional methods in GOHR, demonstrating the effectiveness of representation strategies.
DataRater learns which data points are most valuable for training models.
problem Training model efficiency depends on high-quality training data.
method Meta-learning to estimate the value of data points for training.
result Meta-learning improves compute efficiency by filtering data effectively.
Paper compares ETF and futures carry rates in segmented Bitcoin markets.
problem Limitations in cross-margining between spot Bitcoin and CME futures.
method Estimates carry rates from IBIT options and CME futures, uses put-call parity and daily ETF holdings.
result Mean and median wedge in carry rates is 2.58 and 2.52 percent, respectively.
Model estimates non-reported GHG emissions for companies using machine learning.
problem Incomplete GHG emissions reporting by companies.
method Interpretable machine learning model tailored for non-reporting companies.
result Model accurately estimates emissions for diverse company groups.
Norm-range partition improves MIPS search efficiency by reducing query complexity.
problem Efficiently searching for maximum inner product in large datasets.
method Norm-range partition technique that divides datasets into sub-datasets with similar norms and builds independent hash indexes.
result Significantly reduces the number of probed buckets for LSH-based MIPS algorithms.
Study variance-optimal hedging of forward curve derivatives under stochastic volatility.
problem Variance-optimal hedging of forward curve derivatives with stochastic volatility.
method Assumes HJM-Musiela dynamics modulated by stochastic covariance, uses Galtchouk-Kunita-Watanabe projection.
result Density of finite-maturity strategies, convergence of finite-rank projections, decomposition of hedging error.
Contextual bandit algorithms have become popular for online recommendation systems such as Digg, Yahoo! Buzz, and news recommendation in general. \emph{Offline} evaluation of the effectiveness of new algorithms in these applications is critical for protecting online user experiences but very challenging due to their "p…
Paper introduces active and passive causal inference techniques.
problem Causal inference in machine learning.
method Categorizes causal inference techniques into active and passive approaches.
result Describes and discusses various causal inference methods.
Backtests of structured strategies lose much of their predictive power in live trading.
problem Uncertainty in how marketed backtests predict live performance of structured strategies.
method Analysis of 1,726 structured strategies from ten global institutions.
result Raw backtests have limited portability into live trading and deteriorate sharply.
This paper examines how ads on LinkedIn affect user behavior over time.
problem Understanding long-term impact of ads on user engagement and revenue.
method Conducted experiments with randomized member buckets to measure short and long-term effects of ads density.
result Long-term impact of ads is much smaller than short-term impact, and different user cohorts react differently over time.
Novel OTT method for cryptocurrency trading offers high annualized profit.
problem Quantifying and exploiting trading opportunities in cryptocurrency markets.
method Bi-objective convex optimization for balancing profit and risk.
result Annualized profit of 15.49% in cryptocurrency market from 2020 to 2022.
New metrics quantify implementation risk in portfolio backtesting, revealing systematic differences in engine implementations.
problem Systematic divergence in backtested portfolio metrics due to differences in engine implementations.
method Formalized implementation risk, proposed four metrics, executed 15 strategies through five engines, analyzed source-code defects.
result Implementation risk introduces measurable ambiguity in performance attribution, but does not alter investment decisions.
After the release of the final accounting standards for impairment in July 2014 by the IASB, banks will face the next significant methodological challenge after Basel 2. In this paper, first methodological thoughts are presented, and ways how to approach underlying questions are proposed. It starts with a detailed disc…
VAIOM models financial returns using continuous input and categorical output.
problem Modeling continuous, noisy, and heterogeneous financial data.
method VAIOM is a decoder-only Transformer that separates input representation from output likelihood.
result VAIOM models outperform fixed single-bar LightGBM baseline in both Test halves.
This paper improves prediction uncertainty estimation by inferring variation from neuron activation strength.
problem Estimating prediction uncertainty from ensemble methods is expensive and inaccurate.
method Introduced randomness into model training and inferred prediction variation from neuron activation strength.
result Average R squared on MovieLens is 0.56 and on Criteo is 0.81, with strong performance in variation detection.
Market Microstructure is the investigation of the process and protocols that govern the exchange of assets with the objective of reducing frictions that can impede the transfer. In financial markets, where there is an abundance of recorded information, this translates to the study of the dynamic relationships between o…
Study evaluates information leakage in Polymarket markets, finding limited applicability and resolution ambiguity.
problem Limited applicability of ILS-dl framework across Polymarket markets.
method Scaling from single-case to population-scale evaluation using ILS-dl framework.
result Only 0.7% of candidate markets yield computable ILS-dl values, and resolution semantics are the main obstacle.
Proves HNN extensions of nilpotent groups are left-orderable, constructs non-left-orderable examples.
problem Characterizing left-orderability in HNN extensions of groups.
method Analyzes HNN extensions of torsion-free nilpotent groups and left-orderable groups.
result Constructs examples of non-left-orderable HNN extensions of left-orderable groups.
New criterion links circular-orderability to left-orderability.
problem Understanding when circularly-orderable groups are left-orderable.
method Introducing a criterion based on the product of a group with integers.
result Groups are left-orderable if their product with integers is circularly-orderable.
The paper characterizes L-space 3-manifolds using quandle orderability.
problem Characterizing L-space 3-manifolds using quandle orderability.
method Using quandles and their extensions, the paper investigates orderability and circular orderability conditions.
result The n-quandle Qn(L) of a link quandle is not right circularly orderable.