Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,694 papers · 148 categories

Trend · papers per month

1122 · Nov 201419922001200920172026
45 results for Bucketing

Whereas most dimensionality reduction techniques (e.g. PCA, ICA, NMF) for multivariate data essentially rely on linear algebra to a certain extent, summarizing ranking data, viewed as realizations of a random permutation ΣΣ on a set of items indexed by i{1,,  n}i\in \{1,\ldots,\; n\}, is a great statistical challenge, due to…

2018-10-15abs ↗pdf ↗

A new hashing framework learns multiple hash codes for each image to improve hash bucket search efficiency.

problem Existing hashing methods fail to handle complex image retrieval scenarios efficiently.
method Multiple Code Hashing (MCH) framework with deep reinforcement learning.
result Significant improvement in hash bucket search performance compared to single-code methods.

The aim of this paper is to present a dual-term structure model of interest rate derivatives in order to solve the two hardest problems in financial modeling: the exact volatility calibration of the entire swaption matrix, and the calculation of bucket vegas for structured products. The model takes a series of long-ter…

2016-06-04abs ↗pdf ↗

Computing the partition function ZZ of a discrete graphical model is a fundamental inference challenge. Since this is computationally intractable, variational approximations are often used in practice. Recently, so-called gauge transformations were used to improve variational lower bounds on ZZ. In this paper, we pro…

2018-01-05abs ↗pdf ↗

Probabilistic graphical models are a key tool in machine learning applications. Computing the partition function, i.e., normalizing constant, is a fundamental task of statistical inference but it is generally computationally intractable, leading to extensive study of approximation methods. Iterative variational methods…

2018-03-14abs ↗pdf ↗

We study the problem of maximizing a monotone submodular function subject to a cardinality constraint kk, with the added twist that a number of items ττ from the returned set may be removed. We focus on the worst-case setting considered in (Orlin et al., 2016), in which a constant-factor approximation guarantee was g…

2017-06-15abs ↗pdf ↗

SAFLe solves federated learning's trade-off between non-linearity and scalability.

problem Federated Learning's high communication overhead and performance collapse on non-IID data.
method SAFLe introduces a structured head of bucketed features and sparse, grouped embeddings, mathematically equivalent to a high-dimensional linear regression.
result SAFLe achieves a new state-of-the-art in analytic FL, outperforming linear AFL and multi-round DeepAFL.

We develop an optimal currency hedging strategy for fund managers who own foreign assets to choose the hedge tenors that maximize their FX carry returns within a liquidity risk constraint. The strategy assumes that the offshore assets are fully hedged with FX forwards. The chosen liquidity risk metric is Cash Flow at R…

2019-03-15abs ↗pdf ↗

This paper introduces Zap, a generic machine learning pipeline for making predictions based on online user behavior. Zap combines well known techniques for processing sequential data with more obscure techniques such as Bloom filters, bucketing, and model calibration into an end-to-end solution. The pipeline creates we…

2018-07-16abs ↗pdf ↗

New federated learning protocols resist Byzantine failures and offer privacy guarantees.

problem Resisting Byzantine failures in federated learning.
method Proposes robust federated learning protocols with optimal statistical rates and privacy guarantees.
result Achieves nearly optimal statistical rates and tight rate in terms of all parameters for strongly convex losses.

This work proposes a non-iterative strategy for missing value imputations which is guided by similarity between observations, but instead of explicitly determining distances or nearest neighbors, it assigns observations to overlapping buckets through recursive semi-random hyperplane cuts, in which weighted averages are…

2019-11-15abs ↗pdf ↗

In this paper, we consider the problem of classification of MM high dimensional queries y1,,yMBSy^1,\cdots,y^M\in B^S to NN high dimensional classes x1,,xNASx^1,\cdots,x^N\in A^S where AA and BB are discrete alphabets and the probabilistic model that relates data to the classes P(x,y)P(x,y) is known. This problem has applications …

2019-05-11abs ↗pdf ↗

Study explores reinforcement learning in a complex game environment, analyzing rule inference and policy learning.

problem Learning optimal policies in environments with hidden rules.
method Investigated using the Game Of Hidden Rules (GOHR) environment, employing Feature-Centric and Object-Centric state representations with a Transformer-based A2C algorithm.
result Transformer-based A2C models outperform traditional methods in GOHR, demonstrating the effectiveness of representation strategies.

We present a HJM approach to the projection of multiple yield curves developed to capture the volatility content of historical term structures for risk management purposes. Since we observe the empirical data at daily frequency and only for a finite number of time-to-maturity buckets, we propose a modelling framework w…

2014-11-14abs ↗pdf ↗

Paper compares ETF and futures carry rates in segmented Bitcoin markets.

problem Limitations in cross-margining between spot Bitcoin and CME futures.
method Estimates carry rates from IBIT options and CME futures, uses put-call parity and daily ETF holdings.
result Mean and median wedge in carry rates is 2.58 and 2.52 percent, respectively.

Model predicts Chinese stock market liquidity and customer order behavior.

problem Understanding market liquidity and customer order behavior in the Chinese stock market.
method Dual state-space model using Fourier transform to connect volume-at-price buckets to correlations.
result Customer orders are correlated with market sentiment and stock returns, not with bond returns.

Model estimates non-reported GHG emissions for companies using machine learning.

problem Incomplete GHG emissions reporting by companies.
method Interpretable machine learning model tailored for non-reporting companies.
result Model accurately estimates emissions for diverse company groups.

Study variance-optimal hedging of forward curve derivatives under stochastic volatility.

problem Variance-optimal hedging of forward curve derivatives with stochastic volatility.
method Assumes HJM-Musiela dynamics modulated by stochastic covariance, uses Galtchouk-Kunita-Watanabe projection.
result Density of finite-maturity strategies, convergence of finite-rank projections, decomposition of hedging error.

Organic updates (from a member's network) and sponsored updates (or ads, from advertisers) together form the newsfeed on LinkedIn. The newsfeed, the default homepage for members, attracts them to engage, brings them value and helps LinkedIn grow. Engagement and Revenue on feed are two critical, yet often conflicting ob…

2019-01-29abs ↗pdf ↗

Backtests of structured strategies lose much of their predictive power in live trading.

problem Uncertainty in how marketed backtests predict live performance of structured strategies.
method Analysis of 1,726 structured strategies from ten global institutions.
result Raw backtests have limited portability into live trading and deteriorate sharply.

We study the effect of impairment on stochastic multi-armed bandits and develop new ways to mitigate it. Impairment effect is the phenomena where an agent only accrues reward for an action if they have played it at least a few times in the recent past. It is practically motivated by repetition and recency effects in do…

2018-11-22abs ↗pdf ↗

During recent years the counterparty risk subject has received a growing attention because of the so called Basel Accord. In particular the Basel III Accord asks the banks to fulfill finer conditions concerning counterparty credit exposures arising from banks' derivatives, securities financing transactions, default and…

2015-03-05abs ↗pdf ↗

We investigate the behavior of limit order books on the meso-scale motivated by order execution scheduling algorithms. To do so we carry out empirical analysis of the order flows from market and limit order submissions, aggregated from tick-by-tick data via volume-based bucketing, as well as various LOB depth and shape…

2017-08-09abs ↗pdf ↗

Novel OTT method for cryptocurrency trading offers high annualized profit.

problem Quantifying and exploiting trading opportunities in cryptocurrency markets.
method Bi-objective convex optimization for balancing profit and risk.
result Annualized profit of 15.49% in cryptocurrency market from 2020 to 2022.

seMCD computes depth functions with statistical guarantees using sequential Monte Carlo.

problem Computing depth functions is computationally challenging, especially in high dimensions.
method Sequential Monte Carlo methodology with theoretical and empirical guarantees.
result The seMCD method provides accurate depth approximations with fewer samples than traditional methods.

New metrics quantify implementation risk in portfolio backtesting, revealing systematic differences in engine implementations.

problem Systematic divergence in backtested portfolio metrics due to differences in engine implementations.
method Formalized implementation risk, proposed four metrics, executed 15 strategies through five engines, analyzed source-code defects.
result Implementation risk introduces measurable ambiguity in performance attribution, but does not alter investment decisions.

VAIOM models financial returns using continuous input and categorical output.

problem Modeling continuous, noisy, and heterogeneous financial data.
method VAIOM is a decoder-only Transformer that separates input representation from output likelihood.
result VAIOM models outperform fixed single-bar LightGBM baseline in both Test halves.

This paper improves prediction uncertainty estimation by inferring variation from neuron activation strength.

problem Estimating prediction uncertainty from ensemble methods is expensive and inaccurate.
method Introduced randomness into model training and inferred prediction variation from neuron activation strength.
result Average R squared on MovieLens is 0.56 and on Criteo is 0.81, with strong performance in variation detection.

Market events such as order placement and order cancellation are examples of the complex and substantial flow of data that surrounds a modern financial engineer. New mathematical techniques, developed to describe the interactions of complex oscillatory systems (known as the theory of rough paths) provides new tools for…

2013-07-27abs ↗pdf ↗

Study of Polymarket's prediction market microstructure using tick-level order book data.

problem Understanding the microstructure of decentralized prediction markets.
method Analysis of a continuous tick-level order book feed and on-chain trade records.
result Trade direction inferred from Polymarket's public order-book feed disagrees with on-chain data in ~59% of cases.

Study evaluates information leakage in Polymarket markets, finding limited applicability and resolution ambiguity.

problem Limited applicability of ILS-dl framework across Polymarket markets.
method Scaling from single-case to population-scale evaluation using ILS-dl framework.
result Only 0.7% of candidate markets yield computable ILS-dl values, and resolution semantics are the main obstacle.

ClusterLOB clusters market events to identify different trading behaviors.

problem Understanding market microstructure and participant behavior in financial markets.
method ClusterLOB uses K-means++ algorithm to cluster market events based on six time-dependent features.
result ClusterLOB identifies three distinct trading behaviors: directional, opportunistic, and market-making participants.