Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

87174261348 · Jun 202019922001200920172026
48 results for value extraction

This work introduces uncertainty principles to mitigate Maximal Extractable Value in blockchain systems.

problem Maximal Extractable Value (MEV) in decentralized systems due to transaction submission privacy and monopolist power.
method Unified approaches via uncertainty principles, akin to harmonic analysis and physics, to quantify trade-offs between transaction flexibility and user economic payoff.
result Demonstrates a quantitative trade-off between transaction flexibility and user economic payoff, analogous to the Nyquist-Shannon sampling theorem.

Defines cost of MEV and shows its relevance in various settings.

problem Excess value miners can realize by manipulating transaction order.
method Introduces a simple theoretical definition of cost of MEV, proves properties, and provides examples.
result Reveals the cost of MEV is related to the 'smoothness' of a function over the symmetric group.

Maximal extractable value in CFMMs can degrade or improve routing quality, with reordering MEV showing logarithmic impact.

problem Maximal extractable value in constant function market makers (CFMMs) and its impact on routing quality.
method Game theoretic analysis of MEV in CFMMs, constructing price of anarchy and analyzing reordering MEV.
result Conditions under which reordering MEV shows logarithmic impact, and implications for MEV searchers and CFMM designers.

This paper studies the problem of optimally extracting nonrenewable natural resource in light of various financial and economic restrictions and constraints. Taking into account the fact that the market values of the main natural resources i.e. oil, natural gas, copper,...,etc, fluctuate randomly following global and s…

2016-06-10abs ↗pdf ↗

Feature learning forms the cornerstone for tackling challenging learning problems in domains such as speech, computer vision and natural language processing. In this paper, we consider a novel class of matrix and tensor-valued features, which can be pre-trained using unlabeled samples. We present efficient algorithms f…

2014-12-19abs ↗pdf ↗

A problem of paramount importance in both pure (Restricted Invertibility problem) and applied mathematics (Feature extraction) is the one of selecting a submatrix of a given matrix, such that this submatrix has its smallest singular value above a specified level. Such problems can be addressed using perturbation analys…

2018-04-03abs ↗pdf ↗

Supervised linear feature extraction can be achieved by fitting a reduced rank multivariate model. This paper studies rank penalized and rank constrained vector generalized linear models. From the perspective of thresholding rules, we build a framework for fitting singular value penalized models and use it for feature …

2010-07-19abs ↗pdf ↗

TXtract extracts structured knowledge from thousands of product categories.

problem Extracting structured knowledge from diverse product categories in e-commerce.
method TXtract uses a taxonomy-aware model with category conditional self-attention and multi-task learning.
result TXtract outperforms state-of-the-art approaches by up to 10% in F1 and 15% in coverage across all categories.

Machine learning (ML) models may be deemed confidential due to their sensitive training data, commercial value, or use in security applications. Increasingly often, confidential ML models are being deployed with publicly accessible query interfaces. ML-as-a-service ("predictive analytics") systems are an example: Some …

2016-09-09abs ↗pdf ↗

Charts are an excellent way to convey patterns and trends in data, but they do not facilitate further modeling of the data or close inspection of individual data points. We present a fully automated system for extracting the numerical values of data points from images of scatter plots. We use deep learning techniques t…

2017-04-21abs ↗pdf ↗

Decouples critic chunk length from policy to improve policy reactivity and performance.

problem Bootstrapping bias and difficulty in extracting optimal policies from chunked critics.
method Optimizes policy against a distilled critic for partial action chunks, allowing shorter chunks for policy.
result Reliably outperforms prior methods on long-horizon offline goal-conditioned tasks.

Study optimal auction formats for maximizing MEV on Ethereum.

problem Maximizing extractable value from Ethereum auctions.
method Empirical analysis of 2.2 million transactions, modeling affiliation among bidders.
result English and second-price sealed-bid auctions dominate other formats, with significant revenue losses.

Research examines how strategic latency manipulation impacts Ethereum's network efficiency and decentralization.

problem Impact of artificial latency on Ethereum's network efficiency and decentralization.
method Comprehensive analysis of MEV-Boost auction system and empirical validation with a pilot.
result Increased profitability for node operators and significant systemic challenges like heightened network inefficiencies and centralization risks.

This paper studies a finite-fuel two-dimensional degenerate singular stochastic control problem under regime switching that is motivated by the optimal irreversible extraction problem of an exhaustible commodity. A company extracts a natural resource from a reserve with finite capacity, and sells it in the market at a …

2016-02-22abs ↗pdf ↗

DCoM uses deep neural networks to detect semantic data types from raw column values.

problem Detecting semantic data types from dirty and unseen data.
method DCoM employs multi-input NLP-based deep neural networks trained on 686,765 data columns.
result DCoM outperforms existing methods significantly on 78 different semantic data types.

This paper provides a practical method to extract caplet volatilities from quoted data.

problem Extracting caplet volatilities from quoted data is complex and not straightforward.
method The paper presents a constructive algorithm based on criteria and robust outlier detection. It includes direct interpolation, bootstrap methods, and global search methods.
result The paper introduces methods to extract caplet volatilities that are arbitrage-free and consistent with quoted data.

A price-maker company extracts an exhaustible commodity from a reservoir, and sells it instantaneously in the spot market. In absence of any actions of the company, the commodity's spot price evolves either as a drifted Brownian motion or as an Ornstein-Uhlenbeck process. While extracting, the company affects the marke…

2018-12-04abs ↗pdf ↗

The paper analyzes CEX-DEX arbitrage and profitability on Ethereum, revealing centralization trends and market impacts.

problem Ethereum's decentralization and CEX-DEX arbitrages.
method Empirical analysis of 19 months' data from 7.2M CEX-DEX transactions, refining heuristics to identify and estimate arbitrage revenue.
result Three searchers captured three-quarters of volume and extracted value, and profitability is tied to integration with block builders.

Define quiver representation-valued invariants for classical and virtual knots

problem Define quiver representation-valued invariants for classical and virtual knots
method Define an infinite family of quiver representation-valued invariants of classical and virtual knots associated to a choice of data vector consisting of a biquandle, abelian group, set of biquandle arrows weights with values in the abelian group, coefficient ring and set of biquandle endomorphisms.
result Extract four new polynomial invariants as decategorifications

GAS models have been recently proposed in time-series econometrics as valuable tools for signal extraction and prediction. This paper details how financial risk managers can use GAS models for Value-at-Risk (VaR) prediction using the novel GAS package for R. Details and code snippets for prediction, comparison and back…

2016-11-18abs ↗pdf ↗

Training machine learning (ML) models is expensive in terms of computational power, amounts of labeled data and human expertise. Thus, ML models constitute intellectual property (IP) and business value for their owners. Embedding digital watermarks during model training allows a model owner to later identify their mode…

2019-06-03abs ↗pdf ↗

To understand the relationship between news sentiment and company stock price movements, and to better understand connectivity among companies, we define an algorithm for measuring sentiment-based network risk. The algorithm ranks companies in networks of co-occurrences, and measures sentiment-based risk, by calculatin…

2017-06-19abs ↗pdf ↗

Suppose you have one unit of stock, currently worth 1, which you must sell before time TT. The Optional Sampling Theorem tells us that whatever stopping time we choose to sell, the expected discounted value we get when we sell will be 1. Suppose however that we are able to see aa units of time into the future, and ba…

2016-01-22abs ↗pdf ↗

This article introduces a framework to estimate the value of evidence-based decision making.

problem Lack of empirical tools to assess the value of evidence-based decision making and optimize statistical precision.
method Empirical framework using parametric and nonparametric empirical Bayes methods.
result The value of statistical evidence depends on how organizations translate it into policy decisions.

To identify emerging interdependencies between traded stocks we investigate the behavior of the stocks of FTSE 100 companies in the period 2000-2015, by looking at daily stock values. Exploiting the power of information theoretical measures to extract direct influences between multiple time series, we compute the infor…

2016-11-08abs ↗pdf ↗

Proposes a method for forecasting large-scale interval-valued time series.

problem Modeling and forecasting large-scale interval-valued time series.
method Feature extraction procedure involving auto-segmentation, clustering, and precision matrix estimation.
result The method enhances forecasting performance for large-scale interval-valued time series.

Multivariate time series is a very active topic in the research community and many machine learning tasks are being used in order to extract information from this type of data. However, in real-world problems data has missing values, which may difficult the application of machine learning techniques to extract informat…

2019-03-22abs ↗pdf ↗

Thanks to the rise of wearable and connected devices, sensor-generated time series comprise a large and growing fraction of the world's data. Unfortunately, extracting value from this data can be challenging, since sensors report low-level signals (e.g., acceleration), not the high-level events that are typically of in…

2016-09-29abs ↗pdf ↗

In this study, we propose a new statical approach for high-dimensionality reduction of heterogenous data that limits the curse of dimensionality and deals with missing values. To handle these latter, we propose to use the Random Forest imputation's method. The main purpose here is to extract useful information and so r…

2017-07-02abs ↗pdf ↗

Study examines stylized facts in DEX markets vs. traditional exchanges.

problem Comparing stylized facts in decentralized exchanges (DEXs) vs. traditional markets.
method Empirical analysis of 24 most active Uniswap v3 pools.
result New statistical regularities in DEX markets, linked to market structure and activity.

Extends Tanimoto kernel to real-valued functions.

problem Measuring similarity between real-valued functions.
method Unified representation of real-valued functions via sets, derived general form of the kernel, explicit feature representation, and smooth approximation.
result General Tanimoto kernel for real-valued functions.

An unsupervised anomaly detection method for irregularly sampled time-series data.

problem Anomaly detection in irregularly sampled or missing valued time-series data.
method Uses LSTM networks with time modulation gates to extract temporal features and SVDD for anomaly labeling.
result Significantly outperforms standard approaches on real-life datasets.