Extends clustering method to cost-based hierarchies.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper tackles regression with cost-based rejection, balancing prediction and rejection costs.
A new active learning method considers both uncertainty and diversity to minimize labeling and decision costs.
In this paper we address cardinality estimation problem which is an important subproblem in query optimization. Query optimization is a part of every relational DBMS responsible for finding the best way of the execution for the given query. These ways are called plans. The execution time of different plans may differ b…
Sparse RSP routing improves graph exploration and classification.
A method for pricing and superhedging European options under proportional transaction costs based on linear vector optimisation and geometric duality developed by Lohne & Rudloff (2014) is compared to a special case of the algorithms for American type derivatives due to Roux & Zastawniak (2014). An equivalence between …
For several decades, the no-arbitrage (NA) condition and the martingale measures have played a major role in the financial asset's pricing theory. We propose a new approach for estimating the super-replication cost based on convex duality instead of martingale measures duality: Our prices will be expressed using Fenche…
In this paper we present a theoretical framework for determining dynamic ask and bid prices of derivatives using the theory of dynamic coherent acceptability indices in discrete time. We prove a version of the First Fundamental Theorem of Asset Pricing using the dynamic coherent risk measures. We introduce the dynamic …
Recent attempts to achieve fairness in predictive models focus on the balance between fairness and accuracy. In sensitive applications such as healthcare or criminal justice, this trade-off is often undesirable as any increase in prediction error could have devastating consequences. In this work, we argue that the fair…
New method for pricing financial products without no-arbitrage condition.
We propose a family of relaxations of the optimal transport problem which regularize the problem by introducing an additional minimization step over a small region around one of the underlying transporting measures. The type of regularization that we obtain is related to smoothing techniques studied in the optimization…
Traditionally, machine learning algorithms rely on the assumption that all features of a given dataset are available for free. However, there are many concerns such as monetary data collection costs, patient discomfort in medical procedures, and privacy impacts of data collection that require careful consideration in a…
This study compares various superlearner and deep learning architectures (machine-learning-based and neural-network-based) for classification problems across several simulated and industrial datasets to assess performance and computational efficiency, as both methods have nice theoretical convergence properties. Superl…
New geometry for optimal transport cost based on Bregman divergences.
The challenge of efficiently identifying anomalies in data sequences is an important statistical problem that now arises in many applications. Whilst there has been substantial work aimed at making statistical analyses robust to outliers, or point anomalies, there has been much less work on detecting anomalous segments…
There are many industrial situations where rods are used to stir a fluid, or where rods repeatedly stretch a material such as bread dough or taffy. The goal in these applications is to stretch either material lines (in a fluid) or the material itself (for dough or taffy) as rapidly as possible. The growth rate of mater…
Machine learning has automated much of financial fraud detection, notifying firms of, or even blocking, questionable transactions instantly. However, data imbalance starves traditionally trained models of the content necessary to detect fraud. This study examines three separate factors of credit card fraud detection vi…
Cost-effective feature selection improves network model choice.
A framework previously introduced in [3] for solving a sequence of stochastic optimization problems with bounded changes in the minimizers is extended and applied to machine learning problems such as regression and classification. The stochastic optimization problems arising in these machine learning problems is solved…
New algorithms optimize time series classification speed and accuracy.
Machine learning predicts Bitcoin returns but trading performance drops with costs.
We propose a simple yet effective technique to simplify the training and the resulting model of neural networks. In back propagation, only a small subset of the full gradient is computed to update the model parameters. The gradient vectors are sparsified in such a way that only the top-k elements (in terms of magnitude…
Paper studies fundamental limits of communication in distributed learning.
Study optimizes stock portfolios using network analysis and forecasting.
Imbalanced data with a skewed class distribution are common in many real-world applications. Deep Belief Network (DBN) is a machine learning technique that is effective in classification tasks. However, conventional DBN does not work well for imbalanced data classification because it assumes equal costs for each class.…
Recent advances in neural networks have inspired people to design hybrid recommendation algorithms that can incorporate both (1) user-item interaction information and (2) content information including image, audio, and text. Despite their promising results, neural network-based recommendation algorithms pose extensive …
Graphlets are defined as k-node connected induced subgraph patterns. For an undirected graph, 3-node graphlets include close triangle and open triangle. When k = 4, there are six types of graphlets, e.g., tailed-triangle and clique are two possible 4-node graphlets. The number of each graphlet, called graphlet count, i…
Internet companies are facing the need for handling large-scale machine learning applications on a daily basis and distributed implementation of machine learning algorithms which can handle extra-large scale tasks with great performance is widely needed. Deep forest is a recently proposed deep learning framework which …
Active learning suffers from biased non-response, which this paper addresses.
Develops a framework for valuing Asian options with market impact.
Study on newsvendor problem with censored data, showing how much information is lost.
Quantum algorithm speeds up MIP solving by a near-quadratic factor.