The paper reveals the hidden costs of digitizing commodity money and proposes a new stable-coin system.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Training classification models on imbalanced data tends to result in bias towards the majority class. In this paper, we demonstrate how variable discretization and cost-sensitive logistic regression help mitigate this bias on an imbalanced credit scoring dataset, and further show the application of the variable discret…
In this paper, we reveal the attenuation mechanism of anchor of the commodity money from the perspective of logistics warehousing costs, and propose a novel Decayed Commodity Money (DCM) for the store of value across time and space. Considering the logistics cost of commodity warehousing by the third financial institut…
Unified framework for sparse logistic regression with nonconvex regularization.
Small LLMs outperform large ones on simple tasks without extra labelling costs.
New approach improves stock policies for paper companies, reducing waste and costs.
The l1-regularized logistic regression (or sparse logistic regression) is a widely used method for simultaneous classification and feature selection. Although many recent efforts have been devoted to its efficient implementation, its application to high dimensional data still poses significant challenges. In this paper…
Classification is the most important process in data analysis. However, due to the inherent non-convex and non-smooth structure of the zero-one loss function of the classification model, various convex surrogate loss functions such as hinge loss, squared hinge loss, logistic loss, and exponential loss are introduced. T…
The paper analyzes logistic regression for rare events data, deriving new insights on estimator efficiency and sampling strategies.
A method for safe online classification reduces test costs while maintaining low error rates.
Logistic regression is by far the most widely used classifier in real-world applications. In this paper, we benchmark the state-of-the-art active learning methods for logistic regression and discuss and illustrate their underlying characteristics. Experiments are carried out on three synthetic datasets and 44 real-worl…
New research shows logistic regression can achieve optimal error rate for agnostic learning of halfspaces.
Data privacy and security becomes a major concern in building machine learning models from different data providers. Federated learning shows promise by leaving data at providers locally and exchanging encrypted information. This paper studies the vertical federated learning structure for logistic regression where the …
Logistic regression models are a popular and effective method to predict the probability of categorical response data. However inference for these models can become computationally prohibitive for large datasets. Here we adapt ideas from symbolic data analysis to summarise the collection of predictor variables into his…
FOLKLORE algorithm speeds up online multiclass logistic regression.
Improved CRT for sparse logistic regression in high dimensions.
New method predicts customer churn using mixed-penalty logistic regression.
This study aims to predict vessel stay and delay times at ports to optimize logistics.
We consider Online Convex Optimization (OCO) in the setting where the costs are -strongly convex and the online learner pays a switching cost for changing decisions between rounds. We show that the recently proposed Online Balanced Descent (OBD) algorithm is constant competitive in this setting, with competitive rat…
Developed efficient distributed logistic regression for large datasets.
MPNN improves on UniFL approximation with provable guarantees.
New algorithms reduce regret in reinforcement learning with MNL approximations.
A new multi-phase approach improves supply chain forecasting accuracy.
We generated a dataset of 200 GB with 10^9 features, to test our recent b-bit minwise hashing algorithms for training very large-scale logistic regression and SVM. The results confirm our prior work that, compared with the VW hashing algorithm (which has the same variance as random projections), b-bit minwise hashing i…
FisherSFT selects informative examples to fine-tune LLMs efficiently.
ACOWA improves distributed sparse classification with extra communication round.
Bayesian method tackles variable selection in high-dimensional data.
FAB-COST improves cold-start recommendation accuracy with less data.
New method for efficiently deleting data from ML models.
Maximum likelihood estimator performance in logistic regression analyzed.
Tests for Esophageal cancer can be expensive, uncomfortable and can have side effects. For many patients, we can predict non-existence of disease with 100% certainty, just using demographics, lifestyle, and medical history information. Our objective is to devise a general methodology for customizing tests using user pr…
Paper finds a lower bound for estimating low-rank matrices in logistic regression.
Sparsity-constrained optimization has wide applicability in machine learning, statistics, and signal processing problems such as feature selection and compressive Sensing. A vast body of work has studied the sparsity-constrained optimization from theoretical, algorithmic, and application aspects in the context of spars…
Revises logistic-softmax likelihood for Bayesian meta-learning in few-shot classification.
Modern large scale machine learning applications require stochastic optimization algorithms to be implemented on distributed computational architectures. A key bottleneck is the communication overhead for exchanging information such as stochastic gradients among different workers. In this paper, to reduce the communica…
Novel bounds for logistic regression coreset construction and feature selection.
For massive data, the family of subsampling algorithms is popular to downsize the data volume and reduce computational burden. Existing studies focus on approximating the ordinary least squares estimate in linear regression, where statistical leverage scores are often used to define subsampling probabilities. In this p…
Study explores geometric structure and prior for beta-logistic distribution.
We comment on the fact that gradient ascent for logistic regression has a connection with the perceptron learning algorithm. Logistic learning is the "soft" variant of perceptron learning.
We develop unbiased implicit variational inference (UIVI), a method that expands the applicability of variational inference by defining an expressive variational family. UIVI considers an implicit variational distribution obtained in a hierarchical manner using a simple reparameterizable distribution whose variational …
Trans-GCR uses GCR model for node classification, providing theoretical guarantees and superior performance.
Develops a tool to identify abnormal blood smear results based on CBC tests.
AUC (area under ROC curve) is an important evaluation criterion, which has been popularly used in many learning tasks such as class-imbalance learning, cost-sensitive learning, learning to rank, etc. Many learning approaches try to optimize AUC, while owing to the non-convexity and discontinuousness of AUC, almost all …
Paper explains learning property of logistic and softmax losses for balanced and imbalanced class data.
Improved sketching for logistic and regression with near-linear dimensions.
Safe screening rules reduce computation time in logistic regression with regularization.
Synthetic lethality (SL) is a promising concept for novel discovery of anti-cancer drug targets. However, wet-lab experiments for detecting SLs are faced with various challenges, such as high cost, low consistency across platforms or cell lines. Therefore, computational prediction methods are needed to address these is…
Proposes logistic-beta process for modeling dependent probabilities with beta marginals.