The paper reveals the hidden costs of digitizing commodity money and proposes a new stable-coin system.
problem Depreciation of banknotes due to high logistics costs after digitization.
method Analyzing the functions of money from a logistics perspective and comparing commodity money to digital currency.
result There is no honest money that is both a store of value and has negligible logistics costs.
The paper shows how variable discretization and cost-sensitive logistic regression improve credit scoring models on imbalanced data.
problem Bias in classification models on imbalanced datasets.
method Variable discretization and cost-sensitive logistic regression.
result Improves model performance on imbalanced credit scoring data and other domains.
A new currency DCM reduces logistics costs and preserves wealth.
problem High logistics costs associated with commodity money.
method Introducing Decayed Commodity Money (DCM) with an attenuation coefficient based on logistics costs.
result DCM offers a cost-effective and wealth-preserving alternative to traditional currency.
Unified framework for sparse logistic regression with nonconvex regularization.
problem Sparse logistic regression with nonconvex regularization.
method Unified framework, line search criteria for nonconvex terms.
result Effective classification and feature selection at lower computational cost.
Small LLMs outperform large ones on simple tasks without extra labelling costs.
problem Performance of large commercial models in simple classification tasks.
method Logistic Regression on small LLM embeddings.
result Small LLMs equal or outperform large LLMs in 'tens-of-shot' classification tasks.
New approach improves stock policies for paper companies, reducing waste and costs.
problem Improving stock policies for integrated paper companies.
method Developed a new approach to determine near-optimal stock policies.
result Reduction in total waste by 9% and logistics costs.
The l1-regularized logistic regression (or sparse logistic regression) is a widely used method for simultaneous classification and feature selection. Although many recent efforts have been devoted to its efficient implementation, its application to high dimensional data still poses significant challenges. In this paper…
Paper proposes a faster federated learning method for logistic regression.
problem Federated learning for logistic regression with disjoint feature sets.
method Quasi-Newton method under additively homomorphic encryption.
result Significant reduction in communication rounds with minimal additional cost.
Efficient logistic regression for aggregated data reduces computation time.
problem Inference for logistic regression models is computationally expensive for large datasets.
method Adapted symbolic data analysis to summarise predictor variables into histograms and use composite likelihoods.
result The method achieves comparable classification rates to full data analysis but at lower computational cost.
Logitron combines Perceptron and logistic loss for improved classification.
problem Non-convex and non-smooth zero-one loss function in classification models.
method Introduces a Perceptron-augmented convex classification framework with an extended logistic loss function.
result Hinge-Logitron outperforms logistic regression and SVM in classification accuracy.
This paper benchmarks active learning methods for logistic regression and finds uncertainty sampling performs well.
problem Benchmarking and comparing active learning methods for logistic regression.
method State-of-the-art active learning methods for logistic regression were benchmarked and compared.
result Uncertainty sampling performs exceptionally well overall.
The paper analyzes logistic regression for rare events data, deriving new insights on estimator efficiency and sampling strategies.
problem Binary logistic regression for rare events data with significantly fewer events than controls.
method Derives asymptotic distribution of MLE, proves under-sampling advantage, and compares over-sampling efficiency.
result Under-sampling a small proportion of nonevents can improve efficiency in rare events data analysis.
A method for safe online classification reduces test costs while maintaining low error rates.
problem Sequential testing for binary disease outcomes with unknown logistic model parameters.
method Joint estimation of logistic parameter and feature distribution with a conservative threshold.
result Achieves target error with high probability and requires minimal excess tests.
New research shows logistic regression can achieve optimal error rate for agnostic learning of halfspaces.
problem Agnostic learning of homogeneous halfspaces with logistic loss.
method Constructing a well-behaved distribution and using logistic regression with additional convex optimization steps.
result Logistic regression can achieve Ω ( e x t r m O P T ) Ω(\sqrt{ extrm{OPT}}) Ω ( e x t r m O P T ) misclassification risk, matching the upper bound. FOLKLORE algorithm speeds up online multiclass logistic regression.
problem Efficiently solving online multiclass logistic regression without high computational cost.
method Developed FOLKLORE algorithm with improved runtime and regret bound.
result First practical algorithm for online multiclass logistic regression.
Improved CRT for sparse logistic regression in high dimensions.
problem Accurate inference in high-dimensional sparse logistic regression.
method Variable-distillation and decorrelation steps in CRT-logit.
result CRT-logit provides a more powerful solution with theoretical guarantees.
New method predicts customer churn using mixed-penalty logistic regression.
problem Predicting customer churn in CRM systems.
method Mixed-penalty logistic regression for big data analysis.
result Proposed method enhances logistic regression for better predictive analytics.
Proposes efficient subsampling for logistic regression with optimal probabilities.
problem Efficiently approximating maximum likelihood estimate in logistic regression for large datasets.
method Develops subsampling algorithms for logistic regression, derives optimal subsampling probabilities, and proposes two-step approximation.
result Optimal subsampling reduces computing time significantly while maintaining estimator consistency and normality.
SL2MF predicts synthetic lethality using logistic matrix factorization.
problem Predicting synthetic lethality in human cancers from limited experimental data.
method Logistic matrix factorization incorporating biological knowledge.
result SL2MF effectively predicts known and unknown SL interactions.
This study aims to predict vessel stay and delay times at ports to optimize logistics.
problem Uncertainties in maritime logistics, including weather, cargo diversity, and port dynamics, lead to increased costs and inefficiencies.
method Developed predictive analytics to address shortcomings in previous works, using feature analysis and SHAP explanations.
result Predictive analytics can assist in efficient planning and scheduling of port operations, reducing costs and improving logistics.
OBD algorithm optimizes online convex optimization with strong convexity and switching costs.
problem Online convex optimization with strong convexity and switching costs.
method Online Balanced Descent (OBD) algorithm for m m m -strongly convex costs with near-optimal dynamic regret and per-round accuracy for ε ε ε -smooth sequences. result OBD achieves a competitive ratio of 3 + O ( 1 / m ) 3 + O(1/m) 3 + O ( 1/ m ) for m m m -strongly convex costs. Probability calibration trees improve accuracy of probability estimates.
problem Improving accuracy and calibration of probability estimates from classifiers.
method Probability calibration trees modify logistic model trees to learn different models in regions of the input space.
result Probability calibration trees outperform isotonic regression and Platt scaling in terms of root mean squared error.
New method learns optimal cost for machine learning models.
problem Optimizing machine learning models under distributional uncertainty.
method Data-driven approach to define distributional uncertainty neighborhoods.
result Improves upon various machine learning estimators.
Proposes a new loss function for distributional learning.
problem Learning sparse and singular distributions.
method Entropy-regularized optimal transport and Fenchel duality.
result Geometric loss results in unconstrained convex objective functions.
Developed efficient distributed logistic regression for large datasets.
problem Communication inefficiency and sparsity issues in distributed training of large-scale models.
method Iterative local optimization of a surrogate likelihood to improve initial solutions, handling sparsity and diverging updates.
result Learned a communication-efficient distributed logistic regression model for millions of features.
The paper finds active learning is helpful when it reduces error rate.
problem Understanding when active learning improves model performance.
method Empirical study on 21 datasets with logistic regression and uncertainty sampling.
result There is a strong inverse correlation between data efficiency and error rate.
MPNN improves on UniFL approximation with provable guarantees.
problem Uniform Facility Location (UniFL) optimization problem.
method Graph Neural Network (MPNN) incorporating approximation-algorithmic principles.
result Empirically outperforms standard approximation algorithms.
Serverless runtimes boost large-scale optimization efficiency.
problem Efficiently solving large-scale optimization problems.
method Master-worker setup with AWS Lambda, parallel optimization algorithm.
result Relative speedups up to 256 workers and efficiencies above 70% up to 64 workers.
New algorithms reduce regret in reinforcement learning with MNL approximations.
problem Efficient reinforcement learning with MNL function approximation for MDPs.
method Proposed randomized exploration algorithms with frequentist regret guarantees.
result Achieved improved regret bounds for MNL transition models.
A new multi-phase approach improves supply chain forecasting accuracy.
problem Improving forecast accuracy for hierarchical supply chain demands.
method Independent child-level forecasting followed by parent-level estimation.
result 82-90% improvement in forecast accuracy compared to traditional methods.
We generated a dataset of 200 GB with 10^9 features, to test our recent b-bit minwise hashing algorithms for training very large-scale logistic regression and SVM. The results confirm our prior work that, compared with the VW hashing algorithm (which has the same variance as random projections), b-bit minwise hashing i…
The paper proposes gradient sparsification to reduce communication costs in distributed optimization.
problem Reduction of communication overhead in distributed machine learning.
method Formulates a convex optimization problem to minimize gradient coding length, and proposes simple algorithms for approximate solution.
result The proposed sparsification techniques significantly reduce communication costs without sacrificing accuracy.
FisherSFT selects informative examples to fine-tune LLMs efficiently.
problem Adapting large language models to new domains efficiently.
method Selects examples maximizing information gain using Hessian of log-likelihood.
result Empirically demonstrates improved performance with reduced computational cost.
ACOWA improves distributed sparse classification with extra communication round.
problem Efficiently optimizing sparse classification with limited communication.
method Introducing ACOWA, a new technique with an extra communication round.
result ACOWA achieves better approximation quality and higher accuracy.
Bayesian method tackles variable selection in high-dimensional data.
problem Challenges in Bayesian variable selection with large P.
method Efficient MCMC scheme with sublinear cost per iteration, extended to generalized linear models.
result Demonstrated effectiveness on cancer and maize genomic data.
Logistic regression connects to perceptron learning via gradient ascent.
problem No specific problem stated; focuses on connection between algorithms.
method Gradient ascent for logistic regression compared to perceptron learning.
result Gradient ascent for logistic regression is a soft variant of perceptron learning.
FAB-COST improves cold-start recommendation accuracy with less data.
problem Cold-start problem in recommendation systems.
method Contextual bandit algorithm using Expectation Propagation and Assumed Density Filtering.
result FAB-COST outperforms Laplace approximation on real data.
New method for efficiently deleting data from ML models.
problem Efficiently removing data from trained ML models without retraining.
method Approximate deletion method for linear and logistic models.
result Significantly faster than existing methods, with linear time dependence on feature dimension.
Sparsity-constrained optimization has wide applicability in machine learning, statistics, and signal processing problems such as feature selection and compressive Sensing. A vast body of work has studied the sparsity-constrained optimization from theoretical, algorithmic, and application aspects in the context of spars…
Maximum likelihood estimator performance in logistic regression analyzed.
problem Performance of maximum likelihood estimator in logistic regression.
method Sharp non-asymptotic guarantees for existence and excess logistic risk.
result Sharp guarantees for the existence and excess risk of MLE in logistic regression.
Paper finds a lower bound for estimating low-rank matrices in logistic regression.
problem Estimating low-rank coefficient matrices in logistic regression.
method Derives a minimax lower bound on the risk.
result The bound depends on matrix dimensions, rank, and sample size.
Revises logistic-softmax likelihood for Bayesian meta-learning in few-shot classification.
problem Inherent uncertainty in logistic-softmax leads to suboptimal performance in meta-learning.
method Redesigns logistic-softmax likelihood with a temperature parameter for better control of prior confidence.
result Achieves well-calibrated uncertainty estimates and comparable/superior performance on benchmark datasets.
Novel bounds for logistic regression coreset construction and feature selection.
problem Efficiently summarize and reduce logistic regression inputs.
method Feature space sketching for logistic regression.
result Tight bounds for coreset construction and feature selection.
Customizes esophageal cancer tests based on patient preferences.
problem Expensive and uncomfortable esophageal cancer tests.
method Classifiers trained from EHRs for test selection, with a focus on minimizing false abnormals.
result 99.8% accuracy in test selection using kernel SVM and LR, without MP features.
Trans-GCR uses GCR model for node classification, providing theoretical guarantees and superior performance.
problem Challenges in obtaining node classification labels in real-world scenarios.
method Graph Convolutional Multinomial Logistic Regression (GCR) model and transfer learning method based on GCR.
result Trans-GCR provides superior empirical performance and theoretical guarantees.
Study explores geometric structure and prior for beta-logistic distribution.
problem Understanding the geometric structure and prior distributions of the beta-logistic distribution.
method Exploring dual geometric structure and uncovering α \alpha α -parallel prior. result The beta-logistic distribution admits an α \alpha α -parallel prior for any real number α \alpha α . Improves logistic regression performance on imbalanced data.
problem Imbalanced data leads to all labels being estimated as majority class.
method Uses F-measure optimization to estimate relative density ratio and approximate relative F-measure.
result Proposed method improves logistic regression performance on imbalanced data.
Develops a tool to identify abnormal blood smear results based on CBC tests.
problem Manual review of blood smears by technologists is time-consuming and inconsistent.
method Cost-sensitive Lasso-penalized additive logistic regression combined with stability selection.
result The tool correctly identifies true cutoff values for abnormal smear results.