This paper shows neural networks with logistic activations can be trained effectively.
problem Training feedforward neural networks with standard logistic activations is difficult.
method A set of conditions for parameter initialization derived from information-theoretic studies.
result The proposed initialization leads to networks achieving comparable generalization performance to those using hyperbolic tangent activations.
This paper benchmarks active learning methods for logistic regression and finds uncertainty sampling performs well.
problem Benchmarking and comparing active learning methods for logistic regression.
method State-of-the-art active learning methods for logistic regression were benchmarked and compared.
result Uncertainty sampling performs exceptionally well overall.
A greedy active learning algorithm for logistic regression reduces model size and training size.
problem Binary classification with reduced model size and training size.
method Modified batch subject selection strategy with greedy variable selection.
result Competitive performance with smaller training size and model size.
This paper provides an overview of activation functions in neural networks.
problem Confusion in activation function selection and properties in deep learning.
method Analytic review of popular activation functions.
result Clarification of activation function properties and selection.
A fast method for Lasso and Logistic Lasso problems.
problem Solving Lasso and Logistic Lasso regression problems efficiently.
method Iterative active set approach using solver updates.
result 31.41 times faster on average for compressed sensing.
Gradient descent with logistic loss can interpolate deep networks with smoothed ReLU activations under certain conditions.
problem Conditions for gradient descent to drive logistic loss to zero in deep networks with smoothed ReLU activations.
method Gradient descent applied to fixed-width deep networks with smoothed ReLU approximations (e.g., Swish, Huberized ReLU).
result Gradient descent can drive logistic loss to zero under specific conditions, providing bounds on convergence rate.
SGD converges globally to logistic loss minima for two-layer nets.
problem Global convergence of SGD for logistic loss on two-layer neural nets.
method Demonstrates existence of Frobenius norm regularized logistic loss functions as Villani functions, proving convergence and exponential rate.
result SGD converges globally to the global minima of appropriately regularized logistic empirical risk of depth 2 nets.
The paper finds active learning is helpful when it reduces error rate.
problem Understanding when active learning improves model performance.
method Empirical study on 21 datasets with logistic regression and uncertainty sampling.
result There is a strong inverse correlation between data efficiency and error rate.
A new method improves active learning for large batch sizes.
problem Challenges in scaling Bayesian active learning to large batch sizes.
method Derives Partial Batch Label Sampling (ParBaLS) for EPIG algorithm.
result ParBaLS EPIG outperforms top-B selection and BatchBALD. New algorithm for active learning from feedback coding.
problem Efficiently selecting examples for labeling in active machine learning.
method Formalized structural similarities between active learning and feedback channel coding, developed an optimal transport-based algorithm called Approximate Posterior Matching (APM).
result Learning performance comparable to existing methods at reduced computational cost.
Model predicts individual insurance claim reserves using activation patterns.
problem Accurately predicting individual claim reserves in insurance contracts.
method Multinomial logistic regression to model claim activation and development.
result The model generates accurate predictions of total and per coverage reserves.
Study compares forecasting methods for logistics time series.
problem Improving forecasting accuracy in logistics.
method Compared statistical and machine learning methods on simulated time series.
result Statistical methods outperformed machine learning in one-step forecasts.
Improves node classification in graphs with active learning.
problem Difficult or expensive labeling in node classification tasks.
method Graph cognizant logistic regression and preemptive query generation.
result Significant improvement over state-of-the-art approaches.
FIRAL is a scalable active learning algorithm for multiclass classification.
problem Scalability issues with FIRAL in large datasets.
method Proposed an approximate algorithm with reduced storage and computational complexity.
result Demonstrated strong scalability and accuracy on large datasets.
YASENN interprets neural networks by partitioning activation sequences.
problem Interpreting complex neural network decisions.
method YASENN uses layer-wise gradient boosting decision trees to distill and partition neuron activation sequences.
result YASENN provides interpretable partitions of the input space, revealing neural network decision artifacts.
Spatial metamodels improve root finding in uncertain oracle responses.
problem Finding roots in noisy, location-dependent oracle responses.
method Propose spatial metamodels to infer oracle distribution and update Bayesian knowledge.
result Spatial PBA algorithm outperforms earlier models in synthetic and real-world problems.
Proposes a new active learning criterion to maximize classifier instability.
problem Efficiently train classifiers with minimal labeled data.
method Maximizes variance of output changes for unlabeled data.
result Achieves state-of-the-art performance in experiments.
Deep learning method improves regression accuracy.
problem Nonparametric regression challenges.
method Over-parametrized deep neural networks with logistic activation, gradient descent, special topology, random initialization, and data-dependent learning rate.
result Theoretical bound on L2 error and improved finite sample performance. With the rapid development of social media sharing, people often need to manage the growing volume of multimedia data such as large scale video classification and annotation, especially to organize those videos containing human activities. Recently, manifold regularized semi-supervised learning (SSL), which explores th…
New activation functions mimic neuronal biology to improve deep learning performance.
problem Vanishing gradients and suboptimal learning in deep learning models.
method Introducing bionodal root unit (BRU) activation functions based on neuronal cell properties.
result BRU activation functions lead to faster training and better generalization in deep learning models.
Anomaly detection model for large networks identifies attacks with reduced false positives.
problem Detecting anomalies in large, sparse, directed networks.
method Dynamic logistic model with latent factors, variational Bayesian estimation, case-control approximation.
result Model reduces false positives by identifying red team attack with half the detection rate.
The paper proposes a method to automatically discover effective spatial filters for hyperspectral image classification.
problem Discovering an effective set of spatial filters for hyperspectral image classification problems.
method An active set feature learner that includes only features improving the classifier, using multiclass logistic classification.
result A simple classifier can reach state-of-the-art performance with learned filters.
Adversarial training makes logistic regression weight loss landscapes sharper.
problem Understanding why adversarial training sharpens the weight loss landscape in logistic regression.
method Theoretical analysis of linear logistic regression model with L2 norm constraints, and experiments on ResNet18.
result Adversarial training sharpens the weight loss landscape in linear logistic regression models.
Analysis of gradient descent on wide neural networks reveals strong generalization.
problem Understanding why wide neural networks trained with logistic loss perform well.
method Characterization of gradient flow limits and comparison to max-margin classifier.
result Margin is independent of ambient dimension, leading to strong generalization.
A new method for Bayesian neural networks using probabilistic backpropagation.
problem Approximating posterior distributions in Bayesian neural networks.
method Variational Expectation Propagation (VEP) with probabilistic backpropagation.
result Efficient algorithm for approximate integration over posterior distributions.
Proposes active sampling for improving fairness in machine learning.
problem Improving fairness in machine learning models, especially for disadvantaged groups.
method Simple active sampling and reweighting strategies for min-max fairness.
result Proves the rate of convergence to a min-max fair solution for convex problems.
The paper shows over-confidence in models isn't just due to over-parametrization.
problem Over-confidence in machine learning models, especially in binary classification.
method Theoretical analysis of logistic regression and other binary classification problems.
result Logistic regression is inherently over-confident in certain settings, but over-confidence is not always the case.
A robust multiclass logistic regression method using Tsallis divergence.
problem Noise robustness in multiclass logistic regression.
method Two-temperature logistic regression with Tsallis divergence.
result Significant robustness to outliers and noise.
This paper models how features influence event triggers in high-dimensional networks.
problem Estimating context-dependent networks in high-dimensional marked point processes.
method Leveraging compositional time series and regularization methods, the paper considers autoregressive multinomial and logistic-normal models for network estimation.
result The logistic-normal model leads to a convex negative log-likelihood objective and captures dependence across categories.
Study fractal and regular geometry in deep neural networks.
problem Investigate geometric properties of neural networks.
method Analyze boundary volumes of excursion sets for different activations.
result Hausdorff dimension increases with depth for non-regular activations.
Social media enhances or diminishes scientific status, depending on usage.
problem Impact of social media on scientific stratification and mobility.
method Logistic Attribution Analysis combining statistical and machine learning methods.
result Social media promotes stratification and mobility, but beyond a threshold, it negatively impacts status.
We extend neural networks with fractional and mixed activation functions for better function approximation.
problem Limitations in approximating higher-order smooth functions in complex spaces.
method Incorporating fractional exponents in activation functions and defining new density functions.
result Improved accuracy and broader applicability of neural network approximation theory.
We propose a novel general algorithm LHAC that efficiently uses second-order information to train a class of large-scale l1-regularized problems. Our method executes cheap iterations while achieving fast local convergence rate by exploiting the special structure of a low-rank matrix, constructed via quasi-Newton approx…
Dropout training improves neural networks' performance.
problem Improving neural network convergence and generalization.
method Two-layer neural networks with ReLU activations, overparametrization, and positive margin assumption.
result Dropout training achieves ε-suboptimality in test error in O(1/ε) iterations.
New method learns SIMs with arbitrary monotone activations without strong distributional assumptions.
problem Learning Single-Index Models with arbitrary monotone activations.
method Based on omniprediction with calibrated multiaccuracy and Bregman divergences.
result First agnostic learning result for SIMs with arbitrary monotone activations.
We consider sequential decision making problems for binary classification scenario in which the learner takes an active role in repeatedly selecting samples from the action pool and receives the binary label of the selected alternatives. Our problem is motivated by applications where observations are time consuming and…
Active learning improves ordering of items with contextual attributes.
problem Learning accurate item orderings from pairwise comparisons, especially when exhaustive comparisons are impractical.
method Proposes an active learning strategy that samples items to minimize expected ordering error, accounting for uncertainty in comparisons.
result Superior sample efficiency and generalization compared to non-contextual ranking approaches and active preference learning baselines.
A new method uses active learning to improve bile duct stone evaluation.
problem Efficiently collecting necessary patient data in sequential healthcare decisions.
method Developed an active learning-based multistage sequential decision-making model.
result Improves estimation efficiency by 62%-1838% compared to baseline methods.
A Qini-based uplift model improves retention marketing campaign performance.
problem Isolating the marketing effect of a campaign and identifying responsive customers.
method Qini-based uplift regression model using logistic regression.
result Qini-optimized uplift models improve performance and provide interpretable models.
Picasso is a new library for sparse learning problems in R and Python.
problem Sparse learning problems in high-dimensional data analysis.
method Unified framework of pathwise coordinate optimization with efficient active set selection strategies.
result picasso can efficiently handle large-scale problems.
High-dimensional feature selection arises in many areas of modern science. For example, in genomic research we want to find the genes that can be used to separate tissues of different classes (e.g. cancer and normal) from tens of thousands of genes that are active (expressed) in certain tissue cells. To this end, we wi…
Gradient descent in neural networks maximizes margin.
problem Optimizing neural networks using gradient descent.
method Gradient descent or gradient flow on homogeneous neural networks.
result Normalized margin increases over time if training loss decreases below a threshold.
Neural networks encode inputs deterministically and categorically, behaving like hash encoders.
problem Understanding the encoding properties of neural networks.
method Analyzed the input space partitioned by ReLU-like activations in neural networks.
result Neural networks can be represented by unique activation patterns, similar to hash encoders.
An active learner is given a class of models, a large set of unlabeled examples, and the ability to interactively query labels of a subset of these examples; the goal of the learner is to learn a model in the class that fits the data well. Previous theoretical work has rigorously characterized label complexity of activ…
New lower bounds improve logistic log-likelihood optimization and inference.
problem Designing computationally tractable lower bounds for logistic log-likelihoods.
method Developed a piece-wise quadratic lower bound that uniformly improves tangent quadratic minorizers.
result Improves the speed of convergence and accuracy of variational Bayes approximations.
Study finds farmers are willing to pay higher premiums for higher coverage in agricultural insurance.
problem Determining the demand factors and WTP for agricultural insurance.
method Conducted a survey of 200 farmers to analyze the impact of socio-demographic variables and premium on insurance purchase decisions.
result Farmers are willing to pay higher premiums for higher coverage in agricultural insurance.
Model stores many more patterns than neurons, improving pattern recognition.
problem Storing and retrieving many more patterns than neurons in a network.
method Constructs a family of models interpolating between feature-matching and prototype modes, corresponding to neural networks with various activation functions.
result Higher rectified polynomials can be used in neural networks for improved pattern recognition.
The problem of human activity recognition is central for understanding and predicting the human behavior, in particular in a prospective of assistive services to humans, such as health monitoring, well being, security, etc. There is therefore a growing need to build accurate models which can take into account the varia…