Paper proposes an active multi-step TD algorithm for reinforcement learning.
problem Challenging decision making and control tasks in reinforcement learning.
method Active stepsize learning and adaptive multi-step TD algorithm with context-aware mechanism.
result Competitive results compared to other reinforcement learning baselines on discrete and continuous space tasks.
Framework for active and adaptive ML solving changing problems.
problem Solving changing machine learning problems over time.
method Active querying of informative samples and adaptation to changes.
result Near-optimal excess risk performance for maximum likelihood estimation.
New method reduces PDE surrogate model training costs by selectively acquiring time steps.
problem High computational cost of generating training data for PDE surrogate models.
method STAP (Selective Time-Step Acquisition for PDEs) framework that acquires only important time steps.
result Demonstrated effectiveness on several benchmark PDEs, reducing training costs.
Proposes PKM for soft K-means clustering.
problem Soft K-means (m=1) unsolved since 1981.
method Probabilistic K-Means (PKM) via nonlinear programming.
result Proposed methods solve PKM efficiently.
Paper uses RL to optimize daily step distribution for better health biomarkers.
problem Lack of personalized PA distribution recommendations for health biomarkers.
method Developed an offline reinforcement learning algorithm to learn optimal PA distributions.
result Learned optimal policy suggests more consistent daily steps and tailored recommendations.
Adaptive smoothing in fMRI improves brain activity analysis.
problem Optimizing spatial smoothing in fMRI data processing pipelines.
method Integrating adaptive spatial smoothing as a neural network layer.
result Adaptive smoothing enhances brain activity analysis accuracy.
Study fractal and regular geometry in deep neural networks.
problem Investigate geometric properties of neural networks.
method Analyze boundary volumes of excursion sets for different activations.
result Hausdorff dimension increases with depth for non-regular activations.
RAN model recognizes multiple activities from unlabeled sensor data.
problem Handling weakly labeled multi-activity data from wearable sensors.
method Recurrent Attention Networks (RAN) for sequential multi-activity recognition and localization.
result RAN model can infer multiple activities and determine activity locations from unlabeled data.
A new algorithm improves SSC clustering accuracy with low complexity.
problem Sparse Subspace Clustering accuracy loss in time efficiency.
method Active Orthogonal Matching Pursuit (Active OMP-SSC) for improved clustering accuracy.
result Improves clustering accuracy of OMP-SSC with low computational complexity.
GOLS finds activation functions affect training robustness, especially ReLU.
problem Investigate how different activation functions impact GOLS in neural network training.
method Identify SNN-GPPs for GOLS, analyze activation function effects on gradient continuity.
result GOLS robust for most activation functions but sensitive to ReLU.
Study learning from multiple thinkers providing step-by-step solutions to problems.
problem Learning from multiple, possibly different, thinkers providing step-by-step solutions to problems.
method Active learning algorithm that uses CoT data from multiple thinkers and end-result data.
result Learning can be hard from CoT supervision provided by two or a few different thinkers, but a generic algorithm can learn efficiently.
Most prior work on active learning of classifiers has focused on sequentially selecting one unlabeled example at a time to be labeled in order to reduce the overall labeling effort. In many scenarios, however, it is desirable to label an entire batch of examples at once, for example, when labels can be acquired in para…
In this paper we present a new method for motion tracking of tumors in liver ultrasound image sequences. Our algorithm has two main steps. In the first step, we apply mean shift algorithm with multiple features to estimate the center of the target in each frame. Target in the first frame is defined using an ellipse. Ed…
Improves neural network performance by normalizing activation functions.
problem Improving convergence speed and robustness of neural networks.
method Transforming existing activation functions into ones with better properties.
result Significantly promotes convergence robustness, maximum training depth, and anytime performance.
Batch normalization improves deep learning by enabling larger learning rates.
problem Improving accuracy and speeding up training in deep neural networks.
method Empirical experiments and analysis of gradient and activation behavior.
result Batch normalization primarily enables training with larger learning rates, leading to faster convergence and better generalization.
Collaborative filtering is a useful technique for exploiting the preference patterns of a group of users to predict the utility of items for the active user. In general, the performance of collaborative filtering depends on the number of rated examples given by the active user. The more the number of rated examples giv…
Improved solver maintains positivity and accuracy across all time steps.
problem Linear second-order schemes for Fokker-Planck equation cannot preserve positivity.
method Flux-Corrected Diagonal Frog (FCDF) framework using nonlinear extension and iterative limiter.
result FCDF schemes are unconditionally positive across all time steps and maintain second-order accuracy.
Proves existence of optimal shallow neural networks with ReLU activation.
problem Proving the existence of optimal shallow feedforward networks with ReLU activation.
method Proves existence of global minima in the loss landscape for continuous target functions using shallow feedforward neural networks with ReLU activation.
result Existence of global minima in the loss landscape for shallow feedforward networks with ReLU activation.
Automatically discovers effective activation functions for deep learning.
problem Inconsistent performance of novel activation functions in deep learning networks.
method Evolutionary search for general form, gradient descent for parameters.
result Significant performance improvements over ReLU and other functions.
Two-step model estimates DLMO using both daily and frequent data.
problem Expensive and time-consuming DLMO measurement.
method Two-step framework combining daily and frequent data.
result Two-step model with two time-scale features has lower errors.
Method infers dynamics from incomplete time series data.
problem Challenges in inferring stochastic dynamics from time series with missing data.
method Expectation Maximization (EM) algorithm that iterates between E-step and M-step.
result The EM algorithm effectively recovers missing data points and infers underlying network models from real neuronal activities.
Dynamic ensemble active learning tackles non-stationary criteria in active learning.
problem Active learning's effectiveness varies across datasets and sessions, leading to suboptimal results.
method Developed a dynamic ensemble active learner based on a non-stationary multi-armed bandit with expert advice.
result Dynamic ensemble selects the best criteria at each step, improving overall performance.
D-CSC framework reveals how ReLU activation functions recover activation paths in neural networks.
problem Understanding how ReLU activation functions recover activation paths in neural networks.
method Deep Convolutional Sparse Coding (D-CSC) framework, omitting dictionary learning, to analyze activation paths.
result Uniform guarantees for recovery of true activation paths with high probability for greater activation densities.
Paper tackles black-box machine teaching with cross-space models, proposing an active teacher model.
problem Teaching a learner with different feature representations and without full observation.
method Proposes an active teacher model that queries the learner to estimate its status and guide faster convergence.
result Active teacher model achieves faster convergence rate than traditional passive learning.
We develop a primal dual active set with continuation algorithm for solving the \ell^0-regularized least-squares problem that frequently arises in compressed sensing. The algorithm couples the the primal dual active set method with a continuation strategy on the regularization parameter. At each inner iteration, it fir…
In this work we investigate intra-day patterns of activity on a population of 7,261 users of mobile health wearable devices and apps. We show that: (1) using intra-day step and sleep data recorded from passive trackers significantly improves classification performance on self-reported chronic conditions related to ment…
New approach models computer network activity as mixtures of sources.
problem Malicious activity detection in computer networks using standard algorithms is ineffective.
method Source separation approach to model short-term dynamics of computer network activity.
result Qualitative and quantitative experiments validate the approach.
Gradient descent memorizes many Gaussians efficiently.
problem Memorizing many Gaussians with minimal parameters.
method Gradient descent on a depth-two neural network.
result One step of gradient descent memorizes $Ω\left(\frac{dq}{\log^4(d)}
ight)$ Gaussians.
Study uses simulation-based inference to decode brain activity from synthetic stimuli.
problem Reversing the process of brain activity emulation to recover stimuli or their properties.
method Pairing brain emulator with LLMs to learn a probabilistic mapping from brain maps to stimulus parameters.
result LLMs can serve as controllable stimulus generators and parameters can be recovered from brain maps.
Incremental methods for structure learning of pairwise Markov random fields (MRFs), such as grafting, improve scalability by avoiding inference over the entire feature space in each optimization step. Instead, inference is performed over an incrementally grown active set of features. In this paper, we address key compu…
New method protects sensitive data in deep learning training.
problem Protecting sensitive data in deep learning training.
method Distributed layer-partitioned training with step-wise activation functions.
result Experimental results show the method is simple and effective.
A novel probabilistic approach forecasts imbalance prices in Belgium.
problem Forecasting imbalance prices in short-term energy markets.
method Two-step approach: compute net regulation volume state transition probabilities, then infer imbalance prices.
result The probabilistic approach outperforms deterministic and Gaussian Process models.
A method improves deep network accuracy with low precision quantization.
problem Maintaining high accuracy in low precision deep networks.
method Learned Step Size Quantization, improving quantizer configuration and gradient estimation.
result Achieves highest accuracy on ImageNet with 2-4 bit precision models.
We propose an active set selection framework for Gaussian process classification for cases when the dataset is large enough to render its inference prohibitive. Our scheme consists of a two step alternating procedure of active set update rules and hyperparameter optimization based upon marginal likelihood maximization.…
BatchGFN uses generative flow networks for efficient batch active learning.
problem Efficiently selecting informative batches for active learning.
method Generative flow networks to sample batches proportional to a batch reward.
result Constructs highly informative batches with a single forward pass per point.
Sideways trains video models by overwriting activations as new frames arrive, potentially improving generalization.
problem Training deep video models synchronously slows down and requires storing activations, limiting parallelism.
method Sideways trains video models by overwriting activations as new frames arrive, breaking the precise correspondence between gradients and activations.
result Sideways training can converge and potentially generalize better than standard synchronized backpropagation.
The study connects deep neural networks with statistical mechanics, revealing natural activation functions.
problem Understanding the activation functions in deep neural networks.
method Statistical Mechanics model of deep neural networks, focusing on encoding, validation, and propagation steps.
result A set of natural activations including Sigmoid, tanh, ReLU, and Swish are identified.
Neural networks plateau during training, identified and quantified.
problem Plateau phenomenon in gradient descent training of ReLU networks.
method Identification and quantification of plateau phenomenon; new iterative training method ANLS.
result Plateaux correspond to periods of constant activation patterns; quantification of gradient flow dynamics; characterization of stationary points.
Our work proves convergence to low robust training loss for polynomial width ReLU networks.
problem Understanding why adversarial training leads to low robust training loss in over-parameterized neural nets.
method Extending convergence theory for standard supervised training to adversarial training, using tools from online learning and showing ReLU networks can approximate the step function.
result Convergence to low robust training loss for polynomial width ReLU networks under natural assumptions.
Proposes an active RBI framework using Rényi information measures for more informed decision-making.
problem Optimal latent variable estimates in real-time settings with streaming noisy observations.
method Unified inference and query selection steps through Rényi entropy and α-divergence; new objective called Momentum for exploration.
result Analytically demonstrates superior performance compared to conventional methods like mutual information.
Improved EXACT strategy reduces GNN memory consumption and runtime.
problem Efficiently training large-scale GNNs with reduced memory usage.
method Block-wise quantization of intermediate activation maps with improved variance minimization.
result Further reduction in memory consumption (>15%) and runtime speedup (5%) with similar performance trade-offs.
Novel neural network models using convex optimization for improved training.
problem Improving the training of neural networks.
method Representing activation functions as argmin of convex optimization problems, applying block-coordinate descent methods.
result Proposed models provide excellent initial guesses and avenues for extensions.
The paper improves Gaussian process models for efficient batch optimization.
problem Poor scaling and optimization loop issues in Gaussian process models.
method Dual GP parameterization for linear scaling and non-Gaussian likelihood updates.
result Extends sparse models to greedy batch fantasizing acquisition functions.
Proposes Hebbian-descent for neural network learning, addressing Hebbian and gradient descent issues.
problem Learning issues with correlated data and vanishing error term in gradient descent.
method Integrates Hebbian and gradient descent principles without activation function derivatives, centering neural activities.
result Biologically plausible, convergent, and effective in online learning with correlated data.
New activation functions improve neural network stability and generalize well.
problem Proving theoretical generalization of non-parametric activation functions.
method Stability analysis of non-convex models trained with SGD.
result Neural networks with kernel activation functions generalize well with SGD.
Automated quality control for seismic data reduces human labor and time.
problem Costly and time-consuming manual QC of seismic data.
method Active learning to select and label relevant seismic data.
result Active learning technique reduces QC time and improves accuracy.
Improved OSV with active transfer learning for Persian signatures.
problem Challenges in OSV with skilled forgeries and limited labeled data.
method Active transfer learning using pre-trained CNN and SVM for active learning.
result Near 13% improvement over random selection and 1% over state-of-the-art.
Three-hidden-layer neural networks can approximate Hölder continuous functions uniformly with exponential rate.
problem Approximating Hölder continuous functions with neural networks.
method Introduced Floor-Exponential-Step (FLES) networks with three hidden layers.
result Uniform approximation of Hölder continuous functions with an exponential rate.