DNNet slices network for efficient inference.
problem Resource constraints in inference.
method Doubly nested network with channel-wise nesting and channel-causal convolutions.
result Resource-efficient inference with sub-models.
RecNets use RNNs to process image channels in a compact, recurrent way.
problem Creating efficient neural network architectures for computer vision.
method Introducing RecNets with CRC layers that simulate recurrent processing of image channels.
result RecNets achieve superior size-accuracy trade-off compared to other compact models.
A new method for early stopping in neural networks without validation sets.
problem Determining when to stop training neural networks to avoid overfitting.
method Channel-wise DeepNNK (CW-DeepNNK) using non-negative kernel regression and polytope interpolation.
result The proposed early stopping criterion based on CW-DeepNNK performs better than standard validation-based methods.
PixelHop++ improves image classification with a smaller model size.
problem Improving image classification models with smaller sizes.
method Decomposing input tensor, channel-wise Saab transform, successive subspace learning, feature ranking.
result PixelHop++ offers a flexible tradeoff between model size and performance.
New graph attention operators improve performance and reduce computational costs.
problem Excessive computational resources in graph attention operators.
method Proposed hGAO and cGAO using hard and channel-wise attention mechanisms.
result Improved performance and computational savings with new operators.
A new distillation method transfers channel information from teacher to student.
problem Transfer knowledge from teacher to student with fewer parameters and calculations.
method Channel Distillation (CD) and Guided Knowledge Distillation (GKD) with loss decay.
result Achieved 27.68% top-1 error on ImageNet with ResNet18, outperforming state-of-the-art methods.
New method prunes neural networks while conserving resources.
problem Design efficient neural networks for edge devices with limited computational power.
method Iterative channel-wise pruning with holonomic constraint in Lagrangian multipliers framework.
result Reduction of 15.47 GMAC to 3.87 GMAC for VGG-16 with only 1% top-1 accuracy loss.
FEM improves attention mechanisms by applying value-driven log-linear tilts.
problem Standard attention mechanisms read via convex average, limiting channel-wise selection.
method Free Energy Mixer (FEM) applies a value-driven, per-channel log-linear tilt to a fast prior over indices.
result FEM outperforms strong baselines on NLP, vision, and time-series tasks.
SAMformer improves transformer performance in time series forecasting.
problem Transformers struggle with multivariate long-term forecasting.
method Sharpness-aware minimization and channel-wise attention.
result SAMformer surpasses state-of-the-art methods in multivariate time series forecasting.
New technique trains deep neural networks without normalization or minibatch statistics.
problem Training deep neural networks at high learning rates without normalization.
method Channel-wise zero-mean initialization and gradient modification to maintain common mode rejection.
result Achieves higher accuracy compared to batch normalization and shows minibatches are unnecessary.
Study cobordisms of nested manifolds and their invariants.
problem Understanding cobordisms of nested manifolds and their invariants.
method Identify a nested analog of the Pontryagin-Thom construction and find spaces homotopy equivalent to nested Pontryagin-Thom spaces.
result Discover nested cobordism invariants and provide an alternative proof of Wall's splitting result.
New methods for estimating nested expectations in machine learning.
problem Nested expectations in machine learning and statistics.
method Investigation and analysis of statistical implications of nesting Monte Carlo estimators.
result Established conditions for convergence of nested MC estimators and derived corresponding rates.
Quantum algorithm speeds up nested expectation estimation by nearly quadratically.
problem Estimating repeatedly nested expectations with quantum computing.
method Proposes a quantum algorithm achieving nearly quadratic speedup over classical methods.
result Achieves nearly quadratic speedup for RNEs, up to logarithmic factors.
Improved nested simulation for financial risk measurement.
problem Efficiently estimating nested risk measures in financial engineering.
method Reusing inner simulation outputs to improve efficiency and accuracy.
result The proposed approach outperforms standard nested simulation and regression methods.
Scalable tools for nested optimization in deep learning.
problem Solving nested optimization problems on a large scale in deep learning.
method Building scalable tools for bilevel optimization.
result Tools for nested optimization scale to deep learning setups.
Unified SGD method improves convergence for nested optimization problems.
problem Stochastic nested optimization problems.
method ALTERNATE dESCEN (ALSET) method leveraging hidden smoothness.
result Requires O ( ε − 2 ) {\cal O}(ε^{-2}) O ( ε − 2 ) samples to achieve an ε ε ε -stationary point. Paper proposes a new estimator for nested expectations with faster convergence.
problem Estimating nested expectations is computationally challenging.
method Nested kernel quadrature estimators with proof of faster convergence rate.
result The proposed method requires fewer samples for accurate estimation.
Let R be an o-minimal expansion of the real field, and let L(R) be the language consisting of all nested Rolle leaves over R. We call a set nested subpfaffian over R if it is the projection of a boolean combination of definable sets and nested Rolle leaves over R. Assuming that R admits analytic cell decomposition, we …
Develops a method for learning proposals in nested importance samplers.
problem Improving sampling quality in complex distributions.
method Nested Variational Inference (NVI) using forward or reverse KL divergence.
result Optimizing nested objectives leads to improved sample quality.
Nested model averaging improves high-dimensional linear regression performance.
problem High-dimensional linear regression with predictor ordering impact.
method Combining model averaging with regularized estimators on the solution path.
result Nested model averaging with lasso and SLOPE outperforms competing methods.
Nested Slice Sampling accelerates Nested Sampling for GPU acceleration.
problem Challenging inference for complex, multimodal targets.
method Vectorized Nested Slice Sampling using Hit-and-Run Slice Sampling.
result NSS maintains accurate evidence estimates and high-quality posterior samples, robust on multimodal problems.
Gradient-guided nested sampling improves posterior inference efficiency.
problem Efficiently sampling from complex posterior distributions.
method Gradient-guided nested sampling combining differentiable programming, Hamiltonian slice sampling, clustering, mode separation, dynamic nested sampling, and parallelization.
result Significantly faster mode discovery and more accurate partition function estimates.
We formalize nesting in probabilistic queries and correct inconsistent estimates.
problem Inconsistent estimates in probabilistic query nesting.
method Formalized nesting, introduced online nested Monte Carlo estimator, proved correctness and asymptotic variance.
result Corrected inconsistent estimates through new estimator and conditions.
New discrete cobordism category for nested manifolds and relations to algebraic structures.
problem Discrete cobordism category for nested manifolds.
method Stratified Morse theory, Cyl-objects, doubling construction, cylindrical bar construction.
result Relations between Cyl-objects and algebraic structures like Temperley-Lieb algebras.
Nested sampling improved for arbitrary priors.
problem Technical obstacle to using nested sampling with arbitrary priors.
method Parametric bijectors trained on samples from a desired prior density.
result Nested sampling can be used with arbitrary priors.
Improves predictive performance of nested dichotomies.
problem Improving the performance of multi-class classification problems.
method A simple, general method for improving nested dichotomies produced by random subset selection techniques.
result Improves root mean squared error of nested dichotomies.
Nested dichotomies for multiclass tasks often fail to calibrate probabilities.
problem Poor probability calibration in nested dichotomies for multiclass classification.
method Transforming multiclass problems into binary ones using a tree structure, and applying various calibration strategies.
result Improving accuracy and log-loss by calibrating both internal base models and the nested dichotomy structure.
Study of loops in sums of Laplace eigenfunctions on surfaces.
problem Uniform bound for the number of nested loops in sums of Laplace eigenfunctions.
method Real-analytic category analysis and biharmonic function construction.
result Uniform bound for the number of rooted double nests in terms of surface, root, and spectral cutoff.
Enhances Bayesian model selection for high-dimensional problems.
problem Bayesian model selection for high-dimensional problems.
method Proximal nested sampling with data-driven priors.
result Improves model selection for log-convex likelihood models.
Nested learning improves model performance on multi-granular tasks.
problem Overconfident models and lack of fine-grained confidence in predictions.
method Introducing nested learning with a sequence of nested feature embeddings and explicit combination of outputs.
result Nested learning outperforms standard end-to-end training on various datasets.
New AD algorithms speed up learning in integer models.
problem Efficient inference and learning in integer latent variable models.
method Nested automatic differentiation for exact inference and learning.
result Significantly faster and more accurate learning algorithms.
NEST optimizes deep learning training by placing devices efficiently across networks and memory.
problem Inefficient device placement in distributed deep learning leads to high communication and memory overhead.
method NEST uses network-, compute-, and memory-aware dynamic programming to optimize device placement.
result NEST achieves up to 2.43 times higher throughput and better memory efficiency.
Spiking-YOLO improves object detection with low power and fast convergence.
problem Challenging object detection tasks with spiking neural networks.
method Channel-wise normalization and signed neuron with imbalanced threshold.
result Spiking-YOLO achieves comparable results to Tiny YOLO but with significantly less energy consumption.
New methods for estimating complex causal effects in econometrics.
problem Estimating causal parameters in short panel data models using nested nonparametric instrumental variable regression.
method Introducing techniques to limit ill-posedness in nested NPIV, providing explicit mean square rates and efficient inference.
result Explicit mean square rates for nested NPIV and efficient inference for causal parameters.
Sharp lower bound on GHHs' representation power of CPWL functions.
problem Proving the minimum number of nestings for GHHs to represent arbitrary CPWL functions.
method Using a key lemma about finite sums of periodic functions, proving necessity of n nestings.
result Proving necessity of n nestings for GHHs to achieve universal representation power.
End-to-end model extracts nested terms without extra features.
problem Automatic term extraction for nested terms.
method Deep learning model that predicts conceptual terms within fixed sentence lengths.
result High recall and comparable precision on term extraction task.
There is an increasing interest in estimating expectations outside of the classical inference framework, such as for models expressed as probabilistic programs. Many of these contexts call for some form of nested inference to be applied. In this paper, we analyse the behaviour of nested Monte Carlo (NMC) schemes, for w…
Paper tackles robust model training with a new stochastic algorithm.
problem Training robust models against data distribution shift.
method Derives a novel dual formulation and proposes a nested stochastic gradient descent algorithm.
result Establishes polynomial iteration and sample complexities for large-scale DRO problems.
EENNs improve inference efficiency but need nested prediction sets for reliable uncertainty estimates.
problem Non-nested prediction sets from standard uncertainty quantification methods in EENNs.
method Introduced anytime-valid confidence sequences (AVCSs) tailored for EENNs.
result AVCSs generate nested prediction sets across EENN exits, addressing the issue of non-nested sets.
Paper tackles robust optimization under uncertainty using nested distance.
problem Optimizing under distributionally robust uncertainty with nested distance.
method Equivalent recursive and dynamic programming reformulations for tractable optimization.
result Optimal robust policies can be found efficiently using convex optimization.
New formulas compare total mean curvatures of nested hypersurfaces.
problem Computing total mean curvatures of nested hypersurfaces.
method Developed differential forms based on Chern's work to compare curvatures.
result Quicker proof of recent result on total mean curvatures.
Proposes Dirichlet Simplex Nest for probabilistic modeling of various data types.
problem Modeling and inference for diverse data types.
method Probabilistic models based on Dirichlet distribution and Voronoi tessellation, with fast and accurate inference algorithms exploiting convex geometry and simplicial structure.
result Inference algorithms achieve consistency and strong error bounds across various settings and data distributions.
New unbiased gradient estimators for complex optimization problems.
problem Unbiased and variance-limited gradient estimation for conditional stochastic optimization.
method Developed multilevel Monte Carlo gradient estimators for conditional stochastic optimization problems.
result Unbiased and finite variance gradient estimators for conditional stochastic optimization problems.
ESE-FN improves elderly activity recognition accuracy.
problem Recognizing individual actions and human-object interactions in elderly activities.
method Exploits multi-modal features from RGB videos and skeleton sequences using ESE attentions and a new Multi-modal Loss.
result ESE-FN achieves best accuracy on ETRI-Activity3D dataset.
New estimator reduces nested expectation estimation costs.
problem Estimating repeatedly nested expectations is computationally expensive.
method Recursive Estimator for Arbitrary Depth (READ) using randomized multilevel Monte Carlo.
result Optimal computational cost of O(ε^(-2)) for every fixed D.
Study dynamic assortment planning under nested logit models for revenue maximization.
problem Maximize revenue by dynamically selecting assortments of products during a selling season.
method Developed a novel UCB policy that learns and makes decisions based on customers' choice behavior.
result Achieved accumulated regret of i l d e O ( M N T ) ilde{O}(\sqrt{MNT}) i l d e O ( M N T ) with a lower bound of Ω ( M T ) Ω(\sqrt{MT}) Ω ( M T ) . A system of nested dichotomies is a method of decomposing a multi-class problem into a collection of binary problems. Such a system recursively splits the set of classes into two subsets, and trains a binary classifier to distinguish between each subset. Even though ensembles of nested dichotomies with random structure…
Study geodesics on nested non-holonomic systems.
problem Interplay between geodesics on related non-holonomic systems.
method Hamiltonian formalism and geometric preliminaries.
result Presented several geometric examples, including exotic spheres and twistor space.