New neural stack and Turing Machine architectures prove stability and computational power.
problem Designing stable neural network architectures for Turing Machine simulation.
method Introducing neural stack and Turing Machine architectures, proving stability and computational equivalence.
result Differentiable nnTM with bounded neurons can simulate Turing Machine in real-time and is equivalent to UTM.
RocketStack integrates predictions from multiple base learners using a recursive stacking architecture up to ten levels.
problem Feature redundancy, complexity, and computational burden in deep stacking.
method Level-aware recursive stacking with pruning and compression techniques.
result Increasing accuracy with depth and outperforming standalone ensembles at later levels.
StackGAN++ generates high-quality images from text descriptions.
problem Generating high-quality photo-realistic images from text descriptions.
method Two-stage and multi-stage generative adversarial networks (GANs) with stacked architecture.
result StackGAN++ significantly outperforms other methods in generating photo-realistic images.
Tree-SMU enables strong compositional generalization in neural networks.
problem Zero-shot generalization to novel compositions of concepts.
method Tree Stack Memory Units (Tree-SMU) with Stack Memory Units (SMU).
result Tree-SMU achieves strong empirical results on mathematical reasoning benchmarks.
Study uses stacked hourglass networks to improve facial landmark detection for medical diagnosis.
problem Improving accuracy of facial landmark detection for medical diagnosis.
method Conducted a study on landmark localisation methods using stacked hourglass networks.
result State-of-the-art stacked hourglass architecture outperforms traditional methods.
Delayed-RNN approximates stacked and bidirectional RNNs.
problem Improving RNN expressiveness and representational capacity.
method Weight-constrained delayed-RNN, equivalent to stacked-RNNs, with partial acausality.
result Delayed-RNN can approximate stacked and bidirectional RNNs, outperforming them in some tasks.
Proposes SVM-based Deep Stacking Network for improved deep learning.
problem Improving deep learning performance and interpretability.
method Uses stacked SVM classifiers within a DSN architecture and a BP-like layer tuning scheme.
result Demonstrates superior performance compared to benchmark models on image and text data.
Two potential bottlenecks on the expressiveness of recurrent neural networks (RNNs) are their ability to store information about the task in their parameters, and to store information about the input history in their units. We show experimentally that all common RNN architectures achieve nearly the same per-task and pe…
Autostacker uses EA to evolve machine learning pipelines without domain knowledge.
problem Finding optimal machine learning models and hyperparameters without expert knowledge.
method Hierarchical stacking architecture + Evolutionary Algorithm (EA) for efficient parameter search.
result Autostacker achieves state-of-the-art performance on 15 datasets.
VTA offers flexible DL specialization for evolving workloads.
problem Inflexible specialized DL hardware accelerators.
method Parametrizable architecture, two-level ISA, JIT compiler.
result Flexible deep learning specialization on edge-class FPGAs.
Deep learning boosts building energy load forecasting.
problem Short-term load forecasting in buildings.
method Stacked Boosters Network architecture with sparse interactions, parameter sharing, and equivariant representations.
result Outperforms state-of-the-art models in short-term load forecasting tasks.
Haar scattering networks improve pattern recognition across various tasks.
problem Improving pattern recognition in diverse tasks like regression and classification.
method Stacking convolutional filters based on Haar wavelets followed by non-linear operators.
result Outperformed best algorithms in 4 out of 18 data classification problems.
This paper optimizes deep neural networks for resource-constrained devices.
problem Efficient deployment of deep neural networks on resource-constrained devices.
method Across-stack optimization of CNNs using weight pruning, channel pruning, quantization, and parallel execution.
result Comprehensive Pareto curves for trade-offs between accuracy, execution time, and memory space.
iPrescribe offers fast online offer recommendations using deep learning.
problem Online offer recommendation in real-time.
method Ensemble of deep learning and machine learning algorithms, optimized streaming technology stack, and efficient LSTM deployment.
result 90th percentile recommendation latency of 38 milliseconds.
We present a novel architecture, the "stacked what-where auto-encoders" (SWWAE), which integrates discriminative and generative pathways and provides a unified approach to supervised, semi-supervised and unsupervised learning without relying on sampling during training. An instantiation of SWWAE uses a convolutional ne…
In this work, we propose a novel recurrent neural network (RNN) architecture. The proposed RNN, gated-feedback RNN (GF-RNN), extends the existing approach of stacking multiple recurrent layers by allowing and controlling signals flowing from upper recurrent layers to lower layers using a global gating unit for each pai…
Generates biomedical abstracts from titles, years, and keywords.
problem Difficulties in understanding biomedical research papers due to specialized language.
method Conditional transformer-based language model with metadata conditioning.
result Generated abstracts are more relevant and non-trivial than GPT-2.
Paper combines multiple ETA models into a stacked ensemble for better ETA predictions.
problem Improving ETA predictions for taxi schedules and trips.
method Developed a two-level stacked ensemble model and applied XAI methods to explain it.
result The stacked ensemble model outperforms previous ETA approaches.
This paper introduces early exits in neural networks for faster inference.
problem Reducing inference time and preventing overfitting in neural networks.
method Designing and training multi-output neural networks with early exits.
result Significant reductions in inference time and improved robustness.
LATTE tackles heterogeneous network embedding challenges with layer-stacked attention.
problem Aggregating higher-order indirect relations in heterogeneous networks.
method Layer-stacked ATTention Embedding (LATTE) that decomposes meta relations at each layer.
result LATTE achieves state-of-the-art performance on benchmark datasets.
This paper shows that explicitly learning motion improves reinforcement learning in dynamic environments.
problem Learning controllers for dynamic environments without explicit motion representation.
method Explicitly learning motion representation using image difference or temporal stacks of frames.
result Explicit motion learning improves the quality of learned controllers in dynamic scenarios.
Dual neural network architecture improves accuracy and interpretability.
problem Improving neural network interpretability and accuracy.
method Stacked recurrent and feedforward layers, binary activation function.
result Binary activation leads to simpler, more interpretable models with higher accuracy.
A graph-based evolutionary algorithm automates machine learning workflows.
problem Automated machine learning to reduce manual operations.
method Graph-based architecture for flexible model combinations, evolutionary algorithm with mutation and heredity operators, Bayesian hyper-parameter optimization.
result State-of-the-art performance compared to other AutoML systems.
Capsule Networks improve clothing retrieval without landmark info.
problem In-shop clothing retrieval performance improvement.
method Triplet-based Capsule Network architecture with SC and RC blocks.
result Triplet Capsule Networks outperform FashionNet and SOTA architectures.
SDQL uses modular deep Q networks to efficiently learn multi-stage optimal control tasks.
problem Training complex deep reinforcement learning models for multi-stage control tasks is inefficient and unstable.
method Stacked Deep Q Learning (SDQL) with modular Q networks and backward training.
result SDQL efficiently learns optimal control policies for multi-stage tasks with high-dimensional state and action spaces.
Recurrent Neural Networks (RNNs) have long been recognized for their potential to model complex time series. However, it remains to be determined what optimization techniques and recurrent architectures can be used to best realize this potential. The experiments presented take a deep look into Hessian free optimization…
The paper analyzes how stacking improves model stability.
problem Lack of theoretical insight into how stacking works.
method Stability analysis of learning algorithms, focusing on hypothesis stability.
result The hypothesis stability of stacking is a product of base models and combiner.
We investigate unsupervised pre-training of deep architectures as feature generators for "shallow" classifiers. Stacked Denoising Autoencoders (SdA), when used as feature pre-processing tools for SVM classification, can lead to significant improvements in accuracy - however, at the price of a substantial increase in co…
We review the basic definition of a stack and apply it to the topological and smooth settings. We then address two subtleties of the theory: the correct definition of a ``stack over a stack'' and the distinction between small stacks (which are algebraic objects) and large stacks (which are generalized spaces).
In this article, we derive many properties of étale stacks in various contexts, and prove that étale stacks may be characterized categorically as those stacks that arise as prolongations of stacks on a site of spaces and local homeomorphisms. Moreover, we show that the bicategory of étale differentiable stacks and loca…
STAR-GCN improves recommender systems by learning node representations.
problem Cold start problem in recommender systems.
method Stacked and reconstructed Graph Convolutional Networks (GCN) with intermediate supervision and node embedding reconstruction.
result Significant improvements in predicting ratings, especially in the cold start scenario.
Large receptive field CNNs improve distant speech recognition.
problem Degrading performance of ASR systems in noisy environments.
method Investigated large receptive field CNN variants including recursive, dilated, and hourglass networks.
result Stacked hourglass networks show significant improvements in distant speech recognition.
Grow and prune LSTM to reduce parameters and latency.
problem Model redundancy, increased run-time delay, and overfitting in deep LSTM models.
method Hidden-layer LSTM (H-LSTM) with grow-and-prune (GP) training.
result Significant reduction in parameters, latency, and improvement in accuracy.
Paper introduces HISA for efficient FHE computations.
problem Efficiently evaluating encrypted neural networks.
method Developed HISA for FHE applications, including compiler and runtime.
result Generated code is faster than hand-optimized implementations.
Constructs cohomology decompositions for symmetric stacks.
problem Cohomology of symmetric stacks.
method Constructs decompositions of cohomology, Borel--Moore homology, and vanishing cycle cohomology.
result Defines BPS cohomology and proves its equivalence to intersection cohomology for smooth stacks.
Bayesian stacking improves model performance with varying model weights.
problem Improving model predictions with heterogeneous input performance.
method Bayesian hierarchical stacking with varying model weights inferred via Bayesian inference.
result Hierarchical stacking yields better predictions than linear averaging.
The paper compares different deep architectures for feature learning from EHRs.
problem Extracting meaningful insights from high-dimensional, sparse clinical data.
method Uses stacked sparse autoencoders, deep belief networks, adversarial autoencoders, and variational autoencoders for feature representation.
result Stacked sparse autoencoders perform better for small data sets, while variational autoencoders outperform for large data sets.
GCN and GPCA are mathematically connected, leading to improved node classification performance.
problem Improving node classification performance in semi-supervised settings.
method Established a mathematical connection between GCN and GPCA, demonstrating their equivalence and using this to design an effective initialization strategy.
result GPCA paired with a simple MLP achieves similar or better performance than GCN on semi-supervised node classification tasks.
This thesis explores geometric stacks and Poisson manifolds, proving new results in their classification and equivalence.
problem Classifying and understanding geometric stacks and Poisson manifolds.
method Rigorous proofs and new site constructions for geometric stacks and Poisson manifolds.
result Classification and equivalence results for b-symplectic manifolds.
Relates discrete group actions to orbit spaces as differentiable stacks.
problem Understanding dynamics of discrete groups on manifolds.
method Relating discrete group actions to orbit spaces as differentiable stacks.
result Orbit stack encodes dynamics up to conjugation and inversion.
Paper combines machine learning and model averaging for robust parameter estimation.
problem Estimating structural parameters with partially unknown functional forms.
method Pairing double/debiased machine learning with stacking for model averaging.
result DDML with stacking is more robust to unknown functional forms than single learners.
Improved traffic forecasting model handles missing data.
problem Short-term traffic forecasting with missing values.
method Proposed SBU-LSTM architecture with bidirectional and unidirectional LSTM.
result Superior performance in accuracy and robustness for network-wide traffic prediction.
Develops theory of differential graded schemes for derived stacks.
problem Creating a theory for derived stacks using dg schemes.
method Formulates dg schemes as homotopy sites, equates to stacks on dg algebras.
result Infinity category of stacks represented by dg schemes is derived schemes.
We generalize the notion of a small sheaf of sets over a topological space or manifold to define the notion of a small stack of groupoids over an étale topological or differentiable stack. We then provide a construction analogous to the étalé space construction in this context, establishing an equivalence of 2-categori…
ORIGAMI accelerates ML algorithms by splitting compute tasks between in-memory and off-chip accelerators.
problem Memory bandwidth bottleneck in ML processing.
method Heterogeneous in-memory accelerators and off-chip compute platform, pattern-matching for compute patterns, computation-splitting compiler.
result ORIGAMI outperforms state-of-the-art accelerators in performance and energy-efficiency.
NN-Stacking improves predictive power of regression models by adjusting stacking coefficients with features.
problem Low predictive power of linear stacking methods.
method NN-Stacking uses neural networks to estimate adaptive stacking coefficients.
result NN-Stacking leads to better predictive power, especially in large datasets.
Meta-learning framework for credit risk assessment of SMEs, aligning financial statement dates with evaluation dates.
problem Temporal misalignment of credit scoring models leading to bias and inconsistent predictions.
method Two-step temporal decomposition: static model for annual PDs, dynamic model for monthly PDs; stacking architecture to aggregate multiple models.
result Framework effectively captures credit risk evolution over time, improving temporal consistency and predictive stability.
When using deep, multi-layered architectures to build generative models of data, it is difficult to train all layers at once. We propose a layer-wise training procedure admitting a performance guarantee compared to the global optimum. It is based on an optimistic proxy of future performance, the best latent marginal. W…