GSANet improves semantic segmentation accuracy with selective and global attention.
problem Semantic segmentation accuracy improvement.
method Global and selective attention mechanism with ASPP and sparsemax.
result GSANet achieves state-of-the-art accuracy on ADE20k and Cityscapes datasets.
Generates realistic pedestrian data for training detectors.
problem Lack of labeled pedestrian data for training detectors.
method Pedestrian-Synthesis-GAN using GAN with multiple discriminators and SPP layer.
result Synthetic pedestrians improve detector performance.
Proposes SPE for robust speaker verification.
problem Improving text-independent speaker verification accuracy.
method Spatial pyramid encoding and deep length normalization.
result Proposed system outperforms i-vector and d-vector baselines.
Paper tackles fixed-size representation learning for variable-sized signatures.
problem Learning feature representations for signatures of varying sizes.
method Modified Spatial Pyramid Pooling to learn fixed-sized representations from variable-sized signatures.
result Comparable performance to state-of-the-art on GPDS dataset, removing size constraint.
Pyramidal GNN combines RC and pooling for efficient graph embeddings.
problem Efficiently embedding graphs while maintaining accuracy.
method Alternates RC layers with pooling to reduce complexity.
result Formally shows how pooling reduces complexity and speeds convergence.
Improves speaker verification for variable-duration utterances using a feature pyramid module.
problem Improving robustness for variable-duration utterances in speaker verification.
method Integrates a feature pyramid module into multi-scale aggregation to enhance speaker-discriminative information from multiple layers.
result Improves performance for both short and long utterances compared to state-of-the-art approaches.
We define invariants for colored oriented spatial graphs by generalizing CM invariants, which were defined via non-integral highest weight representations of Uq(sl2). We apply the same method to define Yokota's invariants, and we call these invariants Yokota type invariants. Then we propose a volume conjecture of t…
Improved vehicle classification using ResNets and spatial pooling.
problem Fine-grained vehicle classification using ResNet architectures.
method Training ResNet-18, -34, and -50 on Comprehensive Cars dataset. Adding Spatially Weighted Pooling and localisation.
result Combining Spatially Weighted Pooling and localisation increases top-1 accuracy to 96.351%.
SPF uses a hierarchical approach to efficiently emulate climate changes.
problem Slow and unstable climate emulation for long horizons.
method Spatiotemporal Pyramid Flows (SPF) model data hierarchically across spatial and temporal scales.
result SPF outperforms flow matching baselines and pre-trained models on ClimateBench.
ASTPN improves video-based person re-identification by jointly attending to spatial and temporal features.
problem Video-based person re-identification in surveillance and HCI.
method Joint Spatial and Temporal Attention Pooling Network (ASTPN).
result ASTPN outperforms state-of-the-art methods on multiple datasets.
Improved TSC with BOSS and SP techniques.
problem Comparing BOP and BOSS for time series classification.
method Deconstructed and measured components of BOP and BOSS, adapted CV techniques.
result SP with BOSS significantly more accurate than benchmarks.
Deep learning improves automatic image segmentation.
problem Automatic object localization and boundary delineation in images and medical scans.
method Proposed and evaluated novel dilated dense encoder-decoder architectures for salient object segmentation and lesion localization in medical images.
result Proposed architectures outperform state-of-the-art models in accuracy and efficiency.
The challenge of object categorization in images is largely due to arbitrary translations and scales of the foreground objects. To attack this difficulty, we propose a new approach called collaborative receptive field learning to extract specific receptive fields (RF's) or regions from multiple images, and the selected…
Simpler CNN model with spatial attention and temporal pooling outperforms complex models.
problem Emotion recognition from videos with small face deformations and identity variations.
method Spatial attention mechanism and temporal softmax pooling applied to a pre-trained CNN.
result The approach achieves higher accuracy than state-of-the-art methods on the EmotiW dataset.
NDP improves GNN efficiency by coarsening graphs without losing structure.
problem Efficiently summarize graph data for deep learning models.
method Node Decimation Pooling (NDP) reduces graph density while preserving topology.
result NDP achieves comparable performance to state-of-the-art pooling methods but with improved efficiency.
The paper analyzes Laplacian pyramids for extending and denoising discrete functions.
problem Analyzing conditions for convergence and stability of Laplacian pyramids.
method Investigates Laplacian pyramids for extension and denoising, providing convergence conditions and stability bounds.
result Mild conditions are provided under which the Laplacian pyramids algorithm converges and stability bounds are proven.
Spatial graph representation improves GNN performance.
problem GNNs struggle with distinguishing similar local structures in different graph locations.
method Proposes a spatial graph representation method to distinguish local structures and simplify graph downsampling.
result Proposed graph pooling method achieves competitive results.
Improves instance segmentation accuracy by integrating low-level features.
problem Low-level features enhance instance localization accuracy.
method Integrates low-level features into all layers of a feature pyramid network.
result Consistent improvement in precision on COCO Dataset.
New image classifier uses hierarchical max-pooling with local pooling.
problem Improving image classification accuracy with variable spatial relationships.
method Introduces a hierarchical max-pooling model with additional local pooling for convolutional neural networks.
result Demonstrates improved performance in estimating image features.
Paper proposes CNN with SIFT for rotation invariant feature extraction.
problem Max-pooling layer discards rotational information, leading to rotation invariance issues.
method Uses SIFT descriptor to capture orientation and spatial relationships.
result Improves feature extraction on MNIST and fashionMNIST datasets.
Proposes a new pooling operator for CNNs to handle spatially varying information.
problem Need to treat spatial locations in non-uniform manner for better image classification.
method Introduces an extended pooling operator that can learn different weights for each pixel location.
result The proposed pooling operator improves generalization and robustness in image classification tasks.
Proposes SimPool for graph pooling using structural similarity features.
problem Challenges in graph pooling due to lack of spatial locality.
method Integrates structural similarity features with a revised pooling layer to propose SimPool.
result SimPool produces node cluster assignments resembling CNN's locality preserving pooling.
A novel approach for augmenting histopathological images by blending Gaussian-Laplacian pyramids.
problem Data imbalance and inter-patient variability in histopathological images.
method Image blending using Gaussian-Laplacian pyramids to distribute inter-patient variability.
result Promising gains in performance compared to existing data augmentation techniques.
Pyramid Attention Networks improve image restoration by leveraging self-similarities across scales.
problem Lack of full exploitation of self-similarities in image restoration by recent deep learning methods.
method Introduces a Pyramid Attention module that processes multi-scale feature correspondences to borrow clean signals from coarser levels.
result Pyramid Attention module achieves state-of-the-art results in various image restoration tasks.
Simulation reveals relationships in stock market pyramid schemes.
problem Understanding pyramid scheme behavior in stock markets.
method Agent-based simulation with four investor types and parameters.
result Relationships between main fund's rate of return and trend investors' proportion.
A new pyramidal diffusion model speeds up image generation.
problem Training and evaluating diffusion models are time-consuming.
method Introduces a pyramidal diffusion model with a single score function.
result Generates high-resolution images from low-resolution starting points efficiently.
Single wide layer followed by a pyramidal structure ensures global convergence in deep networks.
problem Ensuring global convergence in deep neural networks with limited width constraints.
method Proves that a single wide layer followed by a pyramidal structure guarantees global convergence for over-parameterized networks.
result Single wide layer of width N suffices for global convergence in deep networks with constant-width remaining layers. We generalize the observable diameter and the separation distance for metric measure spaces to those for pyramids, and prove some limit formulas for these invariants for a convergent sequence of pyramids. We obtain various applications of our limit formulas as follows. We have a criterion of the phase transition proper…
By formulating N = 1, 2, 4, 8, D = 3, Yang-Mills with a single Lagrangian and single set of transformation rules, but with fields valued respectively in R,C,H,O, it was recently shown that tensoring left and right multiplets yields a Freudenthal-Rosenfeld-Tits magic square of D = 3 supergravities. This was subsequently…
TP-AIS improves sampling efficiency over existing methods.
problem Efficient sampling from complex probability distributions.
method Iterative sampling using a tree pyramid structure.
result TP-AIS outperforms DM-PMC, M-PMC, and LAIS.
Study smoothings of ellipsoid intersections with singularities.
problem Smooth singular intersections of ellipsoids.
method Study 3D manifolds with singularities as small covers of Coxeter polyhedral orbifolds.
result Introduce n-pyramitoid to generalize n-pyramids. Faster CNN training with downsampling and pooling.
problem Reducing training time in CNNs.
method Interleaved training with downsampling and pooling.
result Up to 23% reduction in training time with minimal loss.
Deep convolutional networks can be understood through kernel methods, providing insights into their inductive bias.
problem Understanding the functional space and inductive bias of deep convolutional networks.
method Using kernel methods to analyze simple hierarchical kernels with convolution and pooling layers.
result The RKHS consists of additive models of interaction terms between patches, and pooling layers encourage spatial similarities.
The paper develops a new Ricci flow method in higher dimensions.
problem Constructing a Ricci flow on non-collapsed IC1-limit spaces. method Constructing a pyramid Ricci flow on a subset of space-time.
result Non-collapsed IC1-limit spaces are globally homeomorphic to smooth manifolds. Global homeomorphism constructed from 3D Ricci limit spaces.
problem Global regularity of 3D Ricci limit spaces.
method Local pyramid Ricci flows and partial Ricci flows.
result Global homeomorphism to a smooth manifold.
iPool selects informative nodes for pooling in arbitrary graphs.
problem Pooling in graph neural networks is often overlooked.
method iPool uses a criterion based on neighborhood conditional entropy to select nodes for pooling.
result iPool achieves state-of-the-art performance on graph classification tasks.
A new histogram layer improves texture analysis performance.
problem Extracting features for texture analysis from local spatial regions.
method Directly computes local spatial distribution of features during backpropagation.
result Improves performance on three material/texture datasets.
A new neural network learns from acoustic scenes by suppressing irrelevant patterns.
problem Acoustic scenes are rich and redundant, making classification challenging.
method Spatio-temporal attention pooling layer coupled with a convolutional recurrent neural network.
result The method outperforms a strong convolutional neural network baseline and sets new state-of-the-art performance.
A novel approach predicts long-term stock price trends using 2D-convolutional encoders and semantic segmentation.
problem Predicting long-term daily stock price changes with deep learning models.
method Proposes a hierarchical CNN structure with Atrous Spatial Pyramid Pooling blocks to capture both long and short-term temporal relationships.
result Achieved overall accuracy and AUC of 78.18% and 0.88 for predicting trends over the next 20 days.
WavPool improves deep neural networks with wavelet-based pooling.
problem Improving efficiency and performance of deep neural networks.
method Introducing WavPool, a wavelet-transform-based pooling layer.
result WavPool outperforms existing network architectures by 10% on CIFAR-10.
PyFi uses adversarial agents to train VLMs on financial image understanding.
problem Training VLMs to understand complex financial questions.
method PyFi-600K dataset and adversarial MCTS mechanism.
result Fine-tuned VLMs improve by 19.52% and 8.06% on financial question accuracy.
ST-UNet models spatio-temporal graphs by pooling and unpooling operations.
problem Lack of effective means to extract dynamic features from spatio-temporal graphs.
method Designing a multi-scale architecture, Spatio-Temporal U-Net (ST-UNet), with paired sampling operations.
result Achieves substantial improvements in spatio-temporal prediction tasks.
Infinite CNNs lose spatial correlations, but can be restored by correlated weights.
problem Infinite CNNs lose spatial correlations, which are crucial for their performance.
method Introduced correlated weights to restore spatial correlations in infinite CNNs.
result Optimal performance is achieved with a moderate level of weight correlation.
Spatial smoothing improves BNNs' accuracy, uncertainty, and robustness without increasing computational cost.
problem Large ensembles in BNNs increase computational cost and reduce performance.
method Spatial smoothing adds blur layers to convolutional neural networks to ensemble neighboring feature map points.
result Spatial smoothing improves BNNs' performance with fewer ensembles and enhances robustness.
Enhances neural networks' robustness against adversarial samples without sacrificing clean sample generalization.
problem Limited generalization and time complexity of adversarial training.
method Feature Pyramid Decoder (FPD) framework that integrates denoising and image restoration modules into CNNs and constrains the Lipschitz constant.
result FPD-enhanced CNNs achieve sufficient robustness against general adversarial samples on various datasets.
Deep neural network detects changepoints at multiple scales in multivariate time series.
problem Detecting gradual and abrupt changes in multivariate time series data.
method Proposed a pyramid recurrent neural network (PRN) for multi-scale changepoint detection.
result PRN outperforms state-of-the-art methods in detecting changepoints at multiple scales.
Detects parking spaces in parcels using satellite images.
problem Locating parking spaces in urban areas using satellite imagery.
method Used Feature Pyramid based Mask RCNN for localizing parking spaces and vehicles in parking lots.
result Average class accuracy of 97.56% for parking spaces and vehicles.
The GANs are generative models whose random samples realistically reflect natural images. It also can generate samples with specific attributes by concatenating a condition vector into the input, yet research on this field is not well studied. We propose novel methods of conditioning generative adversarial networks (GA…