Proposes SPE for robust speaker verification.
problem Improving text-independent speaker verification accuracy.
method Spatial pyramid encoding and deep length normalization.
result Proposed system outperforms i-vector and d-vector baselines.
We define invariants for colored oriented spatial graphs by generalizing CM invariants, which were defined via non-integral highest weight representations of Uq(sl2). We apply the same method to define Yokota's invariants, and we call these invariants Yokota type invariants. Then we propose a volume conjecture of t…
SPF uses a hierarchical approach to efficiently emulate climate changes.
problem Slow and unstable climate emulation for long horizons.
method Spatiotemporal Pyramid Flows (SPF) model data hierarchically across spatial and temporal scales.
result SPF outperforms flow matching baselines and pre-trained models on ClimateBench.
GSANet improves semantic segmentation accuracy with selective and global attention.
problem Semantic segmentation accuracy improvement.
method Global and selective attention mechanism with ASPP and sparsemax.
result GSANet achieves state-of-the-art accuracy on ADE20k and Cityscapes datasets.
The paper analyzes Laplacian pyramids for extending and denoising discrete functions.
problem Analyzing conditions for convergence and stability of Laplacian pyramids.
method Investigates Laplacian pyramids for extension and denoising, providing convergence conditions and stability bounds.
result Mild conditions are provided under which the Laplacian pyramids algorithm converges and stability bounds are proven.
Low level features like edges and textures play an important role in accurately localizing instances in neural networks. In this paper, we propose an architecture which improves feature pyramid networks commonly used instance segmentation networks by incorporating low level features in all layers of the pyramid in an o…
Deep learning improves automatic image segmentation.
problem Automatic object localization and boundary delineation in images and medical scans.
method Proposed and evaluated novel dilated dense encoder-decoder architectures for salient object segmentation and lesion localization in medical images.
result Proposed architectures outperform state-of-the-art models in accuracy and efficiency.
The key idea of variational auto-encoders (VAEs) resembles that of traditional auto-encoder models in which spatial information is supposed to be explicitly encoded in the latent space. However, the latent variables in VAEs are vectors, which can be interpreted as multiple feature maps of size 1x1. Such representations…
A family of algorithms for time series classification (TSC) involve running a sliding window across each series, discretising the window to form a word, forming a histogram of word counts over the dictionary, then constructing a classifier on the histograms. A recent evaluation of two of this type of algorithm, Bag of …
A novel approach for augmenting histopathological images by blending Gaussian-Laplacian pyramids.
problem Data imbalance and inter-patient variability in histopathological images.
method Image blending using Gaussian-Laplacian pyramids to distribute inter-patient variability.
result Promising gains in performance compared to existing data augmentation techniques.
Pyramid Attention Networks improve image restoration by leveraging self-similarities across scales.
problem Lack of full exploitation of self-similarities in image restoration by recent deep learning methods.
method Introduces a Pyramid Attention module that processes multi-scale feature correspondences to borrow clean signals from coarser levels.
result Pyramid Attention module achieves state-of-the-art results in various image restoration tasks.
Simulation reveals relationships in stock market pyramid schemes.
problem Understanding pyramid scheme behavior in stock markets.
method Agent-based simulation with four investor types and parameters.
result Relationships between main fund's rate of return and trend investors' proportion.
Computer-aided assessment of physical rehabilitation entails evaluation of patient performance in completing prescribed rehabilitation exercises, based on processing movement data captured with a sensory system. Despite the essential role of rehabilitation assessment toward improved patient outcomes and reduced healthc…
A novel approach predicts long-term stock price trends using 2D-convolutional encoders and semantic segmentation.
problem Predicting long-term daily stock price changes with deep learning models.
method Proposes a hierarchical CNN structure with Atrous Spatial Pyramid Pooling blocks to capture both long and short-term temporal relationships.
result Achieved overall accuracy and AUC of 78.18% and 0.88 for predicting trends over the next 20 days.
A new pyramidal diffusion model speeds up image generation.
problem Training and evaluating diffusion models are time-consuming.
method Introduces a pyramidal diffusion model with a single score function.
result Generates high-resolution images from low-resolution starting points efficiently.
We construct a series of finitely presented semigroups. The centers of these semigroups encode uniquely up to rigid ambient isotopy in 3-space all non-oriented spatial graphs. This encoding is obtained by using three-page embeddings of graphs into the product of the line with the cone on three points. By exploiting thr…
Single wide layer followed by a pyramidal structure ensures global convergence in deep networks.
problem Ensuring global convergence in deep neural networks with limited width constraints.
method Proves that a single wide layer followed by a pyramidal structure guarantees global convergence for over-parameterized networks.
result Single wide layer of width N suffices for global convergence in deep networks with constant-width remaining layers. We generalize the observable diameter and the separation distance for metric measure spaces to those for pyramids, and prove some limit formulas for these invariants for a convergent sequence of pyramids. We obtain various applications of our limit formulas as follows. We have a criterion of the phase transition proper…
PriorVAE uses VAEs to efficiently encode spatial priors for small-area estimation.
problem Efficiently encoding spatial priors for small-area estimation using Gaussian processes.
method Approximating Gaussian process priors with a variational autoencoder (VAE).
result Efficient spatial inference through a low-dimensional latent Gaussian space representation.
Time-aware deep learning methods improve spatial downscaling of atmospheric pollutants.
problem Transform coarse satellite data of atmospheric pollutants into high-resolution fields.
method Super-resolution deep residual networks and UNet architectures are extended with a temporal module encoding observation time.
result Temporal modules significantly improve downscaling performance and convergence speed.
By formulating N = 1, 2, 4, 8, D = 3, Yang-Mills with a single Lagrangian and single set of transformation rules, but with fields valued respectively in R,C,H,O, it was recently shown that tensoring left and right multiplets yields a Freudenthal-Rosenfeld-Tits magic square of D = 3 supergravities. This was subsequently…
TP-AIS improves sampling efficiency over existing methods.
problem Efficient sampling from complex probability distributions.
method Iterative sampling using a tree pyramid structure.
result TP-AIS outperforms DM-PMC, M-PMC, and LAIS.
Co-PLNet combines point and line predictions to improve wireframe parsing accuracy and efficiency.
problem Separate line and point predictions lead to inconsistent wireframes.
method Co-PLNet uses a Point-Line Prompt Encoder to convert early point detections into spatial prompts, which guide line refinement.
result Co-PLNet achieves better accuracy and robustness in wireframe parsing compared to existing methods.
MRCNet tackles crowd counting and density mapping in aerial imagery.
problem Accurate crowd counting and density estimation in aerial imagery.
method MRCNet is a novel encoder-decoder CNN that combines VGG-16 with FPN-inspired lateral connections.
result MRCNet outperforms state-of-the-art methods in aerial and CCTV-based crowd counting.
Study smoothings of ellipsoid intersections with singularities.
problem Smooth singular intersections of ellipsoids.
method Study 3D manifolds with singularities as small covers of Coxeter polyhedral orbifolds.
result Introduce n-pyramitoid to generalize n-pyramids. Improves speaker verification for variable-duration utterances using a feature pyramid module.
problem Improving robustness for variable-duration utterances in speaker verification.
method Integrates a feature pyramid module into multi-scale aggregation to enhance speaker-discriminative information from multiple layers.
result Improves performance for both short and long utterances compared to state-of-the-art approaches.
PE-GQNN improves spatial data prediction and uncertainty quantification.
problem Poor calibration of predictive distributions in spatial data models.
method Combines PE-GNNs with Quantile Neural Networks and recalibration techniques.
result PE-GQNN outperforms existing methods in predictive accuracy and uncertainty quantification.
iREPA shows spatial structure, not global semantic, drives generation performance in REPA.
problem Understanding what aspect of the target representation matters for generation.
method Empirical analysis of 27 vision encoders, two modifications to REPA.
result Spatial structure, not global semantic, drives generation performance.
Multi-step passenger demand forecasting is a crucial task in on-demand vehicle sharing services. However, predicting passenger demand over multiple time horizons is generally challenging due to the nonlinear and dynamic spatial-temporal dependencies. In this work, we propose to model multi-step citywide passenger deman…
The paper develops a new Ricci flow method in higher dimensions.
problem Constructing a Ricci flow on non-collapsed IC1-limit spaces. method Constructing a pyramid Ricci flow on a subset of space-time.
result Non-collapsed IC1-limit spaces are globally homeomorphic to smooth manifolds. This work creates a system for understanding human movement in spaces.
problem Simplify communication and interaction between robots and humans in spatial tasks.
method Uses unsupervised learning with neural autoencoding to learn continuous representations of spatio-temporal trajectory data.
result Proposes a method to form prototypical representations of movement based on spatial context.
We present a machine learning-based approach to lossy image compression which outperforms all existing codecs, while running in real-time. Our algorithm typically produces files 2.5 times smaller than JPEG and JPEG 2000, 2 times smaller than WebP, and 1.7 times smaller than BPG on datasets of generic images across all …
PyFi uses adversarial agents to train VLMs on financial image understanding.
problem Training VLMs to understand complex financial questions.
method PyFi-600K dataset and adversarial MCTS mechanism.
result Fine-tuned VLMs improve by 19.52% and 8.06% on financial question accuracy.
A new histogram layer improves texture analysis performance.
problem Extracting features for texture analysis from local spatial regions.
method Directly computes local spatial distribution of features during backpropagation.
result Improves performance on three material/texture datasets.
Hybrid model integrates GATv2 and geostatistics for better spatial prediction and uncertainty.
problem Accurate spatial prediction and uncertainty quantification in epidemiology and risk analysis.
method Integrates Graph Attention Network (GATv2) with model-based geostatistics (MBG) to capture relational and spatial dependencies.
result Hybrid model improves predictive accuracy and uncertainty quantification compared to standalone models.
State-of-the-art pedestrian detection models have achieved great success in many benchmarks. However, these models require lots of annotation information and the labeling process usually takes much time and efforts. In this paper, we propose a method to generate labeled pedestrian data and adapt them to support the tra…
ADAVI tackles variational inference for large HBM models in neuroimaging.
problem Large, pyramidally-organized HBM models in neuroimaging studies.
method Automatic dual amortized variational inference using neural networks and attention-based hierarchical encoding.
result Significantly reduced parameterization of the variational family, maintaining expressivity.
This work proposes a novel autoencoder for fusing visible and infrared images.
problem Challenging task to combine spatial and spectral information from visible and infrared images.
method Spatially constrained adversarial autoencoder with residual architecture and adversarial regularizer.
result Generates a more realistic fused image with enhanced spatial and spectral information.
Model place cells as spatial embeddings for efficient path planning and cognitive map construction.
problem Encoding spatial navigation in the hippocampus.
method Model place cells using spectral decomposition of multi-step random walk transition kernels, inducing sparsity and adjacency.
result Place cells encode spatial information through non-negativity and inner-product structure, forming a cognitive map.
Enhances neural networks' robustness against adversarial samples without sacrificing clean sample generalization.
problem Limited generalization and time complexity of adversarial training.
method Feature Pyramid Decoder (FPD) framework that integrates denoising and image restoration modules into CNNs and constrains the Lipschitz constant.
result FPD-enhanced CNNs achieve sufficient robustness against general adversarial samples on various datasets.
We present convolutional neural network (CNN) based approaches for unsupervised multimodal subspace clustering. The proposed framework consists of three main stages - multimodal encoder, self-expressive layer, and multimodal decoder. The encoder takes multimodal data as input and fuses them to a latent space representa…
Pyramidal GNN combines RC and pooling for efficient graph embeddings.
problem Efficiently embedding graphs while maintaining accuracy.
method Alternates RC layers with pooling to reduce complexity.
result Formally shows how pooling reduces complexity and speeds convergence.
Detects parking spaces in parcels using satellite images.
problem Locating parking spaces in urban areas using satellite imagery.
method Used Feature Pyramid based Mask RCNN for localizing parking spaces and vehicles in parking lots.
result Average class accuracy of 97.56% for parking spaces and vehicles.
We introduce a wavelet-domain functional analysis of variance (fANOVA) method based on a Bayesian hierarchical model. The factor effects are modeled through a spike-and-slab mixture at each location-scale combination along with a normal-inverse-Gamma (NIG) conjugate setup for the coefficients and errors. A graphical mo…
This paper focuses on a class of linear Hawkes processes with general immigrants. These are counting processes with shot noise intensity, including self-excited and externally excited patterns. For such processes, we introduce the concept of age pyramid which evolves according to immigration and births. The virtue if t…
Biological neural network mimics CCA for multi-channel data.
problem Implementing CCA in a biologically plausible neural network.
method Derive an online CCA algorithm with local synaptic updates for multi-compartmental neurons.
result The derived neural network architecture and synaptic updates resemble cortical pyramidal neuron behavior.
Space2Vec learns multi-scale spatial representations from grid cell insights.
problem Encoding spatial features with varying scales from GIS data.
method Proposes Space2Vec, a multi-scale representation learning model using grid cell insights.
result Space2Vec outperforms baselines in predicting POI types and image classification with geo-locations.
We present Listen, Attend and Spell (LAS), a neural network that learns to transcribe speech utterances to characters. Unlike traditional DNN-HMM models, this model learns all the components of a speech recognizer jointly. Our system has two components: a listener and a speller. The listener is a pyramidal recurrent ne…