Single directions are key to deep networks' generalization.
problem Understanding why some deep networks generalize well while others do not.
method Examined the reliance on single directions in deep networks and their relation to generalization performance.
result Reliance on single directions is a good predictor of a network's generalization performance.
Study of curves and surfaces from single-direction projections.
problem Obtaining complete shape information from a single view.
method Theoretical study of differential geometric information from multiple orthogonal projections.
result Formulae for recovering certain information on curves or surfaces from their projections.
Proposes a new semi-parametric framework for batched bandits with covariates.
problem Sequential decision-making with batched feedback and contextual information.
method Batched single-Index Dynamic binning and Successive arm elimination (BIDS) using single-index regression.
result Achieves minimax-optimal rates for nonparametric batched bandits.
The paper presents a method for sound event localization and detection using CRNN models.
problem Sound event localization and detection in complex environments.
method Consecutive ensemble of CRNN models for estimating event onset, offset, direction of arrival, and classification.
result The proposed method outperforms other participants in the DCASE2019 task3.
Proposes a multivariate regression model for better analysis of multiple datasets.
problem Insufficient performance of single-dataset analysis in integrative studies.
method Sparse estimation for variable and group selection, alternating direction method of multipliers algorithm.
result Demonstrated improved performance through simulations and real data analysis.
This note demonstrates how both the concept of distance and the concept of holonomy can be constructed from a suitable network with directed edges (and no lengths). The number of different edge types depends on the signature of the metric and the dimension of the holonomy group. If the holonomy group is of dimension on…
MPHL algorithm improves wireless network localization using distance and direction data.
problem Accurate and affordable positioning in wireless networks.
method Hybrid approach combining distance and direction estimates, statistical model, belief propagation, and MCMC sampling.
result Significant reduction in localization error, up to 50% compared to competing algorithms.
Unified routing and arbitrage with concave continuation.
problem Combining routing and arbitrage in financial markets.
method Extending AMM trade functions to negative inputs via concave continuation.
result Unified approach unifies routing and arbitrage.
Neural networks learn faster with correlated latent variables.
problem Efficiently learning from higher-order correlations in neural networks.
method Analytical derivation and simulations of two-layer neural networks.
result Correlations between latent variables speed up learning from higher-order correlations.
Two-layer neural networks learn features through a few gradient descent steps, improving approximation capacity.
problem Improving approximation capacity of two-layer neural networks.
method Theoretical investigation of a two-layer neural network's adaptation to target function through a few gradient descent steps.
result Learning multiple target directions requires a larger batch size and more gradient steps, improving approximation capacity.
Paper discusses directional differentiability of interval-valued functions on Riemannian manifolds.
problem Equivalence of directional differentiability of interval-valued functions and their components.
method Analyzes directional differentiability of interval-valued functions on Riemannian manifolds.
result Directional differentiability of interval-valued functions is not equivalent to the directional differentiability of their components.
DeepCA combines neural networks with component analysis for improved performance.
problem Limited capacity of shallow component analysis in deep learning.
method Deep Component Analysis (DeepCA) with ADNNs for inference.
result Improved performance on various tasks, including depth prediction.
Enhances interpretability of linear latent spaces through automated clustering and ranking.
problem Severe interpretability issues in latent directions of PCA, ICA, CCA, and FA.
method LS-PIE framework automates clustering and ranking of latent vectors.
result Enhanced interpretability of latent vectors through LR, LS, LC, and LCON.
New method uses latent variables to estimate treatment effects from single-arm trials.
problem Estimating treatment effects from single-arm trials due to lack of external control groups.
method Latent-variable modeling with amortized variational inference for patient matching and direct effect estimation.
result Improved performance in direct treatment effect estimation and effect estimation via patient matching compared to previous methods.
New findings show single-treatment effects are unidentifiable in factorial experiments.
problem Identifying the effect of a single intervention in factorial experiments.
method Formalized sufficient conditions for the identifiability of single-treatment effects and developed nonparametric sharp bounds.
result Researchers must justify assumptions for extrapolating single-treatment effects.
Infers causal direction from mixed-type multivariate data using information theory.
problem Inferring causal direction from multivariate and mixed-type data.
method Information theoretic approach based on Kolmogorov complexity and Minimum Description Length (MDL) principle.
result Crack algorithm reliably infers causal direction with high accuracy.
DAGgr aggregates multiple DAGs to stabilize causal structure learning.
problem Stability in learning causal structure from data.
method Model averaging of candidate DAGs weighted by predictive likelihood, with acyclicity enforced.
result DAGgr consistently outperforms individual DAGs and bootstrap-aggregation baselines.
A neural network improves DOA estimation from a single snapshot.
problem Estimating DOAs from a single snapshot with limited aperture.
method Deep learning architecture trained to generate high-resolution spatial spectrum.
result Our (SP)2-Net outperforms classical methods. Framework identifies causal direction from single data setting.
problem Identify causal direction from single observational data.
method VCEI framework based on ICM principle and artificial variation.
result VCEI is competitive to other frameworks in identifying causal direction.
New method constructs tilings of the plane using directed edges and alignments.
problem Modeling tilings of the Euclidean or hyperbolic plane as presheaves over categories.
method Introducing finite categories for polygons with labeled directed edges, constructing reflective alignments.
result Characterizing alignments of tilings by comparing edge directions and generating families with elegant symmetry.
TQF models multivariate uncertainty by learning conditional quantiles.
problem Challenges in fully nonparametric estimation of multivariate conditional distributions.
method Tomographic Quantile Forests (TQF) learns conditional quantiles of directional projections.
result TQF reconstructs multivariate conditional distribution efficiently without convexity restrictions.
New models for causal effect identification without directed cycles.
problem Identifying causal effects in complex graphical models.
method Introduces new graphical models with directed, undirected, and bidirected edges, without cycles. Provides algorithms for identification and learning from data.
result Developed algorithms for identifying causal effects in new models and gated models.
Paper tackles catastrophic overfitting in single-step adversarial training.
problem Catastrophic overfitting leads to sudden drop in robust accuracy.
method Proposes a method to prevent overfitting by using all adversarial examples.
result Demonstrates prevention of catastrophic overfitting and improves robustness.
Study SGD dynamics in sequence models, revealing training phases and influence of sequence length.
problem Understanding SGD in sequence models like attention networks.
method Derived closed-form population loss and analyzed SGD dynamics for SSI models.
result Two distinct training phases: escape from uninformative initialization and alignment with target subspace.
A Bayesian treatment of latent directed graph structure for non-iid data is provided where each child datum is sampled with a directed conditional dependence on a single unknown parent datum. The latent graph structure is assumed to lie in the family of directed out-tree graphs which leads to efficient Bayesian inferen…
A neural network, IHT-Net, improves DOA estimation with sparse arrays.
problem Single-snapshot DOA estimation with sparse arrays in dynamic settings.
method IHT-inspired neural network with recurrent neural network and autoencoders.
result IHT-Net achieves faster convergence and higher accuracy in DOA estimation.
Improved Q&A model with LSTM and bi-directional attention.
problem Enhancing neural network models for effective question answering.
method Implemented a Bi-directional attention flow layer connected to a Multi-layer LSTM encoder, with a new end-index decoder layer conditioning on start-index output.
result Increased model performance by 15.16% on test set.
New algorithm learns sub-task policies from unsegmented demonstrations.
problem Challenges in learning hierarchical policies from unsegmented demonstrations.
method Generative adversarial imitation learning framework with directed information maximization.
result Automatic learning of sub-task policies from unsegmented demonstrations.
The study analyzes how neural reward models learn features for policy optimization in a Gaussian single-index model.
problem Reward modeling in policy optimization and its impact on downstream value.
method Two-stage neural reward model: first learns hidden direction, then fits readout layer.
result For any feature-learning temperature above a dimension-free threshold, a constant fraction of neurons recover the hidden direction.
Single-spike neurons can approximate as well as multi-spike neurons.
problem Limitation of single-spike neurons in spiking neural networks.
method Comparison of single-spike and multi-spike neural networks.
result Single-spike and multi-spike neural networks are equivalent in approximation capabilities.
Proposes a new loss function for better handling mislabeling costs.
problem Handling mislabeling costs in machine learning models.
method Introduces Real-World-Weight Crossentropy loss function.
result Demonstrates improved performance in scenarios of mislabeling.
Bayesian l0-regularized least squares for high-dimensional predictors.
problem Optimizing a non-convex objective function over model space.
method Spike-and-slab priors with single Best Replacement (SBR) for scalability.
result SBR can find the spike-and-slab estimator, bridging Bayesian regularization and proximal updating.
Develops a method to efficiently learn causal DAGs using directed clique trees.
problem Efficiently learning causal DAGs in the presence of large cliques.
method Decomposes DAGs into independently orientable components using directed clique trees and designs a two-phase intervention algorithm.
result Proves that the number of single-node interventions necessary to orient any DAG in an EC is at least the sum of half the size of the largest cliques in each chain component of the essential graph.
We prove that the topology, smooth structure, and metric of a compact Lorentzian manifold with boundary is uniquely determined by data at the boundary. The data consists of the lengths and directions of future-directed once-broken geodesics connecting points on the boundary, which are first timelike and then lightlike.…
New method learns cell trajectories and network interactions from single-cell data.
problem Network inference in systems biology from steady-state data.
method Min-entropy estimation for stochastic dynamics, leveraging both temporal and perturbational data.
result Jointly learns cellular trajectories and network interactions.
Paper improves DOA estimation in sparse arrays using Siamese neural networks.
problem Challenges in DOA estimation with limited snapshots in sparse linear arrays.
method Introduces a Siamese neural network with a sparse augmentation layer for enhanced signal feature embedding.
result Demonstrates improved DOA estimation accuracy in sparse arrays.
Mamba efficiently learns low-dimensional targets in-context via feature extraction.
problem Learning low-dimensional targets in context for computational efficiency.
method Test-time feature learning of a single-index model using Mamba's pretrained linear-time sequence model.
result Mamba achieves efficient in-context learning of low-dimensional targets via feature extraction.
DG improves policy gradients by weighting actions with a sigmoid of advantage and surprisal.
problem Pathologies in standard policy gradients, leading to poor updates and over-allocation of gradient budget.
method Introduces Delightful Policy Gradient (DG) that gates each term with a sigmoid of advantage and surprisal.
result DG provably improves directional accuracy in a single context and shifts the expected gradient closer to the oracle across multiple contexts.
Single-stage neural architecture improves text-to-image synthesis.
problem Combining text and image generation is challenging.
method Used deep residual networks and sentence interpolation.
result Achieved state-of-the-art performance with single-stage training.
Automates selection and visualization of model responses in various directions.
problem Manual selection and limited visualization of PDPs.
method Formalizes method for automating PDP selection and extends to arbitrary directions.
result Demonstrates usefulness across model selection, bias detection, and latent space exploration.
Deep single-index Fréchet regression for metric space-valued outputs
problem Predicting outputs in non-Euclidean spaces
method DeSI (Deep Single-Index Fréchet Regression)
result Interpretable index direction for inputs
A method learns to imitate from a single video demonstration using contrastive training and Siamese networks.
problem Learning to imitate from video demonstrations without direct access to state or action information.
method Contrastive training with Siamese recurrent neural networks to learn rewards and an RL policy to minimize distance.
result The method significantly improves policy learning and outperforms current techniques in various environments.
A new framework uses directed information to efficiently select context chunks.
problem Efficiently selecting relevant context chunks for query understanding.
method Directed Information γ-covering framework, formulated as a γ-cover problem, with a greedy algorithm for context selection. result The γ-covering algorithm provides clear advantages in hard-decision regimes like context compression and single-slot prompt selection. Bidirectional VAE reduces parameters and improves image tasks.
problem Improving image reconstruction, classification, interpolation, and generation.
method Uses a single neural network for both encoding and decoding in both forward and backward directions.
result Bidirectional VAEs reduce parameters by almost 50% and slightly outperform unidirectional VAEs.
New algorithm reduces regret for single-index bandits to nearly optimal.
problem Optimizing rewards from unknown projections of high-dimensional contexts.
method Two-phase algorithm: estimate projection direction, reduce to 1D bandit, use UCB.
result Achieved optimal regret of ildeO(T2/3). Transformers without skip connections collapse token representations to a single direction.
problem Rapid convergence of token representations to a single direction in self-attention-only Transformers.
method Analysis of layer normalization, residual connections, and multi-head attention mechanisms.
result Residual connections prevent rank collapse in real Transformers, while MLPs generate new feature directions.
Improved regret for single-index bandits with optimal algorithm.
problem Optimizing rewards from unknown one-dimensional projections of high-dimensional contexts.
method Two-phase algorithm: first estimate projection direction, then discretize and use UCB.
result Achieved optimal regret of ildeO(T2/3). Direct feedback alignment reduces data movement in neural networks.
problem Efficiency and energy-efficiency in training large neural networks.
method Sparse feedback matrix for local learning, reducing data movement and compute.
result Orders of magnitude improvement in data movement and 2x improvement in multiply-and-accumulate operations.