Improved Johnson-Lindenstrauss lemma for preserving distances in Euclidean space.
problem Preserving distances in Euclidean space for arbitrary sets.
method Tighter analysis of embedding recipe from [MMMR18].
result Stronger version of Johnson-Lindenstrauss lemma with improved dimensionality.
Improves retrieval accuracy for hierarchical documents, especially for distant matches.
problem Limited expressive power of dual encoder models in hierarchical retrieval.
method Proves feasibility of DEs for HR, introduces pretrain-finetune recipe to improve long-distance retrieval.
result Pretrain-finetune boosts recall on long-distance pairs from 19% to 76%.
We present a uniform description of SU(3)-structures in dimension 6 as well as G2-structures in dimension 7 in terms of a characterising spinor and the spinorial field equations it satisfies. We apply the results to hypersurface theory to obtain new embedding theorems, and give a general recipe for bu…
A recipe recommendation system suggests missing ingredients using collaborative filtering.
problem Encouraging healthy diets through personalized ingredient suggestions.
method Item-based collaborative filtering applied to a sparse dataset of recipes.
result Best method achieves a recall@10 of circa 40%.
The paper constructs unbounded symplectic embeddings of rational homology balls into surfaces.
problem Bounding symplectic embeddings of rational homology balls into surfaces.
method Using Mori's theory of flips and mutations of polygons.
result Unbounded sequences of symplectically embedded rational homology balls into surfaces.
Transformers use ReLUs to approximate softmax efficiently.
problem Analyzing resource usage in softmax transformer models.
method Translating ReLU approximation results to softmax attention mechanisms.
result Economic resource bounds for softmax attention mechanisms.
Probabilistic embeddings improve speaker diarization accuracy.
problem Improving speaker diarization accuracy using embeddings.
method Extracting x-vectors and precision matrices from speech segments, interfacing with PLDA model, applying agglomerative clustering, joint training of PLDA and extractor.
result Joint training of PLDA and probabilistic x-vector extractor yields accuracy gains.
Improved robustness in ASR systems with speaker adaptation.
problem Improving robustness in automatic speech recognition systems.
method Weighted-Simple-Add method for adding weighted speaker information vectors to the conformer-based acoustic model.
result Achieved 11% relative improvement in WER on Switchboard 300h Hub5'00 dataset.
Proposes EVE for efficient exploration in reinforcement learning.
problem Efficient exploration in reinforcement learning.
method EVE: a recipe for posterior over parameters, facilitating efficient exploration.
result Competitive performance on benchmarks, efficient exploration confirmed.
Proposes a new algorithm to estimate invariant subspaces across multilayer networks.
problem Estimating invariant subspaces across heterogeneous multiple networks.
method Bias-corrected joint spectral embedding algorithm that recursively calibrates diagonal bias and iteratively updates the subspace estimator.
result Established entrywise subspace perturbation bound and entrywise eigenvector central limit theorem for the algorithm.
Simplified Variational Bayes for easier inference.
problem Complex derivation of Variational Bayes.
method 3-step recipe to identify posterior form and directly write updates.
result Easier, faster, shorter derivation of Variational Bayes.
The EM training algorithm of the classical i-vector extractor is often incorrectly described as a maximum-likelihood method. The i-vector model is however intractable: the likelihood itself and the hidden-variable posteriors needed for the EM algorithm cannot be computed in closed form. We show here that the classical …
In a given 4d spacetime bakcground, one can often construct not one but a family of distinct N=2 string theories. This is due to the multiple ways N=2 superconformal algebra can be embedded in a given worldsheet theory. We formulate the principle of obtaining different physical theories by gauging different embeddings …
Dynamic model improves static economics by incorporating time effects.
problem Static economics overlooks time-dependent phenomena, limiting model accuracy.
method Signals-based approach to reinterpret microeconomic theory, using utility function.
result Dynamic models provide better comparisons with empirical observations.
We give a simple explicit construction of pseudo-Anosov mapping classes using an improvement of the homological criterion of Casson-Bleiler.
Improved graph attention model for noisy graphs.
problem Understanding and improving graph attention in noisy graphs.
method Proposes SuperGAT, a self-supervised graph attention network.
result SuperGAT learns more expressive attention by encoding edges.
We give a recipe for constructing families of distinct knots that have identical Khovanov homology and give examples of pairs of prime knots, as well as infinite families, with this property.
Geometrically computes superpotentials for certain 4D N=2 theories.
problem Computing effective twisted superpotentials for 4D N=2 theories.
method Spectral networks and abelianization to compute generating functions of brane opers.
result Geometric recipe for computing effective twisted superpotentials.
KitcheNette predicts and recommends food ingredient pairings.
problem Limited study of food ingredient pairings despite many existing pairings.
method Siamese neural networks trained on a dataset of 300K scores.
result KitcheNette outperforms other models and discovers novel pairings.
Many recent Markov chain Monte Carlo (MCMC) samplers leverage continuous dynamics to define a transition kernel that efficiently explores a target distribution. In tandem, a focus has been on devising scalable variants that subsample the data and use stochastic gradients in place of full-data gradients in the dynamic s…
New graph foundation models respect symmetries for broader applicability.
problem Tailored graph machine learning architectures limit broader applicability.
method Investigates symmetries for label and feature permutations, proving network universal approximator.
result Universal approximator on multisets respecting node and feature permutations.
Improves model accuracy for neural nets in stochastic dynamics with partial prior knowledge.
problem Stability and accuracy in neural nets modeling stochastic dynamics with many parameters.
method Three steps: probabilistic weights, partial knowledge incorporation, and PAC-Bayesian training.
result Improved model fit with partial and noisy prior knowledge.
A new dynamical method calculates shear-bend coordinates for surfaces.
problem Computing shear-bend coordinates for twisted SL2C local systems.
method Dynamics-based approach to abelianization of local systems.
result A dynamical recipe for shear-bend parameterization.
We give a recipe to compute the geometric intersection number of an integral lamination with a particular type of integral lamination on an n-times punctured disk. This provides a way to find the geometric intersection number of two arbitrary integral laminations when combined with an algorithm of Dynnikov and Wiest.
LeDeepChef learns to play multiple cooking games well.
problem Designing a general RL agent for multiple games of the same family.
method Actor-critic framework, action-space pruning, hierarchical RL, specialized module.
result LeDeepChef outperformed competitors on a diverse set of cooking games.
GNMR controls runtime stability in low-precision language model training.
problem Efficient low-precision training faces numerical risks at specific operators.
method GNMR compares gradient norms to historical means, applying bounded recovery actions.
result GNMR preserves high-fidelity quality with sparse, budgeted recovery.
Modified RV-coefficient reveals how training affects neural network representations.
problem Understanding how training affects intermediate representations in convolutional neural networks.
method Experimented with modified RV-coefficient (RV2) to compare activation patterns in deep networks trained on varying amounts of data and layers.
result RV2 successfully recovered expected similarity patterns and provided interpretable similarity matrices.
We list up all the possible local orbit types of hyperbolic or elliptic orbits for the isotropy representations of semisimple pseudo-Riemannian symmetric spaces. It is key to give a recipe to determine the local orbit types of hyperbolic principal orbits by using three kind of restricted root systems and Satake diagram…
Unified framework for ranking-and-selection with multiple correct answers and non-answerable estimates
problem Fixed-precision ranking-and-selection in structured settings with non-unique answers and non-answerable estimates
method Unified framework based on answer-wise acceptance sets, restricted generalized likelihood ratio stopping, and answer-pitfall decomposition
result Unified recipe performs well across a broad range of pure-exploration problems
This paper develops tools for nonreversible MCMC with convergence guarantees.
problem Designing nonreversible MCMC kernels with convergence guarantees.
method Develops tools for nonreversible Markov kernels using conditional invertible transforms.
result Ensures nonreversible kernels have the desired invariance property and lead to convergent algorithms.
DecompKAN improves time series forecasting accuracy and transparency.
problem Accurate and transparent time series forecasting in scientific domains.
method Combines decomposition, patching, normalization, and B-spline KAN edge functions.
result Achieves best or tied-best MSE on 20 of 36 comparisons across 9 datasets.
Enhanced trivalent tangles and handlebody-tangles invariants created.
problem Creating invariants for trivalent and handlebody-tangles.
method Using enhanced trivalent tangles and classical knot theory.
result Constructed invariants for trivalent and handlebody-tangles.
Bayesian SPLDA adapts model parameters for database transfer.
problem Adapting SPLDA for new databases with limited data.
method Variational Bayes estimation of SPLDA parameters.
result Adaptation of SPLDA model parameters for database transfer.
Alternative method proposed for handling uncertainty in i-vector extraction.
problem Uncertainty in i-vector extraction for spoken language recognition.
method Proposes an alternative method to propagate uncertainty into the Gaussian back-end.
result Alternative method effectively handles uncertainty in i-vector extraction.
We prove that the so-called t algebra of braids and ties supports a Markov trace. Further, by using this trace in the Jones' recipe, we define invariant polynomials for classical knots and singular knots. Our invariants have three parameters. The invariant of classical knots is an extension of the Homflypt polynomial a…
This paper tackles multilingual speech processing by optimizing conflicting objectives hierarchically.
problem Training models for multilingual, multi-task speech processing is hampered by conflicting objectives.
method Investigates three multi-objective MSP formulations and introduces a lightweight layer-selection mechanism.
result A bi-level recipe outperforms standard flat optimization in state-of-the-art MSP models.
Unified theory of measure-preserving diffusions on manifolds.
problem Deriving a complete recipe for measure-preserving diffusions on manifolds.
method Developed a geometric theory that unifies and generalizes previous constructions, relying on intrinsic geometry of the target measure.
result The completeness result is a direct consequence of manifold topology and target measure geometry.
Q-chunking improves RL for long tasks by chunking actions.
problem Improving sample efficiency and exploration in offline-to-online RL.
method Action chunking in TD-based RL methods.
result Q-chunking outperforms prior methods on long-horizon tasks.
A new method for diffusion generative models improves sample quality and speed.
problem Improving the efficiency and quality of diffusion generative models.
method Developed a complete recipe for forward diffusion processes in SGMs, introducing PSLD.
result PSLD achieves superior sample quality and speed compared to existing methods.
New framework predicts AMP behavior in spiked models for finite iterations.
problem Understanding AMP dynamics in high-dimensional spiked models.
method Developed a non-asymptotic framework for AMP in spiked matrix estimation.
result Predicted AMP behavior for up to O(polylognn) iterations in Z2 synchronization. Paper proposes tensor-based method for semiconductor manufacturing process control.
problem Challenges of traditional process control methods in high-dimensional image-based overlay errors.
method Builds a high-dimensional process model, proposes tensor-on-vector regression algorithms, designs EWMA controller for tensor data.
result The method reduces overlay errors using limited control recipes and is superior especially when disturbances are not stable.
New method trains deep networks robustly without adaptive methods.
problem Training deep networks with robustness and efficiency.
method Scale invariant architecture + SGD + weight decay + gradient clipping.
result SGD can achieve similar performance to adaptive methods like Adam.
This paper considers the stability of online learning algorithms and its implications for learnability (bounded regret). We introduce a novel quantity called {\em forward regret} that intuitively measures how good an online learning algorithm is if it is allowed a one-step look-ahead into the future. We show that given…
Researchers create exact minimal surfaces with helical motifs in biological structures.
problem Analyzing helical motifs in minimal surfaces of biological structures.
method Developed a method to construct exact minimal surfaces with arbitrary helical motifs.
result Exact minimal surfaces with helical motifs can be created and analyzed.
Mixed-precision CA-SGD for generalized linear models on GPUs
problem SGD communication bottleneck
method Mixed-precision CA-SGD
result Matches FP32 SGD loss within 0.5% on various problems
Predicts food ingredient amounts from images.
problem Predicting relative amounts of ingredients from food images.
method Proposes two deep learning models for sparse and dense predictions, with semi-automatic data pre-processing.
result Encouraging experimental results on a recipe dataset.
Given any generating set of any pseudo-Anosov-containing subgroup of the mapping class group of a surface, we construct a pseudo-Anosov with word length bounded by a constant depending only on the surface. More generally, in any subgroup G we find an element f with the property that the minimal subsurface supporting a …
Reinforcement learning after next-token prediction aids in learning from diverse sequence lengths.
problem Learning from sequences of varying lengths and complexity.
method Introducing a framework to study reinforcement learning with autoregressive transformers, focusing on next-token prediction and mixture distributions of short and long sequences.
result Reinforcement learning after next-token prediction enables autoregressive transformers to generalize from long sequences, even when they are rare.