Higher granularity in MoE models boosts expressivity exponentially.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This work explores the relationship between expressivity and generalization in GNNs.
Researchers create a star product on a Grassmannian with separation of variables.
Token-adaptive FFN design improves LLM expressivity.
Improves latent variable learning for complex data.
SimCD simultaneously clusters cells and identifies differential gene expression in scRNA-seq data.
Paper explains why robust generalization is hard in deep learning models.
Paper studies the expressivity of Convolutional Neural Networks (CNNs).
While on some natural distributions, neural-networks are trained efficiently using gradient-based algorithms, it is known that learning them is computationally hard in the worst-case. To separate hard from easy to learn distributions, we observe the property of local correlation: correlation between local patterns of t…
Optimal portfolios are found for a wide range of utility functions under hyperbolic returns.
We propose a new blind source separation algorithm based on mixtures of alpha-stable distributions. Complex symmetric alpha-stable distributions have been recently showed to better model audio signals in the time-frequency domain than classical Gaussian distributions thanks to their larger dynamic range. However, infer…
Blind Source Separation (BSS) has proven to be a powerful tool for the analysis of composite patterns in engineering and science. We introduce Convex Analysis of Mixtures (CAM) for separating non-negative well-grounded sources, which learns the mixing matrix by identifying the lateral edges of the convex data scatter p…
We present a scalable Gaussian process model for identifying and characterizing smooth multidimensional changepoints, and automatically learning changes in expressive covariance structure. We use Random Kitchen Sink features to flexibly define a change surface in combination with expressive spectral mixture kernels to …
Localized sum-of-norms clustering separates balls in data.
In this work we show that randomized (block) coordinate descent methods can be accelerated by parallelization when applied to the problem of minimizing the sum of a partially separable smooth convex function and a simple separable convex function. The theoretical speedup, as compared to the serial method, and referring…
Quantum correlations enhance generative models, providing a new resource for machine learning.
Feature selection is an important and challenging task in high dimensional clustering. For example, in genomics, there may only be a small number of genes that are differentially expressed, which are informative to the overall clustering structure. Existing feature selection methods, such as Sparse K-means, rarely tack…
Boosting strategies for merging vs. ensembling studies analyzed.
An audiovisual speaker conversion method is presented for simultaneously transforming the facial expressions and voice of a source speaker into those of a target speaker. Transforming the facial and acoustic features together makes it possible for the converted voice and facial expressions to be highly correlated and f…
New non-separable covariance kernels for spatiotemporal data derived from harmonic oscillator physics.
Convolution Neural Network (CNN) has gained tremendous success in computer vision tasks with its outstanding ability to capture the local latent features. Recently, there has been an increasing interest in extending convolution operations to the non-Euclidean geometry. Although various types of convolution operations h…
The study quantifies how many objects can be linearly classified under all views.
Test assesses if a linear classifier is random or significant.
At critical coupling, the interactions of Ginzburg-Landau vortices are determined by the metric on the moduli space of static solutions. The asymptotic form of the metric for two well separated vortices is shown here to be expressible in terms of a Bessel function. A straightforward extension gives the metric for N vor…
Recently, digital music libraries have been developed and can be plainly accessed. Latest research showed that current organization and retrieval of music tracks based on album information are inefficient. Moreover, they demonstrated that people use emotion tags for music tracks in order to search and retrieve them. In…
Tensors play a central role in many modern machine learning and signal processing applications. In such applications, the target tensor is usually of low rank, i.e., can be expressed as a sum of a small number of rank one tensors. This motivates us to consider the problem of low rank tensor recovery from a class of lin…
Enhanced FastMNMF for better speech separation.
Motivation: Public and private repositories of experimental data are growing to sizes that require dedicated methods for finding relevant data. To improve on the state of the art of keyword searches from annotations, methods for content-based retrieval have been proposed. In the context of gene expression experiments, …
Derives asymptotic generalization error for large-margin classifiers.
TensorGuide improves LoRA efficiency and expressivity through joint tensor-train optimization.
Existing depth separation results for constant-depth networks essentially show that certain radial functions in , which can be easily approximated with depth networks, cannot be approximated by depth networks, even up to constant accuracy, unless their size is exponential in . However, the func…
MPNNs generalize poorly, study reviews current research.
A new method speeds up overlapping group lasso computations.
The article derives some novel independence measures and contrast functions for Blind Source Separation (BSS) application. For the order differentiable multivariate functions with equal hyper-volumes (region bounded by hyper-surfaces) and with a constraint of bounded support for , it proves that equality …
We classify all surfaces with constant Gaussian curvature in Euclidean -space that can be expressed as an implicit equation of type , where , and are real functions of one variable. If , we prove that the surface is a surface of revolution, a cylindrical surface or a conical sur…
Method learns shared and specific factors in multi-study gene expression data.
New methods detect continuous variation in single-cell data.
Hierarchical neural networks are exponentially more efficient than their corresponding "shallow" counterpart with the same expressive power, but involve huge number of parameters and require tedious amounts of training. By approximating the tangent subspace, we suggest a sparse representation that enables switching to …
The likelihood model of high dimensional data can often be expressed as , where is a collection of hidden features shared across objects, indexed by , and is a non-negative factor loading vector with entries where indicates the strength of …
Modern machine learning classifiers often exhibit vanishing classification error on the training set. They achieve this by learning nonlinear representations of the inputs that maps the data into linearly separable classes. Motivated by these phenomena, we revisit high-dimensional maximum margin classification for line…
This work extends alpha-beta divergences to complex data and finds closed-form solutions.
Identifying changes in model parameters is fundamental in machine learning and statistics. However, standard changepoint models are limited in expressiveness, often addressing unidimensional problems and assuming instantaneous changes. We introduce change surfaces as a multidimensional and highly expressive generalizat…
We make two theoretical contributions to disentanglement learning by (a) defining precise semantics of disentangled representations, and (b) establishing robust metrics for evaluation. First, we characterize the concept "disentangled representations" used in supervised and unsupervised methods along three dimensions-in…
Study on size and depth of neural networks for approximating benign functions, showing barriers and explicit results.
A critical decision point when training predictors using multiple studies is whether studies should be combined or treated separately. We compare two multi-study prediction approaches in the presence of potential heterogeneity in predictor-outcome relationships across datasets: 1) merging all of the datasets and traini…
It is notoriously difficult to control the behavior of reinforcement learning agents. Agents often learn to exploit the environment or reward signal and need to be retrained multiple times. The multi-objective reinforcement learning (MORL) framework separates a reward function into several objectives. An ideal MORL age…
Understanding the power of depth in feed-forward neural networks is an ongoing challenge in the field of deep learning theory. While current works account for the importance of depth for the expressive power of neural-networks, it remains an open question whether these benefits are exploited during a gradient-based opt…
This study compares GNNs and GA-MLPs, finding GA-MLPs can distinguish graphs but not count walks.