New loss function calibrates WW-hinge loss for multiclass SVM.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper corrects GIRP algorithm to ensure isotonic models.
Graph partitioning is the problem of dividing the nodes of a graph into balanced partitions while minimizing the edge cut across the partitions. Due to its combinatorial nature, many approximate solutions have been developed, including variants of multi-level methods and spectral clustering. We propose GAP, a Generaliz…
DEEPLY improves cloud service partitioning across multiple datasets and goals.
Paper proposes a new method to learn EBMs and their partition function.
Let G=(V,E) be an undirected graph, lambda_k be the k-th smallest eigenvalue of the normalized laplacian matrix of G. There is a basic fact in algebraic graph theory that lambda_k > 0 if and only if G has at most k-1 connected components. We prove a robust version of this fact. If lambda_k>0, then for some 1\leq \ell\l…
Unsupervised space partitioning improves ANNS performance without pre-processing.
Enhances POU-Nets with probabilistic noise model for efficient spatial data clustering.
Exact partitioning of high-order planted models achieved through convex optimization.
This work proposes a method to optimize hyperparameters without validation data.
In this paper we propose an algorithm for exact partitioning of high-order models. We define a general class of -degree Homogeneous Polynomial Models, which subsumes several examples motivated from prior literature. Exact partitioning can be formulated as a tensor optimization problem. We relax this high-order combi…
Paper develops a high-order recombination algorithm for financial modeling.
The ability to integrate information in the brain is considered to be an essential property for cognition and consciousness. Integrated Information Theory (IIT) hypothesizes that the amount of integrated information () in the brain is related to the level of consciousness. IIT proposes that to quantify information i…
Efficiently calculates PL model likelihood for partitioned preference data.
FairGP uses graph partitioning to make Graph Transformers fair and scalable.
There are two natural simplicial complexes associated to the noncrossing partition lattice: the order complex of the full lattice and the order complex of the lattice with its bounding elements removed. The latter is a complex that we call the noncrossing partition link because it is the link of an edge in the former. …
Paper recovers lattice signal partitions efficiently.
This paper finds a unique partition of a sample space for estimating continuous distributions.
New method for distributed online learning with communication constraints reduces joint regret.
Min-cut clustering, based on minimizing one of two heuristic cost-functions proposed by Shi and Malik, has spawned tremendous research, both analytic and algorithmic, in the graph partitioning and image segmentation communities over the last decade. It is however unclear if these heuristics can be derived from a more g…
The paper examines topological features of ReLU networks and their relation to decision boundaries and training loss.
Solves open problem on universally consistent online learning with unbounded losses.
Multi-view clustering is an important yet challenging task due to the difficulty of integrating the information from multiple representations. Most existing multi-view clustering methods explore the heterogeneous information in the space where the data points lie. Such common practice may cause significant information …
Bayesian approach for multifile record linkage and duplicate detection.
Ehrenborg and Jung recently related the order complex for the lattice of d-divisible partitions with the simplicial complex of pointed ordered set partitions via a homotopy equivalence. The latter has top homology naturally identified as a Specht module. Their work unifies that of Calderbank, Hanlon, Robinson, and Wach…
We propose using five data-driven community detection approaches from social networks to partition the label space for the task of multi-label classification as an alternative to random partitioning into equal subsets as performed by RAkELd: modularity-maximizing fastgreedy and leading eigenvector, infomap, walktrap an…
In this paper we develop and analyze Hydra: HYbriD cooRdinAte descent method for solving loss minimization problems with big data. We initially partition the coordinates (features) and assign each partition to a different node of a cluster. At every iteration, each node picks a random subset of the coordinates from tho…
In this paper, we consider unsupervised partitioning problems, such as clustering, image segmentation, video segmentation and other change-point detection problems. We focus on partitioning problems based explicitly or implicitly on the minimization of Euclidean distortions, which include mean-based change-point detect…
This paper addresses the general problem of modelling and learning rank data with ties. We propose a probabilistic generative model, that models the process as permutations over partitions. This results in super-exponential combinatorial state space with unknown numbers of partitions and unknown ordering among them. We…
We show that the discretized configuration space of points in the -simplex is homotopy equivalent to a wedge of spheres of dimension . This space is homeomorphic to the order complex of the poset of ordered partial partitions of with exactly parts. We compute the exponential generating…
Standard bubbles and partitions are stable in various model spaces.
Centroid-based methods including k-means and fuzzy c-means are known as effective and easy-to-implement approaches to clustering purposes in many applications. However, these algorithms cannot be directly applied to supervised tasks. This paper thus presents a generative model extending the centroid-based clustering ap…
Proposes SPFB method for optimizing partition functions in stochastic learning.
We develop a variant of multiclass logistic regression that is significantly more robust to noise. The algorithm has one weight vector per class and the surrogate loss is a function of the linear activations (one per class). The surrogate loss of an example with linear activation vector and class has t…
We study the effect of a relevant double-trace deformation on the partition function (and conformal anomaly) of a CFT at large N and its dual picture in AdS. Three complementary previous results are brought into full agreement with each other: bulk and boundary computations, as well as their formal identity. We show th…
New examples show non-rotational annuli in a ball, solving a uniqueness problem.
Let φ(G) be the minimum conductance of an undirected graph G, and let 0=λ_1 <= λ_2 <=... <= λ_n <= 2 be the eigenvalues of the normalized Laplacian matrix of G. We prove that for any graph G and any k >= 2, φ(G) = O(k) λ_2 / \sqrt{λ_k}, and this performance guarantee is achieved by the spectral partitioning algorithm. …
In a series of recent works, we have generalised the consistency results in the stochastic block model literature to the case of uniform and non-uniform hypergraphs. The present paper continues the same line of study, where we focus on partitioning weighted uniform hypergraphs---a problem often encountered in computer …
Efficiently resolves entities via scaled Ewens--Pitman model.
Differentially private method for synthetic data generation from vertically partitioned data.
This thesis classifies pseudo-Anosov homeomorphisms using geometric Markov partitions.
In this paper we propose a novel Bayesian methodology for Value-at-Risk computation based on parametric Product Partition Models. Value-at-Risk is a standard tool to measure and control the market risk of an asset or a portfolio, and it is also required for regulatory purposes. Its popularity is partly due to the fact …
We study the problem of estimating a temporally varying coefficient and varying structure (VCVS) graphical model underlying nonstationary time series data, such as social states of interacting individuals or microarray expression profiles of gene networks, as opposed to i.i.d. data from an invariant model widely consid…
Structured entropy improves classification performance on structured targets.
A new learning rule consistently reduces error over data samples.
We present a fully-supervized method for learning to segment data structured by an adjacency graph. We introduce the graph-structured contrastive loss, a loss function structured by a ground truth segmentation. It promotes learning vertex embeddings which are homogeneous within desired segments, and have high contrast …
Study on hypermaps and KP hierarchy, proving tau function and enumerative meaning.
Unified framework improves PCA for outliers and distributed data.