Consistent spectral clustering with fairness constraints on representation graphs.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Exact partitioning of high-order planted models achieved through convex optimization.
Hypergraph partitioning lies at the heart of a number of problems in machine learning and network sciences. Many algorithms for hypergraph partitioning have been proposed that extend standard approaches for graph partitioning to the case of hypergraphs. However, theoretical aspects of such methods have seldom received …
Study on detecting hierarchical community structures in networks.
FairGP uses graph partitioning to make Graph Transformers fair and scalable.
The paper uses a graph autoencoder to learn unbiased plant-pollinator interaction embeddings.
New insights link diverse statistical problems via secret leakage planted clique.
Graph clustering involves the task of dividing nodes into clusters, so that the edge density is higher within clusters as opposed to across clusters. A natural, classic and popular statistical setting for evaluating solutions to this problem is the stochastic block model, also referred to as the planted partition model…
The problem of detecting communities in a graph is maybe one the most studied inference problems, given its simplicity and widespread diffusion among several disciplines. A very common benchmark for this problem is the stochastic block model or planted partition problem, where a phase transition takes place in the dete…
We consider two closely related problems: planted clustering and submatrix localization. The planted clustering problem assumes that a random graph is generated based on some underlying clusters of the nodes; the task is to recover these clusters given the graph. The submatrix localization problem concerns locating hid…
The paper extends fairness to hierarchical clustering, finding efficient algorithms with minimal loss.
With the recent popularity of graphical clustering methods, there has been an increased focus on the information between samples. We show how learning cluster structure using edge features naturally and simultaneously determines the most likely number of clusters and addresses data scale issues. These results are parti…
Given the widespread popularity of spectral clustering (SC) for partitioning graph data, we study a version of constrained SC in which we try to incorporate the fairness notion proposed by Chierichetti et al. (2017). According to this notion, a clustering is fair if every demographic group is approximately proportional…
The binary symmetric stochastic block model deals with a random graph of vertices partitioned into two equal-sized clusters, such that each pair of vertices is connected independently with probability within clusters and across clusters. In the asymptotic regime of and for fixe…
Graph energy helps detect communities in networks better than traditional methods.
Fair HAC algorithms ensure clustering fairness across protected groups.
Proposes BA method for unbiased time series anomaly detection evaluation.
New algorithm ensures fairness without sacrificing accuracy.
Spectral algorithms solve optimal community detection and related problems.
The paper uses VGG-19 for plant species classification from leaf images.
Scalable method for regionalizing and extracting temporal patterns from time series data.
This paper tackles fair same-day delivery service by optimizing regional service rates.
Tackles the computational hardness of HPC detection, conjecturing equivalence to PC detection.
Plants monitor their surrounding environment and control their physiological functions by producing an electrical response. We recorded electrical signals from different plants by exposing them to Sodium Chloride (NaCl), Ozone (O3) and Sulfuric Acid (H2SO4) under laboratory conditions. After applying pre-processing tec…
We consider the Degree-Corrected Stochastic Block Model (DC-SBM): a random graph on nodes, having i.i.d. weights (possibly heavy-tailed), partitioned into asymptotically equal-sized clusters. The model parameters are two constants and the finite second moment of the weights $Φ^{…
A new algorithm for fair decision-making in bandit problems with biased feedback.
We address the classical problem of hierarchical clustering, but in a framework where one does not have access to a representation of the objects or their pairwise similarities. Instead, we assume that only a set of comparisons between objects is available, that is, statements of the form "objects and are more …
Study connects database alignment and planted matching using Gaussian features.
Federated learning can propagate bias from a few parties to all participants.
Spectral clustering is a fast and popular algorithm for finding clusters in networks. Recently, Chaudhuri et al. (2012) and Amini et al.(2012) proposed inspired variations on the algorithm that artificially inflate the node degrees for improved statistical performance. The current paper extends the previous statistical…
Chemical plants are complex and dynamical systems consisting of many components for manipulation and sensing, whose state transitions depend on various factors such as time, disturbance, and operation procedures. For the purpose of supporting human operators of chemical plants, we are developing an AI system that can s…
Hydropower plants are one of the most convenient option for power generation, as they generate energy exploiting a renewable source, they have relatively low operating and maintenance costs, and they may be used to provide ancillary services, exploiting the large reservoirs of available water. The recent advances in In…
New method finds balanced clusters in graphs using auxiliary information.
Paper tackles clustering with ordinal comparisons, achieving near-optimal results.
Random Planted Forest interprets tree-based models by keeping some splits, leading to more interpretable predictions.
Understanding the adaptation process of plants to drought stress is essential in improving management practices, breeding strategies as well as engineering viable crops for a sustainable agriculture in the coming decades. Hyper-spectral imaging provides a particularly promising approach to gain such understanding since…
We consider the problem of online learning of optimal control for repeatedly operated systems in the presence of parametric uncertainty. During each round of operation, environment selects system parameters according to a fixed but unknown probability distribution. These parameters govern the dynamics of a plant. An ag…
Study information limits for community detection in sub-hypergraphs.
Power plant is a complex and nonstationary system for which the traditional machine learning modeling approaches fall short of expectations. The ensemble-based online learning methods provide an effective way to continuously learn from the dynamic environment and autonomously update models to respond to environmental c…
Realised pay-offs for discretisation-invariant swaps are those which satisfy a restricted `aggregation property' of Neuberger [2012] for twice continuously differentiable deterministic functions of a multivariate martingale. They are initially characterised as solutions to a second-order system of PDEs, then those pay-…
Paper explores limits of high-order clustering with planted structures.
Develops a new Gaussian process method for efficient Bayesian inference of plant root parameters in the Richards equation.
Interpretable neural network for plant traits and species identification.
A neural collaborative filtering method predicts corn hybrid yield performance.
This paper develops several average-case reduction techniques to show new hardness results for three central high-dimensional statistics problems, implying a statistical-computational gap induced by robustness, a detection-recovery gap and a universality principle for these gaps. A main feature of our approach is to ma…
Fault detection in industrial plants is a hot research area as more and more sensor data are being collected throughout the industrial process. Automatic data-driven approaches are widely needed and seen as a promising area of investment. This paper proposes an effective machine learning algorithm to predict industrial…
We prove lognormal distribution for symmetric perceptron model, solving key conjectures.
Study shows it's impossible to count communities without finding them.