Two algorithms improve K-means clustering speed without sacrificing quality.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study finds weak solutions for complex map flows with optimal lifespan.
New estimators improve efficiency in two-phase designs with coarsened data.
Improved statistical inference for expensive data using machine learning predictions.
Proposes a method to estimate sparse Gaussian graphical models with hidden clustering structure.
In this paper a data analytical approach featuring support vector machines (SVM) is employed to train a predictive model over an experimentaldataset, which consists of the most relevant studies for two-phase flow pattern prediction. The database for this study consists of flow patterns or flow regimes in gas-liquid two…
Simple DP algorithms find approximate solutions for nonconvex ERM.
Common models for two-phase lipid bilayer membranes are based on an energy that consists of an elastic term for each lipid phase and a line energy at interfaces. Although such an energy controls only the length of interfaces, the membrane surface is usually assumed to be at least across phase boundaries. We consi…
Detects anomalies in product health metrics at eBay for better alerts.
This paper proposes a new global optimization algorithm using deep learning.
Study shows thresholding scheme converges for mean curvature flow of convex sets.
We consider a diffuse interface approximation for the lipid phases of rotationally symmetric two-phase bilayer membranes and rigorously derive its -limit. In particular, we prove that limit vesicles are across interfaces, which justifies a regularity assumption that is widely made in formal asymptotic and nume…
Proposes MD-LiNA for multi-domain latent factor causal discovery.
A new neural network improves the accuracy of predicting constants of motion.
We examine the out-of-equilibrium phase reported by Plerou {\it et. al.} in Nature, {\bf 421}, 130 (2003) using the data of the New York stock market (NYSE) between the years 2001 --2002. We find that the observed two phase phenomenon is an artifact of the definition of the control parameter coupled with the nature of …
SATL adapts to varying smoothness in hypothesis transfer learning.
We present, solve and numerically simulate a simple model that describes the consequences of increased longevity on fertility rates, population growth and the distribution of wealth in developed societies. We look at the consequences of the repeated use of life extension techniques and show that they represent a novel …
For better classification generative models are used to initialize the model and model features before training a classifier. Typically it is needed to solve separate unsupervised and supervised learning problems. Generative restricted Boltzmann machines and deep belief networks are widely used for unsupervised learnin…
We explore a computational model of an incompressible fluid with a multi-phase field in three-dimensional Euclidean space. By investigating an incompressible fluid with a two-phase field geometrically, we reformulate the expression of the surface tension for the two-phase field found by Lafaurie, Nardone, Scardovelli, …
In this article we establish a local parabolic almost monotonicity formula for two phase free boundary problems on Riemannian manifolds, which is an extension of a work of Edquist-Petrosyan.
A two-phase algorithm identifies the best arm in sparse linear bandits with fixed budget.
The two phase behavior in financial markets actually means the bifurcation phenomenon, which represents the change of the conditional probability from an unimodal to a bimodal distribution. In this paper, the bifurcation phenomenon in Hang-Seng index is carefully investigated. It is observed that the bifurcation phenom…
Proposes a novel MTL approach based on bias-variance analysis.
We propose the Insertion-Deletion Transformer, a novel transformer-based neural architecture and training method for sequence generation. The model consists of two phases that are executed iteratively, 1) an insertion phase and 2) a deletion phase. The insertion phase parameterizes a distribution of insertions on the c…
We consider the problem of near-optimal arm identification in the fixed confidence setting of the infinitely armed bandit problem when nothing is known about the arm reservoir distribution. We (1) introduce a PAC-like framework within which to derive and cast results; (2) derive a sample complexity lower bound for near…
CLASSIX is a fast and explainable clustering method that sorts data and merges groups.
Reliable training of generative adversarial networks (GANs) typically require massive datasets in order to model complicated distributions. However, in several applications, training samples obey invariances that are \textit{a priori} known; for example, in complex physics simulations, the training data obey universal …
We consider the problem of efficiently computing the maximum likelihood estimator in Generalized Linear Models (GLMs) when the number of observations is much larger than the number of coefficients (). In this regime, optimization algorithms can immensely benefit from approximate second order information.…
The biological plausibility of the backpropagation algorithm has long been doubted by neuroscientists. Two major reasons are that neurons would need to send two different types of signal in the forward and backward phases, and that pairs of neurons would need to communicate through symmetric bidirectional connections. …
To realize efficient computational fluid dynamics (CFD) prediction of two-phase flow, a multi-scale framework was proposed in this paper by applying a physics-guided data-driven approach. Instrumental to this framework, Feature Similarity Measurement (FSM) technique was developed for error estimation in two-phase flow …
MetaNOR learns common nonlocal kernels for efficient metamaterial modeling.
Symplectic forms from two phase spaces are proven equivalent.
Discrimination-aware classification is receiving an increasing attention in data science fields. The pre-process methods for constructing a discrimination-free classifier first remove discrimination from the training data, and then learn the classifier from the cleaned data. However, they lack a theoretical guarantee f…
DP-SGD can update fewer coordinates while maintaining privacy.
We study the stability of partitions in convex domains involving simultaneous coexistence of three phases, viz. triple junctions. We present a careful derivation of the formula for the second variation of area, written in a suitable form with particular attention to boundary and spine terms, and prove, in contrast to t…
In coronary CT angiography, a series of CT images are taken at different levels of radiation dose during the examination. Although this reduces the total radiation dose, the image quality during the low-dose phases is significantly degraded. To address this problem, here we propose a novel semi-supervised learning tech…
New method tackles convergence issues in approximating FBSDEs.
Efficiently estimates hub graphical models with structured sparsity.
We discuss a class of (local and non-local) theories of gravity that share same properties: i) they admit the Einstein spacetime with arbitrary cosmological constant as a solution; ii) the on-shell action of such a theory vanishes and iii) any (cosmological or black hole) horizon in the Einstein spacetime with a positi…
This review reports some key results in theoretical investigations on configurations of lipid membranes and presents several challenges in this field which involve (i) exact solutions to the shape equation of lipid vesicles; (ii) exact solutions to the governing equations of open lipid membranes; (iii) neck condition o…
In big data image/video analytics, we encounter the problem of learning an overcomplete dictionary for sparse representation from a large training dataset, which can not be processed at once because of storage and computational constraints. To tackle the problem of dictionary learning in such scenarios, we propose an a…
We prove nonexistence of nonconstant local minimizers for a class of functionals, which typically appears in the scalar two-phase field model, over a smooth N-dimensional Riemannian manifold without boundary with non-negative Ricci curvature. Conversely for a class of surfaces possessing a simple closed geodesic along …
New method improves statistical inference using machine learning-imputed data.
Improved online Sinkhorn algorithm for large-scale data processing.
Multi-label classification is an approach which allows a datapoint to be labelled with more than one class at the same time. A common but trivial approach is to train individual binary classifiers per label, but the performance can be improved by considering associations within the labels. Like with any machine learnin…
Study uses DRL with Lagrangian relaxation to solve temporal control tasks with STL constraints.
DR-submodular continuous functions are important objectives with wide real-world applications spanning MAP inference in determinantal point processes (DPPs), and mean-field inference for probabilistic submodular models, amongst others. DR-submodularity captures a subclass of non-convex functions that enables both exact…
Outlier detection is a technique in data mining that aims to detect unusual or unexpected records in the dataset. Existing outlier detection algorithms have different pros and cons and exhibit different sensitivity to noisy data such as extreme values. In this paper, we propose a novel cluster-based outlier detection a…