Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

295988117 · Jun 202019922001200920182026
48 results for weighted-data clustering

New EM algorithms for weighted-data clustering improve audio-visual scene analysis.

problem Improving clustering of weighted data in heterogeneous environments.
method Proposed weighted-data Gaussian mixture model and two EM algorithms.
result Validation shows improved clustering in audio-visual scenes.

New method speeds up neural network training by preprocessing weight-data correlation.

problem Slow neural network training due to high time complexity.
method Stores weight-data correlation in a tree structure for quick detection of firing neurons.
result Achieves o(nmd)o(nmd) time per iteration with only O(nmd)O(nmd) preprocessing time.

New ICA method improves on existing techniques.

problem Finding independent components in data.
method Multiple-weighted Independent Component Analysis (MWeICA) based on approximate diagonalization of weighted covariance matrices.
result MWeICA achieves better results than state-of-the-art ICA methods with similar computational time.

As a consequence of the strong and usually violated conditional independence assumption (CIA) of naive Bayes (NB) classifier, the performance of NB becomes less and less favorable compared to sophisticated classifiers when the sample size increases. We learn from this phenomenon that when the size of the training data …

2014-12-21abs ↗pdf ↗

Derives continuum model from discrete ε\varepsilon-graphs with connectivity functional.

problem Modeling diffusion in networks with varying connectivity.
method Energy-based continuum limit derivation, neural-network reconstruction of connectivity.
result Error between discrete and continuum energies is O(ε)O(\varepsilon), valid even with fluctuations.

Framework improves gradient estimation for faster training convergence.

problem Efficiently estimating noisy gradients in stochastic optimization.
method Dynamic adaptive importance sampling combining multiple distributions.
result Adaptively weighted multiple importance sampling yields superior gradient estimates.

Novel neural likelihood ratio estimation for negative data in particle physics.

problem Estimating likelihood ratios with negative probability densities and weights.
method Introducing a novel loss function and a new model architecture based on signed mixture models.
result Demonstrated improved estimation on a real-world example from particle physics.

Optimizes weights for better model performance in shifting data.

problem Improper importance weighting leads to poor model performance in data shifts.
method Interprets weights as a bias-variance trade-off and optimizes them simultaneously with model parameters.
result Optimizing weights significantly improves model generalization performance.

New method learns to weight unlabeled data in semi-supervised learning.

problem Equal weighting of all unlabeled data in semi-supervised learning.
method Adjust weights for each unlabeled example using influence function.
result Technique outperforms state-of-the-art methods on image and language classification tasks.

The paper classifies a space of generalized cusps and its moduli.

problem Classifying the moduli space of generalized cusps.
method Generalized cusp classification, representation theory, and geometric structures.
result The moduli space of generalized cusps is homeomorphic to a subspace of conjugacy classes of representations.

Unweighted matrix factorization can match or outperform weighted methods in recommender systems.

problem Improving recommendation performance with matrix factorization on implicit feedback data.
method Systematic study of various weighting schemes and matrix factorization algorithms.
result Training with unweighted data can perform comparably to, and sometimes outperform, training with weighted data.

Gradient descent with random weights in linear regression analyzed for various noise types.

problem Analyzing the impact of random noise on gradient descent in linear regression.
method Gradient descent with randomly weighted data points, various weighting distributions, geometric moment contraction.
result Characterization of implicit regularization and non-asymptotic convergence bounds.

A new ranking algorithm learns data affinity and ranking scores simultaneously.

problem Retrieving similar objects in large databases is challenging.
method Proposes a ranking algorithm that learns data affinity and ranking scores simultaneously, using adaptive neighbors and smoothness constraints.
result The proposed algorithm outperforms existing methods in synthetic and real datasets.

A new method for efficient Gaussian process regression reduces complexity and improves scalability.

problem Efficient Gaussian process regression for large datasets.
method Learnable coreset-based variational inference for Gaussian processes.
result CVGP reduces the dimensionality of the variational parameter search space to linear complexity.

This paper finds a new way to compress CNN weights, improving on pruning and quantization.

problem Improving performance and storage efficiency of CNNs.
method Identifying and exploiting repeated patterns in CNN weight tensors, using Huffman coding and block sparse matrix formats.
result Achieved compaction ratios of 1.4x to 3.1x in addition to pruning and quantization.

This paper recovers input data from transformer models using attention weights.

problem Recovering input data from transformer models for security and privacy concerns.
method Introducing an algorithm to minimize the loss function between expected and actual outputs of transformers.
result The algorithm successfully recovers input data from attention weights and outputs of transformers.

This paper proposes an automatic neural network compression method.

problem Reducing resource requirements for deep neural networks on resource-constrained devices.
method Jointly prunes and quantizes neural networks without manual hyper-parameter tuning.
result Significant reduction in model size with minimal accuracy loss.

CF-GPS learns policies from logged data by considering counterfactual outcomes.

problem Learning policies from limited real experience in complex environments.
method Assumes logged real experience and models counterfactual outcomes. Uses structural causal models for evaluation.
result Improves policy evaluation and search results on a grid-world task.

Proposes a method to predict cluster number and cluster representatives using cluster stability analysis.

problem Determining the number of clusters in a dataset.
method Analyzes cluster stability using Monte-Carlo simulation to predict cluster number and find cluster representatives.
result Significant improvement in predicting cluster numbers and cluster composition in large datasets.

A new matrix factorization model learns and weights data deviations for better model performance.

problem Stochastic noise causes unreliable data points, leading to suboptimal model fitting.
method Deviation-driven matrix factorization model that learns and weights data deviations.
result Our model outperforms state-of-the-art models in accuracy and efficiency.

Sparse Convex Clustering improves clustering performance in high-dimensional data.

problem Distortion in convex clustering performance with uninformative features.
method Introduces Sparse Convex Clustering with an adaptive group-lasso penalty and a tuning criterion based on clustering stability.
result Demonstrates improved clustering performance through feature selection.

Unified clustering comparison framework for overlapping and hierarchical structures.

problem Critical biases in existing clustering comparison measures.
method Element-centric framework comparing relationships induced by cluster structure.
result Framework does not suffer from biases and provides unique insights.

Mode clustering is a nonparametric method for clustering that defines clusters using the basins of attraction of a density estimator's modes. We provide several enhancements to mode clustering: (i) a soft variant of cluster assignment, (ii) a measure of connectivity between clusters, (iii) a technique for choosing the …

2014-06-06abs ↗pdf ↗

This paper introduces a persistence metric to compare clustering solutions with different numbers of clusters.

problem Determining the true number of clusters in a dataset when prior knowledge is lacking.
method The paper introduces a persistence metric based on the maximum over two-norms of all cluster-covariance matrices.
result The persistence metric accurately identifies clustering solutions with the true number of clusters.

In many practical applications of clustering, the objects to be clustered evolve over time, and a clustering result is desired at each time step. In such applications, evolutionary clustering typically outperforms traditional static clustering by producing clustering results that reflect long-term trends while being ro…

2011-04-11abs ↗pdf ↗

Study examines how cluster number affects short-text clustering, introducing a stability metric.

problem Challenges in finding meaningful clusters in short-text data.
method Introduces a stability metric to determine cluster robustness and visualizes cluster subdivisions.
result Choosing a cluster number involves balancing informativeness and complexity, not seeking a single 'optimal' solution.