New metric addresses issues in hierarchical clustering.
problem Hierarchical clustering trees can be inconsistent with true limits.
method Introduce merge distortion metric to quantify tree differences.
result Merge distortion metric implies separation and minimality.
The paper tightens bounds on distances between Reeb graphs.
problem Certifying quasi-universality of distances between Reeb graphs.
method Establishes tight bi-Lipschitz bounds for various distances.
result Proves strict universality of the functional contortion distance for contour trees and coincides with interleaving distance for merge trees.
LEWIS merges LLMs without training, improving performance on specific tasks.
problem Limited performance improvement of merged models on specific benchmarks.
method Guided model merging using layer-wise sparsity and task-vector pruning.
result Improved model performance by up to 11.3% on math-solving tasks.
New method separates graph structure from node attributes to recover lost signal.
problem Standard representation learning on attributed graphs merges incompatible metric spaces, leading to geometrically flawed alignment.
method Custom variational autoencoder that separates manifold learning from structural alignment.
result Transforms geometric conflict into interpretable structural descriptor, uncovering connectivity patterns and anomalies.
This paper merges deterministic policy gradient estimations to improve deep reinforcement learning performance.
problem The bias-variance tradeoff in estimating and using policy gradients for deep reinforcement learning.
method Introduces elite policy gradients and a two-step merging method to balance bias-variance tradeoffs.
result Two-step merging outperforms interpolation merging and state-of-the-art algorithms on benchmark control tasks.
Study shows training duration impacts model merging quality, suggesting joint selection of duration and method.
problem Impact of expert training duration on model merging quality for large language models (LLMs).
method Systematically fine-tuned experts on five domains across three model sizes, evaluating five merging methods at each duration.
result Training duration affects merging quality, with simple averaging degrading sharply and sparsification-based methods performing well past the validation optimum.
Study shows training duration affects model merging quality, suggesting joint selection of duration and method.
problem Impact of expert training duration on model merging quality for large language models (LLMs).
method Systematically fine-tuned experts on five domains across three model sizes, evaluated five merging methods at each duration.
result Training duration and merging method should be chosen jointly, not independently.
EpiMer merges models by solving Fréchet mean on a Riemannian manifold.
problem Integrating knowledge from multiple models without retraining.
method EpiMer casts model merging as solving the Fréchet mean on a Riemannian manifold, restricting computation to a low-rank subspace.
result EpiMer outperforms flat-geometry methods on image classification tasks.
New Bayesian method improves Pareto front estimation in multitask finetuning.
problem Efficiently estimating Pareto fronts for multitask finetuning.
method Variational Model Merging using non-Gaussian posteriors.
result More flexible posteriors lead to better Pareto front estimates.
Study on merging predictors in causal and anticausal directions using CMAXENT.
problem Comparing merging predictors in causal and anticausal directions.
method Using CMAXENT as inductive bias, study differences in merging predictors.
result CMAXENT solution reduces to logistic regression in causal direction and LDA in anticausal direction.
In this paper, a similarity-driven cluster merging method is proposed for unsuper-vised fuzzy clustering. The cluster merging method is used to resolve the problem of cluster validation. Starting with an overspecified number of clusters in the data, pairs of similar clusters are merged based on the proposed similarity-…
New method merges MCMC samples without distributional assumptions.
problem Efficiently merging MCMC samples from disjoint subsets.
method Diffusion generative modelling for density approximation.
result Outperforms existing methods on high-dimensional problems.
Efficiently merges multiple points to speed up BSGD SVM training.
problem Costly merging of points in BSGD SVM training.
method Merges more than two points at once to reduce training time.
result Significant speed-ups achieved without loss of accuracy.
Securely evaluates the benefits of merging datasets for causal estimation.
problem Challenges in assessing the value of merging datasets for causal treatment effect estimation.
method Cryptographically secure multi-party computation to evaluate Expected Information Gain (EIG) while ensuring privacy.
result Demonstrates the first privacy-preserving method for dataset acquisition tailored to causal estimation.
A new method merges neural networks using CCA to improve model performance.
problem Improving model accuracy through ensembling while reducing computational and storage costs.
method CCA Merge, a new model merging algorithm based on Canonical Correlation Analysis.
result CCA Merge leads to better model performance than past methods, especially in merging more than two models.
This paper reverses a construction by merging boundary critical points into an interior one.
problem Pushing interior critical points to the boundary and splitting them into two boundary points.
method Specific assumptions allow merging two boundary critical points into one interior critical point.
result Merging two boundary critical points into a single interior critical point.
NAMEx merges experts using Nash bargaining for improved performance.
problem Sparse Mixture of Experts merging strategies lack a principled weighting mechanism.
method Reinterpreting expert merging through game theory, introducing Nash Merging and complex momentum.
result NAMEx consistently outperforms competing methods across various tasks and system sizes.
Efficiently estimates longitudinal networks by merging sparse networks.
problem Estimating longitudinal networks with sparse and temporal data.
method Adaptive network merging, tensor decomposition, point process.
result Significantly reduces estimation error and provides guidance for network merging.
IDAS approach for autonomous vehicles to make decisions under merging scenarios.
problem Decision making for autonomous vehicles in merging scenarios with varying driver cooperativeness.
method IDAS approach using multi-agent reinforcement learning (MARL) with curriculum learning and masking mechanism.
result IDAS approach can handle uncertainties in real-world scenarios and make strategic decisions.
A new method reduces task interference in model merging.
problem Model merging overlooks structural information and is susceptible to task interference.
method Task Singular Vectors (TSV) and TSV-Compress for compression and interference reduction.
result TSV-Merge significantly outperforms existing methods.
Discrete knot theory models use lattice-filtered graphs to detect merging knot components.
problem Detecting merging knot components in discrete models.
method Lattice-filtered move graphs to model knot types, identifying connected components and merge scales.
result Merge scale defined by connected components of lattice-filtered move graphs, with specific examples for the figure-eight knot.
Analyzes merging vs. ensembling for multi-study prediction, showing transition point for better performance.
problem Choosing between merging or ensembling multiple studies for prediction.
method Analyzes ridge regression approaches, comparing merging and ensembling methods.
result There is a transition point where ensembling outperforms merging as cross-study heterogeneity increases.
Boosting strategies for merging vs. ensembling studies analyzed.
problem Deciding between merging and ensembling studies for boosting.
method Analytical transition point and bias-variance decomposition for boosting with linear learners.
result Theoretical guidelines for merging vs. ensembling studies.
Method constructs finance LLMs without instruction data using pretraining and model merging.
problem Developing domain-specific LLMs for finance is resource-intensive.
method Continual pretraining on financial data + model merging of instruction-tuned and domain-specific pretrained vectors.
result Successfully constructs instruction-tuned LLMs for finance without additional instruction data.
LiteMORT reduces memory usage for GBDT models by 30% with improved accuracy.
problem Memory inefficiency in GBDT models, especially with adaptive compact distributions.
method Share memory, implicit merging, adaptive histogram resizing.
result Significant reduction in memory usage with improved accuracy.
New split-merge MCMC proposals improve efficiency and speed.
problem Scaling issues in split-merge MCMC for large datasets.
method Locality Sensitive Sampling (LSS) combined with weighted MinHash.
result Significantly faster than state-of-the-art methods on large datasets.
Single global merging boosts decentralized learning performance.
problem Limited communication in decentralized learning hinders performance.
method Scheduled communication, focusing on final step with global merging.
result Single global merging improves global test performance.
Vertex distortion detects if a knot is unknot.
problem Determining if a knot is the unknot.
method Using Denne-Sullivan's bound on Gromov distortion, the vertex distortion of nontrivial lattice knots is bounded. Then, it is shown that trivial vertex distortion implies the unknot.
result The conjecture that trivial vertex distortion implies the unknot is proven.
Parallel neural networks estimate TVD for merging over-clustered datasets.
problem Merging over-partitioned clusters in unsupervised learning.
method Use neural networks to estimate TVD between clusters in parallel.
result Neural network estimates of TVD lead to better merge decisions.
Distorted surfaces in graph manifolds have specific distortion properties.
problem Distortion of surfaces in graph manifolds.
method Analysis of immersed horizontal surfaces in 3D graph manifolds.
result Fundamental group of surfaces is quadratically distorted if virtually embedded, exponentially distorted otherwise.
Deep learning merges dataset columns with similar entities.
problem Joining datasets with different entity representations.
method Creating deep learning models to map surface forms into vectors, indexing for nearest neighbors.
result Models achieve precision@1 of .75-.81 and recall of .74-.81.
DPSM clusters nodes in data and graph spaces via density propagation and subcluster merging.
problem Automatic clustering of nodes in data and graph spaces.
method Density-based node clustering with propagation process and spectral clustering on subclusters.
result DPSM effectively clusters nodes in both data and graph spaces.
MergeNet detects morphological errors in 3D neuron segmentations.
problem High incidence of merge errors in deep learning 3D connectomics.
method Unsupervised training of MergeNet on various datasets.
result MergeNet can detect morphological errors in neuronal shapes.
Bayesian Federated Inference improves survival model analysis without merging data.
problem Accurately estimating survival model parameters requires sufficient data and events, which is often lacking in practice.
method Bayesian Federated Inference (BFI) for survival models, where local centers perform analyses and combine results.
result Results from BFI are similar to those from merged data analyses, demonstrating excellent performance.
New method corrects complex distortions in single view images.
problem Complex distortions in images, especially those caused by refractive surfaces.
method Differentiable image sampling and semantic information augmentation.
result Model can estimate and correct highly complex distortions.
Algorithm finds optimal affine transformation to minimize overall distortion.
problem Minimizing distortion in affine transformations.
method Riemannian geometry approach to define and minimize distortion.
result Mean distorting transformation found for minimizing overall distortion.
New methods merge discrete gradient fields from patches to correct errors.
problem Correctly merging partially defined discrete gradient fields from patches.
method Developed general and lean merging procedures for specific covering patterns.
result Corrected errors in merging discrete gradient fields from patches.
The hierarchical Dirichlet process (HDP) has become an important Bayesian nonparametric model for grouped data, such as document collections. The HDP is used to construct a flexible mixed-membership model where the number of components is determined by the data. As for most Bayesian nonparametric models, exact posterio…
The article improves the display of acceptable exchange ratios for merging companies.
problem Determining feasible exchange ratios for merging companies.
method Exploits a diagrammatic approach to display the bargaining region.
result Shares face upper and lower bounds for acceptable exchange ratios.
Vertex distortion measures how far lattice knots deviate from straight lines.
problem Measuring how much lattice knots deviate from straight paths.
method Analogous to smooth knots, study vertex distortion in lattice knots.
result Vertex distortion is 1 only for the unknot and can be arbitrarily high.
We consider the problem of distortion minimal morphing of n-dimensional compact connected oriented smooth manifolds without boundary embedded in Rn+1. Distortion involves bending and stretching. In this paper, minimal distortion (with respect to stretching) is defined as the infinitesimal relative change in vol…
NetFuse merges different DNN models with varying weights for faster inference.
problem Inference speed of DNN models with different weights cannot be improved using existing techniques.
method NetFuse merges models with the same architecture but different weights and inputs, replacing operations with more general ones.
result NetFuse can speed up DNN inference time up to 3.6x on a NVIDIA V100 GPU.
This paper shows how to calculate risk measures for sums of two counter-monotonic risks.
problem Calculating risk measures for sums of two counter-monotonic risks.
method Using a fixed distortion function and expressing the risk measure of a sum as the sum of two related measures of the marginals.
result The risk measure of a sum of two counter-monotonic risks can be expressed as the sum of two related distortion risk measures of the marginals.
Finite distortion maps cannot have compact branch sets under growth conditions.
problem Understanding the structure of branch sets in mappings of finite distortion.
method Analyzing the asymptotic growth of distortion and constructing specific examples.
result The bound on the size of branch sets is strict and achievable.
A new family of stochastic dominance orders based on distortion functions.
problem Determining a continuum of dominance relations for risk assessment.
method Introducing H-distorted stochastic dominance, a generalized family of stochastic orders.
result Power-distorted stochastic dominance is particularly appealing due to its simplicity and statistical interpretations.
The distortion of a curve measures the maximum arc/chord length ratio. Gromov showed any closed curve has distortion at least pi/2 and asked about the distortion of knots. Here, we prove that any nontrivial tame knot has distortion at least 5pi/3; examples show that distortion under 7.16 suffices to build a trefoil kno…
Study distortion risk measures for step-weighted distributions.
problem Analyzing risk measures for specific distribution types.
method Investigate distortion risk measures of step-weighted distributions.
result Developed methods for calculating risk measures.
The paper calculates subgroup distortions in 3-manifold groups.
problem Understanding subgroup distortions in 3-manifold groups.
method Computed all finitely generated subgroups of finitely generated 3-manifold groups and analyzed their distortions.
result Subgroup distortions in 3-manifold groups are linear, quadratic, exponential, or double exponential.