The paper sets thresholds for testing correlation in hypergraphs, distinguishing between independent and correlated states.
problem Testing correlation between two hypergraphs under different models.
method Derives sharp information-theoretic thresholds for distinguishing between null and alternative hypotheses.
result The testing threshold decreases as the hypergraph's uniformity (m) increases, making correlation testing easier for higher uniformity.
STAT-SVD method reduces high-dimensional data sparsity, achieving optimal estimation.
problem Sparse tensor singular value decomposition for high-dimensional data.
method STAT-SVD method with double projection & thresholding scheme.
result STAT-SVD provides sharp thresholding criterion and minimax rate-optimal estimation.
Optimal iterative thresholding algorithms improve upon hard and soft thresholding.
problem Optimizing sparsity or rank constraints in optimization problems.
method Developed the notion of relative concavity for thresholding operators, finding a new class of operators that are optimal.
result A new class of thresholding operators, including ℓq thresholding and reciprocal thresholding, achieves the strongest convergence guarantee. Locally adaptive clustering for tree delineation.
problem Tree delineation from distance data.
method Locally adaptive hierarchical cluster termination.
result Multi-scale alternative to conventional termination criteria.
Spiking neuronal networks are usually simulated with three main simulation schemes: the classical time-driven and event-driven schemes, and the more recent hybrid scheme. All three schemes evolve the state of a neuron through a series of checkpoints: equally spaced in the first scheme and determined neuron-wise by spik…
Consider the optimal dividend problem for an insurance company whose uncontrolled surplus precess evolves as a spectrally negative Levy process. We assume that dividends are paid to the shareholders according to admissible strategies whose dividend rate is bounded by a constant. The objective is to find a dividend poli…
SoftAD improves classification accuracy with less fine-tuning and fewer computational costs.
problem Improving classification accuracy with less fine-tuning and fewer computational costs.
method SoftAD is a softened, pointwise mechanism that downweights borderline points and limits the effects of outliers.
result SoftAD achieves classification accuracy competitive with flooding and SAM, with a smaller loss generalization gap and model norm.
AIHT improves online high-dimensional quantile regression by separating support discovery and refinement.
problem Online high-dimensional quantile regression with structural sparsity.
method Adaptive Iterative Hard Thresholding (AIHT) alternates stochastic updates with adaptive hard-thresholding steps.
result AIHT achieves logarithmic regret for the sliding-window objective in high-dimensional settings.
Solves a new bandit problem with duels and pulls for crowdsourcing.
problem Finding the best arms with mean rewards above a threshold.
method Alternates between ranking and binary search to solve TBP-DC.
result Proves optimality of the Rank-Search algorithm.
Study proposes machine learning to estimate lactate threshold for runners.
problem Inconvenient and expensive blood lactate measurement for recreational runners.
method Recurrent neural networks and standardized temporal axis.
result 89.52% accuracy in estimating lactate threshold.
Optimal threshold resetting reduces search time for multiple diffusive searchers.
problem Optimizing search time for multiple diffusive searchers in a one-dimensional space.
method Threshold resetting (TR) is introduced as an event-driven optimization strategy, coupling resetting to the internal dynamics of searchers.
result Optimal threshold distance u significantly reduces mean first-passage time for N≥2 searchers, with a minimum at Nopt(u). IHT improves sparse distribution learning.
problem Learning sparse discrete distributions.
method Iterative hard thresholding as a solution, with a greedy approximate projection.
result IHT achieves state of the art results for sparse distribution learning.
Multi-task feature learning aims to identity the shared features among tasks to improve generalization. It has been shown that by minimizing non-convex learning models, a better solution than the convex alternatives can be obtained. Therefore, a non-convex model based on the capped-ℓ1,ℓ1 regularization wa…
Paper analyzes robust matrix completion with efficient nonconvex method and leave-one-out analysis.
problem Robust matrix completion with sparse noise.
method Alternates between projected gradient step for low-rank and thresholding step for sparse noise.
result Achieves linear convergence for general thresholding functions.
Detecting edge correlation between two graphs sharpens a threshold based on densest subgraph.
problem Detecting edge correlation between two Erdős-Rényi graphs.
method Formulated as a hypothesis testing problem, connecting to densest subgraph detection.
result Sharp information-theoretic threshold established for edge correlation detection.
Class imbalance presents a major hurdle in the application of data mining methods. A common practice to deal with it is to create ensembles of classifiers that learn from resampled balanced data. For example, bagged decision trees combined with random undersampling (RUS) or the synthetic minority oversampling technique…
A class of heterogeneous agent models is investigated where investors switch trading position whenever their motivation to do so exceeds some critical threshold. These motivations can be psychological in nature or reflect behaviour suggested by the efficient market hypothesis (EMH). By introducing different propensitie…
Procedure optimizes default thresholds to minimize financial loss in credit risk scenarios.
problem Finding the optimal default threshold to minimize financial loss in loan portfolios.
method Objective comparison and evaluation of default definitions using optimisation procedure.
result Loss minima can exist for a select range of credit risk profiles, suggesting loss optimisation of default thresholds is viable.
GPDFlow models extreme threshold exceedance with flexible dependence using normalizing flows.
problem Challenges in modeling multivariate threshold exceedance probabilities due to infinite parametrizations.
method GPDFlow uses normalizing flows to flexibly represent dependence without explicit parametric assumptions.
result GPDFlow significantly improves modeling accuracy and flexibility compared to traditional parametric methods.
In this work we propose a simple and easily parallelizable algorithm for multiway graph partitioning. The algorithm alternates between three basic components: diffusing seed vertices over the graph, thresholding the diffused seeds, and then randomly reseeding the thresholded clusters. We demonstrate experimentally that…
Paper introduces WWAggr for ensemble CPD, improving accuracy and decision threshold selection.
problem Challenges in detecting abrupt distribution shifts in high-dimensional data streams.
method Introduces WWAggr, a novel task-specific ensemble aggregation method based on Wasserstein distance.
result Demonstrates WWAggr outperforms standard aggregation techniques and decision threshold selection.
We relax demographic parity in regression by enforcing parity at quantile levels and score thresholds.
problem Enforcing full distributional fairness in regression can lead to substantial accuracy loss.
method Introduce (ℓ, Z)-fair predictor, derive closed-form solutions, and develop post-processing algorithm. result The risk gap to the continuous optimum vanishes as the grid is refined, and we enable targeted fairness corrections.
New approach finds analytic interpretation of algebraic invariants for balanced metrics.
problem Finding analytic interpretation of algebraic invariants for balanced metrics.
method Using log canonical thresholds and basis divisors, the approach involves quantized Ding functionals on Bergman spaces.
result Each δ_m is the coercivity threshold of a quantized Ding functional on the m-th Bergman space, characterizing the existence of balanced metrics.
New technique reduces gender discrimination in credit lending models.
problem Bias and unfairness in credit lending predictions.
method Subgroup Threshold Optimizer (STO) technique.
result Reduces gender discrimination by over 90%.
The paper examines fairness issues in decision-making systems when protected class labels are unobserved.
problem Fairness assessment challenges when protected class labels are unavailable.
method Decomposes biases in estimating outcome disparity via threshold-based imputation and proposes a weighted estimator.
result Threshold-based imputation generally overestimates disparities, while the weighted estimator has a simpler negative bias.
We present a framework and analysis of consistent binary classification for complex and non-decomposable performance metrics such as the F-measure and the Jaccard measure. The proposed framework is general, as it applies to both batch and online learning, and to both linear and non-linear models. Our work follows recen…
Bayesian method estimates contamination factor for unsupervised anomaly detection.
problem No good methods for estimating contamination factor in unsupervised anomaly detection.
method Bayesian approach using mixture formulation of anomaly detector outputs.
result Estimated contamination factor distribution is well-calibrated and improves anomaly detection performance.
Classifiers based on sparse representations have recently been shown to provide excellent results in many visual recognition and classification tasks. However, the high cost of computing sparse representations at test time is a major obstacle that limits the applicability of these methods in large-scale problems, or in…
Our goal in this paper is to propose an alternative risk measure which takes into account the fluctuations of losses and possible correlations between random variables. This new notion of risk measures, that we call Copula Conditional Tail Expectation describes the expected amount of risk that can be experienced given …
Paper proves convergence for private FL on non-Lipschitz convex objectives using normalization instead of clipping.
problem Lack of convergence results for differentially private federated learning with non-Lipschitz objectives.
method Developed a convergence result for private FL on smooth convex objectives without assuming Lipschitzness, using normalization instead of clipping.
result Normalization-based private FL algorithm converges better than clipping-based counterpart on smooth convex functions.
New convergence analysis for Lasso l1 reweighting improves practical performance.
problem Theoretical convergence of Lasso l1 reweighting methods is limited.
method Biconvex analysis for an alternated convex search.
result Numerical convergence of the algorithm sequence for practical purposes.
The execution flow drives market dynamics, validated on real data.
problem Understanding the fundamental driving force of market dynamics.
method Developed a numerical framework using the Radon-Nikodym derivative to calculate execution flow and determined thresholds and characteristic time scales.
result Execution flow is the fundamental driving force of market dynamics.
We derive expressions for the first three moments of the decision time (DT) distribution produced via first threshold crossings by sample paths of a drift-diffusion equation. The "pure" and "extended" diffusion processes are widely used to model two-alternative forced choice decisions, and, while simple formulae for ac…
In this paper we study nonconvex penalization using Bernstein functions whose first-order derivatives are completely monotone. The Bernstein function can induce a class of nonconvex penalty functions for high-dimensional sparse estimation problems. We derive a thresholding function based on the Bernstein penalty and di…
Study uses active learning to automate EEG event annotation.
problem Lack of annotated clinical EEG data for machine learning models.
method Active learning algorithm for automated annotation of six types of EEG events.
result Recognition performance improved 2% absolute, capable of auto-annotating.
Motivated by the asset-liability management of a nuclear power plant operator, we consider the problem of finding the least expensive portfolio, which outperforms a given set of stochastic benchmarks. For a specified loss function, the expected shortfall with respect to each of the benchmarks weighted by this loss func…
We consider training probabilistic classifiers in the case of a large number of classes. The number of classes is assumed too large to perform exact normalisation over all classes. To account for this we consider a simple approach that directly approximates the likelihood. We show that this simple approach works well o…
Proposes Likelihood Regret for VAEs to improve OOD detection.
problem VAEs can assign high likelihoods to OOD samples, making traditional likelihood thresholds unreliable.
method Introduces Likelihood Regret, a new OOD score for VAEs.
result Empirical results show Likelihood Regret outperforms existing methods for VAEs.
New method converts conventional ANNs to SNNs with minimal loss and efficiency.
problem Difficulty in training SNNs directly from conventional ANNs due to discreteness.
method Proposes a novel pipeline combining threshold balance and soft-reset mechanisms for efficient conversion.
result Achieves almost no accuracy loss with only 1/10 of typical SNN simulation time.
New method tests weighted networks without thresholding, improving accuracy.
problem Testing and anomaly detection on weighted network data.
method Hierarchical Bayesian hypothesis testing framework for weighted networks.
result Method shows lower Type I error and higher statistical power compared to alternatives.
Paper offers a fast method to assess DeFi liquidation risk.
problem Assessing liquidation risk in DeFi stablecoin lending.
method Modeling collateral exchange rate as zero-drift geometric Brownian motion.
result Derives an exact formula for liquidation probability.
The paper shows exchanging estimates over networks is effective for learning sparse signals.
problem Learning sparse signals over networks with limited communication.
method Iterative algorithm exchanging intermediate estimates over a network, with theoretical and simulation analysis.
result The iterative algorithm provides competitive performance in learning sparse signals.
Paper derives new option pricing formulas and approximations for a local volatility model with discontinuity.
problem Modeling extreme ATM skew in a local volatility model with discontinuity.
method Uses joint distribution of Skew Brownian motion and its functionals to derive option pricing formulas and approximations.
result Derives an approximation of option prices by Black-Scholes prices, simplifying skew behavior.
Developed a new thresholding method that connects soft and hard thresholding.
problem Connecting soft and hard thresholding methods in data analysis.
method Scaled soft thresholding method with empirical scaling values.
result Found two sources of over-fitting in the scaled soft thresholding method.
This paper shows how to combine optimal tests into log-optimal processes.
problem How to combine optimal sequential tests into log-optimal processes.
method Using a new class of WAIT e-processes, the paper aggregates asymptotically optimal sequential tests into asymptotically log-optimal processes.
result It is possible to aggregate asymptotically optimal sequential tests into asymptotically log-optimal e-processes.
Self-paced learning and hard example mining re-weight training instances to improve learning accuracy. This paper presents two improved alternatives based on lightweight estimates of sample uncertainty in stochastic gradient descent (SGD): the variance in predicted probability of the correct class across iterations of …
Proposes COLA, a communication-efficient algorithm for decentralized optimization.
problem Decentralized consensus optimization over a network.
method Linearization and communication-censoring strategy to reduce computation and communication costs.
result Proven convergence and established convergence rates for COLA.
We study the problem of robust time series analysis under the standard auto-regressive (AR) time series model in the presence of arbitrary outliers. We devise an efficient hard thresholding based algorithm which can obtain a consistent estimate of the optimal AR model despite a large fraction of the time series points …