Evaluates change point detection algorithms on real-world data.
problem Insufficient evaluation of change point detection algorithms on real-world time series.
method Developed a data set of 37 time series from various domains, annotated by human experts, and evaluated 14 algorithms using consistency metrics.
result Demonstrates the need for better evaluation methods in change point detection.
Cer-Eval saves LLM evaluation costs while maintaining accuracy.
problem Challenges in evaluating large language models due to large dataset requirements.
method Adapts to different evaluation objectives, uses test sample complexity, and develops a partition-based algorithm.
result Cer-Eval can save 20-40% test points with comparable accuracy and 95% confidence guarantee.
Multi-dimensional state-integrals of products of Faddeev's quantum dilogarithms arise frequently in Quantum Topology, quantum Teichmüller theory and complex Chern--Simons theory. Using the quasi-periodicity property of the quantum dilogarithm, we evaluate 1-dimensional state-integrals at rational points and express the…
Proposes BA method for unbiased time series anomaly detection evaluation.
problem Anomalies in time series data are rare, making F1-score unreliable.
method Introduces Balanced Point Adjustment (BA) to address F1-score bias.
result BA provides fairer evaluation of time series anomaly detectors.
New framework evaluates MIA without retraining, addressing biases.
problem Evaluate MIA without retraining and under distribution shift.
method Causal inference approach to MIA evaluation.
result Practical estimators for MIA metrics without retraining.
The paper evaluates homology for links in a solid torus with special boundary conditions.
problem Evaluating homology for links in a solid torus with specific boundary conditions.
method Using foam evaluation, the paper describes equivariant SL(2) and SL(3) homology for links in the solid torus with a distinguished line.
result Generators of state spaces for annular webs are represented by foams with boundary intersecting a distinguished line, contributing additional terms to the foam evaluation.
In many applications of black-box optimization, one can evaluate multiple points simultaneously, e.g. when evaluating the performances of several different neural network architectures in a parallel computing environment. In this paper, we develop a novel batch Bayesian optimization algorithm --- the parallel knowledge…
PS-BAX uses posterior sampling to select evaluation points for efficient Bayesian algorithm execution.
problem Efficiently selecting evaluation points for expensive functions with limited evaluations.
method Posterior sampling to guide sequential selection of evaluation points.
result PS-BAX is faster, simpler, and more scalable than existing methods.
New algorithm speeds up solving saddle-point problems with large condition numbers.
problem Solving saddle-point problems with large condition numbers.
method Proposes a stochastic proximal point algorithm that accelerates variance reduction methods.
result Reduces logarithmic term of condition number for iteration complexity.
This work compares OmniAnomaly with PCA for MTSAD, finding PCA can match or outperform OmniAnomaly.
problem Comparing deep learning models with classical methods in MTSAD under fair evaluation protocols.
method Systematic comparison of OmniAnomaly and PCA on SMD, using identical thresholding and evaluation procedures.
result PCA can achieve performance comparable to OmniAnomaly and even outperform it under certain conditions.
Look-Ahead-Bench evaluates financial LLMs for lookahead bias, revealing significant differences in model performance.
problem Measuring and mitigating lookahead bias in financial LLMs.
method Standardized benchmark evaluating model behavior in practical financial scenarios, analyzing performance decay across market regimes.
result Standard LLMs exhibit significant lookahead bias, while Pitinf models show improved generalization and reasoning abilities.
Deep RL evaluation underestimates uncertainty, leading to misleading conclusions.
problem Statistical uncertainty in deep RL performance evaluations is underestimated, leading to misleading conclusions.
method Advocates for reporting interval estimates of aggregate performance and proposes performance profiles to account for variability.
result Substantial discrepancies in prior performance comparisons are revealed, highlighting the need for more rigorous evaluation methods.
A new framework evaluates model performance on single input points, revealing insights into data and model structure.
problem Traditional evaluation methods in machine learning are insufficient for understanding model performance and data structure.
method Developed a pointwise framework to measure model performance on individual input points, analyzing the relationship between pointwise and average performance.
result Profiles of data points reveal different types of correlations between pointwise and average performance, challenging existing models of learning.
Policy evaluation is a crucial step in many reinforcement-learning procedures, which estimates a value function that predicts states' long-term value under a given policy. In this paper, we focus on policy evaluation with linear function approximation over a fixed dataset. We first transform the empirical policy evalua…
DPP-BBO diversifies batched Bayesian optimization using DPPs.
problem Efficiently proposing diverse and informative batches in batched Bayesian optimization.
method Introducing DPP-Batch Bayesian Optimization (DPP-BBO) with DPP-Thompson Sampling (DPP-TS).
result Novel Bayesian simple regret bounds for DPP-TS show improved performance over classical methods.
The Minimal Learning Machine improves regression performance with reference point selection.
problem Improving the performance of the Minimal Learning Machine (MLM) in regression tasks.
method Developed theoretical guarantees for MLM's interpolation and approximation capabilities. Proposed clustering-based methods for selecting reference points to enhance MLM's generalization.
result Clustering-based methods for reference point selection outperform random selection, especially with a small number of points.
Recurrent neural networks improve time series forecasting accuracy.
problem Time series forecasting is challenging, especially for sequential data.
method A recurrent neural network framework for feature engineering, prediction, and evaluation is presented.
result The LSTM and GRU networks outperform traditional methods in forecasting accuracy.
Paper proposes a method to create more reliable confidence intervals for off-policy evaluations.
problem Creating reliable confidence intervals for off-policy evaluations.
method Proposes a deeply-debiasing procedure to construct efficient, robust, and flexible confidence intervals.
result Validated by theoretical results and numerical experiments, the method improves the reliability of off-policy evaluations.
Active learning (AL) repeatedly trains the classifier with the minimum labeling budget to improve the current classification model. The training process is usually supervised by an uncertainty evaluation strategy. However, the uncertainty evaluation always suffers from performance degeneration when the initial labeled …
Research evaluates adversarial attacks and defenses on 3D point cloud classifiers.
problem Robustness of 3D object classifiers against adversarial attacks.
method Extending 2D adversarial attacks to 3D point clouds and proposing new defenses.
result 3D point cloud classifiers are weak to adversarial attacks but more defensible.
This work evaluates uncertainty in deep Gaussian processes.
problem Uncertainty quantification in deep Gaussian processes.
method Hierarchical deep Gaussian processes (DGPs) and Deep Sigma Point Processes (DSPPs) evaluated on regression and classification tasks.
result DSPPs provide strong in-distribution calibration but are less robust under distribution shift compared to ensembles.
Study variance-reduced method for estimating fixed points in Banach spaces.
problem Estimating fixed points of contractive operators in Banach spaces with noisy evaluations.
method Variance-reduced stochastic approximation scheme in Banach spaces.
result Establish non-asymptotic bounds for operator defect and estimation error.
REMAL: Residual Equilibrium Manifold Active Learning for Surrogate-Based Multidisciplinary Design Analysis
problem Multidisciplinary design analysis of coupled engineering systems requires solving equilibrium states where all disciplinary coupling variables are consistent.
method Residual manifold surrogate modeling framework for coupled systems.
result REMAL learns a surrogate model of the joint residual manifold via multitask Gaussian process models.
Estimates change point in dynamic stochastic block model.
problem Estimating the location of a single change point in a dynamic stochastic block model.
method Two methods: least squares with clustering and ignoring community structures.
result Established rates of convergence and asymptotic distributions of change point estimators.
Optimistic search speeds up change point detection in large datasets.
problem Efficiently detecting change points in large-scale data with high computational demands.
method Adaptive logarithmic queries to reduce evaluation complexity.
result Asymptotic minimax optimality and fast localization rates for change point detection.
The Neural Testbed evaluates joint predictions of neural agents, revealing their limitations.
problem Evaluating the quality of joint predictions generated by neural agents.
method Developed an open-source benchmark (The Neural Testbed) to assess agents' marginal and joint predictions.
result Popular Bayesian deep learning agents perform poorly on joint predictions, even with accurate marginal predictions.
Active learning reduces SP calculations by 90%.
problem Efficiently calculating saddle points in energy functions.
method Active learning framework with GPR and GAD.
result Significant reduction in the number of expensive evaluations.
Greedy algorithms which use only function evaluations are applied to convex optimization in a general Banach space X. Along with algorithms that use exact evaluations, algorithms with approximate evaluations are treated. A priori upper bounds for the convergence rate of the proposed algorithms are given. These bounds…
New benchmark for earthquake forecasting models shows current neural point processes are not yet suitable.
problem Lack of a modern benchmark for evaluating neural point process models in earthquake forecasting.
method Curated and standardized earthquake catalog, evaluation protocols, and datasets.
result None of the tested NPPs outperformed the classical ETAS model.
New geometric metrics improve Bayesian optimization performance evaluation.
problem Current metrics lack geometric insights and cannot compare algorithms effectively.
method Proposed four geometric metrics: precision, recall, average degree, and average distance.
result Proposed metrics provide more detailed evaluation of Bayesian optimization.
We consider parallel global optimization of derivative-free expensive-to-evaluate functions, and propose an efficient method based on stochastic approximation for implementing a conceptual Bayesian optimization algorithm proposed by Ginsbourger et al. (2007). At the heart of this algorithm is maximizing the information…
We develop parallel predictive entropy search (PPES), a novel algorithm for Bayesian optimization of expensive black-box objective functions. At each iteration, PPES aims to select a batch of points which will maximize the information gain about the global maximizer of the objective. Well known strategies exist for sug…
Improved Transformer language models using dynamic evaluation.
problem Language model perplexity and accuracy improvements.
method Combining Transformers with dynamic evaluation techniques.
result Significant improvement in language model performance (e.g., 0.99 to 0.94 bits/char).
In recent decades, the use of 3D point clouds has been widespread in computer industry. The development of techniques in analyzing point clouds is increasingly important. In particular, mapping of point clouds has been a challenging problem. In this paper, we develop a discrete analogue of the Teichmüller extremal mapp…
Solves challenges in replicating ML/DL model evaluations.
problem Challenges in evaluating and studying ML/DL innovations.
method Proposes MLModelScope for repeatable model evaluation.
result Facilitates rapid adoption of ML/DL innovations.
A new model uses neural networks to efficiently learn multivariate temporal point processes.
problem Efficiently modeling multivariate temporal point processes with low parameter complexity.
method Modeling the cumulative hazard function with neural networks for each variate.
result The proposed model achieves state-of-the-art performance on data fitting and event prediction tasks.
Paper evaluates and proposes a new method for robust curvature estimation from noisy point cloud data.
problem Challenges in discovering a robust manifold structure in point cloud data.
method Proposes a new robust curvature estimation method based on Weingarten map.
result Demonstrates superior performance of the new method in noisy data conditions.
A new method calculates fractional moments using the moment-generating function.
problem Computing fractional moments from probability densities.
method Integral framework based on moment-generating function.
result Exact integral expressions for various types of moments.
New method tackles instance segmentation on 3D point clouds with improved metrics.
problem Evaluation metrics are affected by small regions containing few instances.
method Proposes a new method with O(Np) space complexity that learns embeddings for clusters of instances.
result Achieves state-of-the-art performance using both existing and proposed metrics.
Core-Halo solves large-scale fixed-point problems by decentralizing updates.
problem Large-scale fixed-point equations with block dependencies.
method Core-Halo decomposition separates write ownership from read-only context, aligning with block-dependence structure.
result Core-Halo achieves near-centralized performance while retaining parallelism.
The use of low-precision fixed-point arithmetic along with stochastic rounding has been proposed as a promising alternative to the commonly used 32-bit floating point arithmetic to enhance training neural networks training in terms of performance and energy efficiency. In the first part of this paper, the behaviour of …
ASEs use surrogate estimation to efficiently evaluate model performance with minimal labels.
problem Efficient model evaluation with limited labels.
method Surrogate-based estimation and active learning.
result ASEs offer greater label-efficiency than current methods for deep neural networks.
The paper proposes criteria and methods for evaluating and aggregating feature-based model explanations.
problem Lack of quantitative evaluation criteria for feature-based model explanations.
method Developed quantitative evaluation criteria (low sensitivity, high faithfulness, low complexity), devised a framework for aggregation, and derived a new aggregate Shapley value explanation function.
result A new aggregate Shapley value explanation function that minimizes sensitivity.
MLModelScope streamlines ML/DL model evaluation and benchmarking.
problem Challenges in evaluating and benchmarking ML/DL models.
method Open-source, framework/hardware agnostic platform with distributed design.
result Demonstrates the impact of model evaluation pipelines and HW/SW choices.
Develops deep jump learning for continuous treatment OPE.
problem Estimating mean outcomes under new treatment rules using historical data from different rules.
method Adaptive deep discretization of continuous treatment space using deep learning and multi-scale change point detection.
result Validated method through theoretical results, simulations, and real application to Warfarin Dosing.
FORE evaluates occupancy ratios without requiring Bellman completeness.
problem Offline reinforcement learning occupancy ratio estimation.
method Fitted occupancy-ratio evaluation (FORE) using adjoint Bellman recursion.
result FORE achieves convergence in KL without Bellman completeness.
Given a heterogeneous time-series sample, the objective is to find points in time (called change points) where the probability distribution generating the data has changed. The data are assumed to have been generated by arbitrary unknown stationary ergodic distributions. No modelling, independence or mixing assumptions…
The paper proves Γ-convergence of discrete tangent-point energies to continuous energies and ropelength, with applications to biarc curves.
problem Proving convergence of discrete tangent-point energies to continuous energies and ropelength.
method Using biarc curves and interpolation, the paper proves Γ-convergence of discretized tangent-point energies to the continuous tangent-point energies and ropelength functional. result Discrete almost minimizing biarc curves converge to ropelength minimizers and minimizers of continuous tangent-point energies.