Deleting refusal directions from models leads to systematically more optimistic decisions.
problem The impact of removing refusal directions from models on decision-making outcomes.
method Ablation study using a frozen pipeline of 21,600 weekly equity decisions.
result Ablation of refusal directions makes models more optimistic and justifies themselves more, but also reduces confidence.
Localized Multidirectional Correction improves non-refusal target-response behavior in foundation models.
problem Controlled post-training refusal suppression in routed MoE and hybrid-MoE foundation models.
method Introduce Localized Multidirectional Correction (LoMC), a support-gated intervention framework.
result Substantially improves non-refusal target-response behavior while maintaining general capability under a compact intervention footprint.
ELS framework improves safety alignment by dynamically steering LLMs towards helpful responses.
problem Over-Refusal in Aligned Large Language Models
method Fine-tuning free framework using an Energy-Based Model (EBM) to dynamically steer LLMs during inference.
result Extensive experiments show a significant reduction in false refusals (from 57.3% to 82.6%) while maintaining safety performance.
New benchmark uncovers hidden biases in LLMs that refuse to answer certain queries.
problem Evaluating fairness in LLMs, especially in sensitive applications.
method Silenced Bias Benchmark (SBB) using activation steering to reduce model refusals during QA.
result Exposes hidden unfair preferences in LLMs' latent space, distinguishing direct responses from underlying fairness issues.
Paper explores MM strategies that can refuse to quote or provide single-sided quotes.
problem Overcoming risks in market making due to changing market conditions.
method Adversarial reinforcement learning with new MM agent designs.
result Refusal to quote or providing single-sided quotes can improve MM performance.
LLM safety alignment explained as divergence estimation.
problem Aligning large language models to avoid harmful outputs.
method Presented a theoretical framework showing alignment methods as divergence estimators.
result KLDO method improves safety alignment using compliance-refusal datasets.
Training models to prefer certain responses can unintentionally shift probability to harmful ones.
problem Likelihood displacement in DPO models, leading to unintended unalignment.
method Characterized and mitigated likelihood displacement using CHES score.
result Training models to prefer certain responses can unintentionally shift probability mass to harmful responses.
A new constructivist approach to modeling in economics and theory of consciousness is proposed. The state of elementary object is defined as a set of its measurable consumer properties. A proprietor's refusal or consent for the offered transaction is considered as a result of elementary economic measurement. We were al…
A new protocol corrects confounding effects to measure alignment-induced activation shifts accurately.
problem Confounding effects in measuring alignment-induced activation shifts using naive methods.
method Introduces a four-variant decomposition to separate alignment shift from template effects.
result Correctly measures alignment-induced activation shifts, recovering behaviorally active subspace.
Machine learning algorithms can unintentionally discriminate; tools detect and fix this.
problem Unintentional discrimination in machine learning algorithms.
method Statistical tools to detect and eliminate biases.
result Tools can identify and mitigate algorithmic discrimination.
Weak form of the Efficiency Market Hypothesis (EMH) excludes predictions of future market movements from historical data and makes the technical analysis (TA) out of law. However the technical analysis is widely used by traders and speculators who steadely refuse to consider the market as a "fair game" and survive with…
LLM sandbox and persona dynamics create unethical reality gaps that shift risk to users.
problem Ethical issues arise from LLMs generating reality gaps that shift risk to uninformed users.
method Analyzes the ethical implications of LLM sandbox and persona dynamics, comparing them to financial regulation and compliance.
result Active generation of reality gaps is unethical as it shifts epistemic risk to users.
In this paper we present detailed simulation results on the wealth distribution model with quenched saving propensities. Unlike other wealth distribution models where the saving propensities are either zero or constant, this model is not found to be ergodic and self-averaging. The wealth distribution statistics with a …
Signed-permutation coordinate transport improves model alignment across checkpoints.
problem Improper alignment of coordinate-indexed objects across model checkpoints.
method Introduces sign-marginalized Hungarian matching and coordinate-preserving transport.
result Recovering signed-permutation gauge improves coordinate alignment and model performance.
Model shows bailout stigma affects firm funding and market performance.
problem Bailout stigma impacts firm funding and market stability.
method Developed a model to analyze the effects of bailout stigma on firms and markets.
result Firms avoid stigma by withdrawing or refusing bailouts, leading to market freeze or revival.
New method estimates extreme outcomes in heavy-tailed data, breaking circular dependence.
problem Estimating outcomes for extreme events in heavy-tailed data.
method Proposes an ADRF estimator that includes a structured tail-shape output and a diagnostic to evaluate tail shape.
result Successfully reduces MAE in deep-tail and conditional-shortfall predictions.
A new constructivist approach to modeling in economics and theory of consciousness is proposed. The state of elementary object is defined as a set of its measurable consumer properties. A proprietor's refusal or consent for the offered transaction is considered as a result of elementary economic measurement. Elementary…
Donoho's JCGS (in press) paper is a spirited call to action for statisticians, who he points out are losing ground in the field of data science by refusing to accept that data science is its own domain. (Or, at least, a domain that is becoming distinctly defined.) He calls on writings by John Tukey, Bill Cleveland, and…
Valid certifies LLMs' domain adherence, bounding out-of-domain behavior.
problem Adversarial susceptibility of LLMs to generate out-of-domain outputs.
method VALID approach providing adversarial bounds as a certificate.
result Validates LLMs' domain adherence with meaningful certificates.
Good data stewardship requires removal of data at the request of the data's owner. This raises the question if and how a trained machine-learning model, which implicitly stores information about its training data, should be affected by such a removal request. Is it possible to "remove" data from a machine-learning mode…
Selective removal of data subsets can efficiently unlearn unwanted distributions.
problem Efficiently removing unwanted data subsets without losing important information.
method Formalized as distributional unlearning, using Kullback-Leibler divergence constraints to select a small subset of data.
result Proposed method achieves corresponding log-loss bounds and is quadratically more sample-efficient than random removal.
Unified framework for explaining models by removing features.
problem Unclear relationships and preferences among existing model explanation methods.
method Removal-based explanations characterized by three dimensions.
result Unified framework unifies 26 existing methods.
The paper proves removable singularity for nonlocal minimal graphs.
problem Proving removable singularities for nonlocal minimal graphs.
method Analyzing (s,1)-capacity zero compact sets to ensure graphs are minimal in the entire domain. result Nonlocal minimal graphs are removable in the entire domain if they are minimal in a set of (s,1)-capacity zero. Unified framework for model explanation methods based on feature removal.
problem Unclear relationships and preferences among various model explanation methods.
method Characterizes removal-based explanations along three dimensions.
result Unified 26 existing methods, including widely used approaches.
Removing spurious features can hurt model accuracy and disproportionately affect different groups.
problem Interference from spurious features in robust model performance across different groups.
method Characterization and analysis of spurious feature removal in noiseless overparameterized linear regression.
result Removal of spurious features can decrease accuracy and disproportionately affect different groups, even in balanced datasets.
Removes singularities for Yang-Mills-Higgs fields in higher dimensions.
problem Yang-Mills-Higgs fields with isolated singularities.
method Establishes decay estimates and conformally invariant energy bounds.
result Removable singularity theorem for Yang-Mills-Higgs fields.
New algorithm for mean estimation in add-remove model achieves optimal error.
problem Mean estimation in add-remove model of differential privacy.
method Proposed new algorithm achieving min-max optimality.
result Achieves best possible constant in mean squared error for all ε.
Machine learning confound removal biases results, leading to misleading predictions.
problem Common confound removal methods in machine learning lead to misleading predictions.
method Featurewise removal of confound variance by linear regression before applying ML.
result This common deconfounding approach can leak information, amplifying null or moderate effects.
We call a singularity of a presymplectic form ω removable in its graph if its graph extends to a smooth Dirac structure over the singularity. An example for this is the symplectic form of a magnetic monopole. A criterion for the removability of singularities is given in terms of regularizing functions for pure spinor…
D3M debiases models by selectively removing problematic examples.
problem Model failures on underrepresented subgroups.
method Isolates and removes specific training examples that cause failures.
result Efficiently trains debiased classifiers with minimal example removal.
PUMA augments models to remove unique data points without performance loss.
problem Preserving model performance while removing unique training data points.
method Explicitly models data influence, reweights remaining data optimally.
result PUMA effectively removes unique data points without performance degradation.
In this paper we prove a local removable singularity theorem for certain minimal laminations with isolated singularities in a Riemannian three-manifold. This removable singularity theorem is the key result used in our proof that a complete, embedded minimal surface in R3 with quadratic decay of curvature ha…
This paper proposes a novel Gaussian process approach to fault removal in time-series data. Fault removal does not delete the faulty signal data but, instead, massages the fault from the data. We assume that only one fault occurs at any one time and model the signal by two separate non-parametric Gaussian process model…
Algorithm removes spurious concepts from neural network representations without harming task performance.
problem Spurious correlations hinder neural network out-of-distribution generalization.
method Iterative algorithm that identifies two orthogonal subspaces in neural network representation.
result Algorithm outperforms existing methods on computer vision and natural language processing benchmarks.
Squint bound improved by removing lnlnT term.
problem Improving the Squint bound by removing the lnlnT term. method Using the Krichevsky--Trofimov algorithm to change the prior.
result Removed the lnlnT term from the Squint bound. The study extends removability results for quasiregular curves in Euclidean spaces.
problem Removability of singularities in quasiregular curves.
method Extending a fundamental inequality for volume forms to calibrations and proving a Caccioppoli inequality for quasiregular curves.
result Every non-constant quasiregular curve has infinite energy.
While great progress has been made recently in automatic image manipulation, it has been limited to object centric images like faces or structured scene datasets. In this work, we take a step towards general scene-level image editing by developing an automatic interaction-free object removal model. Our model learns to …
We consider sufficient conditions of local removability of coincidences of maps f,g:N->M, where M,N are manifolds with dimensions dimN>dimM. The coincidence index is the only obstruction to the removability for maps with fibers either acyclic or homeomorphic to spheres of certain dimensions. We also address the normali…
We prove a removal of singularities result for Bach-flat metrics in dimension 4 under the assumption of bounded L^2 norm of curvature, bounded Sobolev constant and a volume growth bound. This result extends the removal of singularities result for special classes of Bach-flat metrics obtained in \cite{TVMOD}. For the pr…
BERT outperforms traditional machine learning in text classification tasks.
problem Comparing BERT to traditional machine learning methods for text classification.
method Empirical testing of BERT against TF-IDF-based machine learning models in various scenarios.
result BERT demonstrates superior performance and independence from text features.
This study examines how removing edges from complete graphs affects Ollivier Ricci curvature.
problem Conditions under which Ollivier Ricci curvature changes sign after edge removal.
method Defined and analyzed graphs obtained by removing matching, vertex incident, and cycle edges from complete graphs.
result Ollivier Ricci curvature remains positive or zero for graphs formed by removing edges from complete graphs.
Minimal surfaces and curves can have singularities removed by isotopy.
problem Removing singularities of minimal surfaces and curves.
method Isotopy through conformal minimal surfaces and null holomorphic curves.
result Branch points and complete ends of finite total curvature can be removed.
Paper explores unsupervised learning for ultrasound image artifact removal.
problem Improving visual quality of ultrasound images from various artifacts.
method Inspired by optimal transport cycleGAN, unsupervised deep learning for artifact removal.
result Unsupervised learning method provides comparable results to supervised learning.
Note removes degeneracy in Kähler geometry estimates.
problem Estimating diameter and inequalities in Kähler geometry with degeneracy.
method Technical improvement of earlier results.
result Established diameter, Green's functions, and Sobolev inequalities without small degeneracy assumption.
Unified analysis of removal-based feature attributions robustness.
problem Robustness of removal-based feature attributions is not well understood.
method Theoretical analysis and upper bounds derivation for removal-based feature attributions under input and model perturbations.
result Upper bounds for the difference between intact and perturbed attributions derived under various perturbation settings.
This paper surveys some recent results on existence, uniqueness and removable singularities for fully nonlinear differential equations on manifolds. The discussion also treats restriction theorems and the strong Bellman principle.
This paper examines the applicability of Random Matrix Theory to portfolio management in finance. Starting from a group of normally distributed stochastic processes with given correlations we devise an algorithm for removing noise from the estimator of correlations constructed from measured time series. We then apply t…
FinanceBench benchmarks LLMs on financial QA, revealing limitations.
problem Evaluating LLMs' performance on financial question answering.
method Developed a comprehensive test suite (FinanceBench) with 10,231 questions, tested 16 models, and manually reviewed answers.
result Existing LLMs have significant limitations for financial QA, especially GPT-4-Turbo.