Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,786 papers · 148 categories

Trend · papers per month

3136269381,251 · Jun 202019922001200920172026
48 results for data input reformulation

New framework improves fairness in machine learning models.

problem Eliminating discrimination in machine learning models for sensitive fields.
method Reformulate data input by removing sensitive features and design adversarial network to minimize dependence.
result Our model achieves better fairness metrics and prediction performance than state-of-the-art methods.

Compressive learning framework adapted for semi-parametric models.

problem Handling large datasets efficiently with semi-parametric models.
method Reformulate compressive learning framework to handle semi-parametric models, capturing their inherent topology and structure.
result Demonstrated robustness and efficiency of the framework in independent component analysis and subspace clustering.

New method certifies neural network robustness under random input noise.

problem Certifying neural network robustness against random input noise.
method Chance-constrained optimization problem reformulated with input-output samples, convex conditions developed.
result Proposed method certifies robustness against various input noise regimes over larger uncertainty regions.

Paper tackles SMPC for linear systems with unknown noise distribution.

problem Stochastic MPC for linear systems with chance state constraints and unknown noise distribution.
method Reformulate chance constraints, design robust benchmark SMPC, and develop adaptive SMPC with online noise statistics learning.
result Adaptive SMPC guarantees time-uniform satisfaction of unknown reformulated state constraints with high probability.

RISE framework unifies and improves time series learning with missing data.

problem Learning from time series with missing data.
method RISE framework unifies and improves time series learning with missing data.
result RISE instances always benefit from encoders that learn representations for numerical values.

We reformulate data-dependent constraints to ensure they are always met with high probability.

problem Ensuring fairness and stability in machine learning models with data-dependent constraints.
method Calibrated reformulation of constraints to guarantee satisfaction with a specified probability.
result Our method guarantees that fairness constraints are met at test time with high probability.

CcGAN tackles conditional image generation for continuous labels.

problem Mathematical challenges in conditioning on continuous, scalar labels.
method Proposes novel empirical losses and label input methods for continuous conditional GANs.
result CcGAN generates diverse, high-quality images from continuous labels.

EMDE efficiently estimates manifold densities for diverse recommendation systems.

problem Efficiently estimating manifold densities for multi-modal recommendation systems.
method EMDE (Efficient Manifold Density Estimator) framework for arbitrary vector representations.
result Established new state-of-the-art results in top-k and session-based recommendation settings.

Nyquist ghost artifacts in EPI are originated from phase mismatch between the even and odd echoes. However, conventional correction methods using reference scans often produce erroneous results especially in high-field MRI due to the non-linear and time-varying local magnetic field changes. Recently, it was shown that …

2018-06-01abs ↗pdf ↗

Optimizes calibration error estimators for better classifier trustworthiness.

problem Lack of guidance on selecting and tuning calibration error estimators.
method Reformulates calibration estimation as a regression problem with i.i.d. input pairs.
result Demonstrates the effectiveness of optimized calibration estimators on image classification tasks.

ROBOT framework solves regression without correspondence for large data and complex models.

problem Regression without known correspondence in large datasets.
method ROBOT framework reformulates regression as a continuous optimization problem and uses hypergradient approach.
result ROBOT achieves better performance than existing methods in linear and nonlinear regression tasks.

Paper reformulates UOT as non-negative penalized linear regression for efficient algorithms.

problem Optimal transport with relaxed marginal conditions.
method Reformulate UOT as non-negative penalized linear regression, propose multiplicative updates.
result Efficient algorithms for UOT with quadratic penalties, continuity of solutions.

LOL-BO improves latent space Bayesian optimization over structured inputs.

problem Optimizing complex functions over high-dimensional, structured search spaces.
method Adapting trust regions from high-dimensional to structured settings, using a DAE to map inputs into a latent space.
result Achieves up to 20x improvement over state-of-the-art methods.

After reconsidering the Dasbach-Hougardy counterexample to the Kauffman Conjecture on alternating knots, we reformulate the conjecture and consider Dasbach-Hougardy counterexample and similar counterexamples in the light of the reformulated conjecture.

2010-05-20abs ↗pdf ↗

In order to cope with the increased data volumes generated by modern radio interferometers such as LOFAR (Low Frequency Array) or SKA (Square Kilometre Array), fast and efficient calibration algorithms are essential. Traditional radio interferometric calibration is performed using nonlinear optimization techniques such…

2013-03-05abs ↗pdf ↗

Cut-DeepONet handles discontinuities and sharp transitions in neural operators.

problem Neural operators struggle with discontinuities and sharp transitions in PDEs.
method Two-stage training framework that explicitly models discontinuities via a lifting strategy and input-dependent discontinuity prediction.
result Cut-DeepONet outperforms state-of-the-art methods on benchmark PDEs with low-resolution datasets.

A new framework using kernel packets overcomes limitations of state space models for multi-dimensional data.

problem Computational limitations of Gaussian process regression in large-scale applications.
method Kernel packet approach, identifying KPs via forward and backward state space representations.
result Exact, memory-efficient inference with linear-time training and logarithmic/predictive time.

We propose a novel adversarial training method in feature space that improves model robustness and computational efficiency.

problem Improving model robustness against adversarial input perturbations with computational efficiency.
method Shift from input to feature-space perturbations, reformulating the adversarial training problem in reproducing kernel Hilbert spaces, enabling exact solution of inner maximization and efficient optimization.
result The feature-perturbed formulation is a relaxation of the original problem and provides a regularized estimator that adapts to noise and function smoothness.

Paper uses AI to predict medications from medical codes, improving accuracy in healthcare.

problem Predicting medications from incomplete or incorrect medical codes is challenging.
method Robust Recurrent Neural Networks (RNNs) with decay mechanism and noise injection.
result The method accurately predicts medication orders from contaminated medical codes.

Simplifies machine learning validation using kNN and conditional probability algorithms.

problem Validating machine learning models in practical applications.
method Reformulated regression and classification problems using kNN and conditional probability algorithms.
result Online capability and reduced memory usage compared to kNN.

Neural networks perform differently when regression is treated as classification.

problem Understanding why neural networks perform better when regression is treated as classification.
method Analyzing two-layer ReLU networks and their feature spaces, focusing on the cross entropy loss vs. square loss.
result The support of the measure induced by the square loss differs from that of the cross entropy loss, indicating optimization difficulties.

The problem of joint feature selection across a group of related tasks has applications in many areas including biomedical informatics and computer vision. We consider the l2,1-norm regularized regression model for joint feature selection from multiple tasks, which can be derived in the probabilistic framework by assum…

2012-05-09abs ↗pdf ↗

In his Ph.D. thesis, Burak Ozbagci described an algorithm computing signatures of Lefschetz fibrations where the input is a factorization of the monodromy into a product of Dehn twists. In this note, we give a reformulation of Ozbagci's algorithm which becomes much easier to implement. Our main tool is Wall's non-addit…

2019-07-26abs ↗pdf ↗

Reformulated sigma models for complex Grassmannians using Gross-Neveu formalism.

problem Classical aspects of N=(2,2)\mathcal{N}=(2,2) supersymmetric sigma models with Hermitian symmetric target spaces.
method Reformulation using Gross-Neveu formalism, proposing two types of equivalent Lagrangians.
result Proposed two types of equivalent Lagrangians for maximal isotropic Grassmannians, making either supersymmetry or geometry manifest.

In this article we dwell into the class of so called ill posed Linear Inverse Problems (LIP) in machine learning, which has become almost a classic in recent times. The fundamental task in an LIP is to recover the entire signal / data from its relatively few random linear measurements. Such problems arise in variety of…

2019-08-16abs ↗pdf ↗

Adversarial training improves linear regression solutions, revealing sparsity and abrupt interpolation.

problem Adversarial attacks on linear regression models.
method Formulated as a convex problem, adversarial training is used to find robust solutions that are sparse and interpolate data.
result Adversarial training with small disturbances gives the solution with the minimum-norm that interpolates the training data, revealing abrupt transition into interpolation.

We establish a characterization of adequate knots in terms of the degree of their colored Jones polynomial. We show that, assuming the Strong Slope conjecture, our characterization can be reformulated in terms of "Jones slopes" of knots and the essential surfaces that realize the slopes .For alternating knots the refor…

2016-01-13abs ↗pdf ↗

DBPA assesses LLM perturbations using frequentist hypothesis testing.

problem Quantifying input perturbation impacts on LLM outputs.
method DBPA reformulates perturbation analysis as frequentist hypothesis testing, using Monte Carlo sampling for empirical null and alternative distributions.
result DBPA provides interpretable p-values and scalar effect sizes for LLM perturbations.

Feature attribution methods, or saliency maps, are one of the most popular approaches for explaining the decisions of complex machine learning models such as deep neural networks. In this study, we propose a stochastic optimization approach for the perturbation-based feature attribution method. While the original optim…

2018-07-12abs ↗pdf ↗

We investigate structured sparsity methods for variable selection in regression problems where the target depends nonlinearly on the inputs. We focus on general nonlinear functions not limiting a priori the function space to additive models. We propose two new regularizers based on partial derivatives as nonlinear equi…

2018-05-16abs ↗pdf ↗