Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

3876113151 · Jun 202019922001200920182026
48 results for ramp activation

Novel ramp loss method improves weakly supervised machine translation and parsing.

problem Training neural models without gold labels in weak supervision scenarios.
method Adapted ramp loss objectives to promote positive outputs and discourage negative ones.
result Bipolar ramp loss objectives outperform other methods on weakly supervised tasks.

Automates detection of fast-ramped flexibility events for DSOs.

problem Monitoring and supervising flexibility activations in power systems.
method Unsupervised detection and open-set classification.
result Automatically identifies critical flexibility activations for early intervention.

Study develops a machine learning-based ramp metering model to improve freeway efficiency.

problem Improving ramp metering to maintain freeway efficiency under various traffic conditions.
method Machine learning approach using historical data to predict and manage traffic flow.
result The novel model outperforms a baseline traffic-responsive ramp metering algorithm.

Deep neural networks generalize well despite having many more parameters than data.

problem Theoretical basis for deep learning's ability to generalize with many more parameters than data.
method Examined the statistical risk of multi-layer networks with 1\ell^1-type parameter controls and ramp activation functions.
result The risk is upper bounded by [(L3logd)/n]1/2[(L^3 \log d)/n]^{1/2}, demonstrating that input dimensions can be much larger than sample sizes.

In this paper we show all possible ramps where an object can move with constant speed under the effect of gravity and friction. The planar ramp are very easy to describe, just rotate a curve with velocity vector (tanh(as),sech(as)). Recall that tanh(as)^2+sech^2(as) = 1. Therefore, the solution of the planar constant s…

2013-05-02abs ↗pdf ↗

The paper finds that circles and logarithmic spirals are the only constant-speed ramps for a specific force field.

problem Determining planar curves for constant-speed motion under specific force conditions.
method Analyzing the motion of a particle under friction and a central force field.
result Every solution to the constant-speed motion problem approaches either a circle or a logarithmic spiral.

Study improves adversarial classification using distributionally robust models.

problem Improving robustness against adversarial attacks in classification models.
method Distributionally robust chance constraints with Wasserstein ambiguity, reformulated as a regularized ramp loss minimization problem.
result Standard descent methods can converge to the global minimizer for the distributionally robust adversarial classification model.

In this paper we propose a tractable quadratic programming formulation for calculating the equilibrium term structure of electricity prices. We rely on a theoretical model described in [21], but extend it so that it reflects actually traded electricity contracts, transaction costs and liquidity considerations. Our nume…

2014-09-23abs ↗pdf ↗

Deep neural nets on 1-D data are convex Lasso models with reflection features.

problem Training neural networks on 1-D data.
method Proving equivalence to convex Lasso problems with discrete, explicitly defined dictionary matrices.
result Reflection features in neural networks with certain activations.

Seesaw optimizes training by balancing learning rate and batch size, accelerating model pretraining.

problem Optimizing training efficiency for large language models with adaptive optimizers.
method Develops a principled framework for batch-size scheduling, introducing Seesaw which multiplies learning rate by 1/√2 and doubles batch size.
result Empirically, Seesaw reduces wall-clock time by approximately 36% compared to cosine decay, matching theoretical limits.

The paper predicts human-like driving behavior of other vehicles for safer AVs.

problem Safe and efficient interaction of AVs with other vehicles.
method Hierarchical inverse reinforcement learning considering both discrete and continuous decisions.
result The proposed approach accurately predicts both discrete and continuous driving behaviors.

Let f f^{\star} be a function on Rd \mathbb{R}^d with an assumption of a spectral norm vf v_{f^{\star}} . For various noise settings, we show that Ef^f2(vf4logdn)1/3 \mathbb{E}\|\hat{f} - f^{\star} \|^2 \leq \left(v^4_{f^{\star}}\frac{\log d}{n}\right)^{1/3} , where n n is the sample size and f^ \hat{f} is either a penalized lea…

2016-07-05abs ↗pdf ↗

We study dynamic hedging of counterparty risk for a portfolio of credit derivatives. Our empirically driven credit model consists of interacting default intensities which ramp up and then decay after the occurrence of credit events. Using the Galtchouk-Kunita-Watanabe decomposition of the counterparty risk price paymen…

2017-09-04abs ↗pdf ↗

The paper constructs minimizers for deep learning networks and analyzes their geometric structure.

problem Underparametrized deep learning networks and their minimizers.
method Direct construction of minimizers without gradient descent, considering specific settings.
result Explicit family of minimizers for the global minimum and a set of degenerate local minima.

In this paper, we systemally study the long time behavior of the curve shortening flow in a closed or non-compact complete locally Riemannian symmetric manifold. Assume that we have a global flow. Then we can exhibit a a limit for the global behavior of the flow. In particular, we show the following results. 1). Let $\…

2003-12-26abs ↗pdf ↗

Case study shows impact of co-optimizing energy and reserve for wind energy.

problem Impact of lack of co-optimization of energy and reserve in high wind penetration scenarios.
method Developed two models with and without co-optimization, calibrated with Spanish market parameters.
result Models show significant differences in energy and reserve management.

Structured learning is appropriate when predicting structured outputs such as trees, graphs, or sequences. Most prior work requires the training set to consist of complete trees, graphs or sequences. Specifying such detailed ground truth can be tedious or infeasible for large outputs. Our main contribution is a large m…

2012-06-27abs ↗pdf ↗

New method combines FMEA and Bayesian Network for root cause analysis in lithium-ion battery production.

problem Complex cause-effect relationships in lithium-ion battery production.
method Combining FMEA with Bayesian Network to detect and resolve inconsistencies.
result Holistic method builds large-scale cross-process Bayesian Failure Network for root cause analysis.

SPADE improves demand forecasting accuracy by 4.5% for post-promotion periods.

problem Overreacting to peak events in demand forecasting leads to biased forecasts.
method SPADE splits forecasting into two tasks: one for peak events and another for post-peak events, using masked convolution filters and a specialized Peak Attention module.
result Overall PPE improvement of 4.5%, 30% improvement for most affected forecasts after promotions and holidays, and 3.9% improvement in PE accuracy.

The paper examines A/B tests in recommendation systems to detect biased algorithm comparisons due to shared data.

problem Bias in comparing recommendation algorithms due to shared data.
method Formalized as a multi-armed bandit problem, analyzed the sign of difference-in-means estimator vs true GTE.
result Data sharing can lead to biased comparisons of recommendation algorithms, and a detection procedure is proposed.

As automotive electronics continue to advance, cars are becoming more and more reliant on sensors to perform everyday driving operations. These sensors are omnipresent and help the car navigate, reduce accidents, and provide comfortable rides. However, they can also be used to learn about the drivers themselves. In thi…

2017-06-09abs ↗pdf ↗

Proposes a framework for predicting traffic situations involving multiple interacting agents.

problem Predicting the behavior of multiple interacting agents in traffic scenarios.
method Two-layer Hidden Markov Model (TLHMM) for situation distribution and learning-based dynamic scene evolution model for trajectory sampling.
result Demonstrates effectiveness and accuracy in highway ramp merging scenario.

This paper characterizes and designs loss functions for robust classification with abstention.

problem Ensuring robustness against adversarial attacks and knowing when to abstain from prediction.
method Proposes adversarial robust reject option loss and characterizes surrogates for calibration.
result Shifted Double Ramp Loss and Shifted Double Sigmoid Loss satisfy the calibration conditions.

A hybrid model combines Q-learning and PID controller for continuous vehicle control.

problem Learning unsatisfactory results with discrete action space in autonomous driving.
method Combining Q-learning and PID controller, using Quadratic Q-function approximation and action network.
result Autonomous vehicle successfully learns smooth and efficient driving behavior.

The support vector machine (SVM) is one of the most successful learning methods for solving classification problems. Despite its popularity, SVM has a serious drawback, that is sensitivity to outliers in training samples. The penalty on misclassification is defined by a convex loss called the hinge loss, and the unboun…

2014-09-03abs ↗pdf ↗

SAT improves adversarial training by smoothing the loss landscape through curriculum learning.

problem Adversarial training sacrifices clean accuracy for robustness and suffers from large generalization error.
method SAT uses curriculum learning to smooth the adversarial loss landscape, improving both clean and robust accuracy.
result SAT models improve clean and robust accuracy significantly compared to adversarial training and other baselines.

MLM models match or exceed RN in generating wind power time series without location info.

problem Generating accurate long-term wind power time series without location information.
method Applied neural networks to MERRA2 wind speed data with and without location info.
result MLM models produce time series of equal or better quality than RN.

iGNN tackles inverse graph prediction using invertible neural networks.

problem Inverse graph prediction problem in data analysis and machine learning.
method Developed invertible graph neural network (iGNN) to solve inverse prediction problem on graphs.
result iGNN model allows efficient generation from output labels and forward prediction.

Many activation functions have been proposed in the past, but selecting an adequate one requires trial and error. We propose a new methodology of designing activation functions within a neural network at each layer. We call this technique an "activation ensemble" because it allows the use of multiple activation functio…

2017-02-24abs ↗pdf ↗

This paper studies activation sparsity in large language models, finding key trends and implications.

problem Activation sparsity in large language models (LLMs) can be improved for efficiency and interpretability.
method Proposes PPL-p%p\% sparsity, analyzes trends with training data, width-depth ratio, and parameter scale.
result ReLU is more efficient for sparsity than SiLU, and deeper architectures can improve sparsity.

BinaryDuo improves BNNs by coupling binary activations, outperforming state-of-the-art models.

problem Gradient mismatch in BNNs due to binarizing activations.
method Using gradient of smoothed loss function to estimate gradient mismatch, proposing BinaryDuo scheme with coupled ternary activations.
result BinaryDuo outperforms state-of-the-art BNNs on various benchmarks.

Evolutionary algorithms improve neural network performance by discovering better activation functions.

problem The choice of activation function affects neural network performance, but ReLU remains dominant.
method Defined a tree-based search space of candidate activation functions and used evolutionary algorithms (mutation, crossover, exhaustive search) to explore and discover better functions.
result Replacing ReLU with evolved activation functions statistically significantly increases network accuracy.