Transformers can learn optimal variable selection in group-sparse classification.
problem Understanding how transformers leverage attention to select relevant variables in group-sparse classification.
method Training a one-layer transformer using gradient descent to select variables from one group of input variables.
result A one-layer transformer can correctly leverage the attention mechanism to select variables, disregarding irrelevant ones.
Develops a framework for consistent loss functions with variable transformations.
problem Lack of theoretical understanding of variable transformations in consistent loss functions.
method Formal characterizations of consistency for transformed loss functions in two cases: realization and prediction variables.
result Establishes new identifiable and elicitable functionals for complex predictive tasks.
In this paper, we propose novel strategies for neutral vector variable decorrelation. Two fundamental invertible transformations, namely serial nonlinear transformation and parallel nonlinear transformation, are proposed to carry out the decorrelation. For a neutral vector variable, which is not multivariate Gaussian d…
Computation of moments of transformed random variables is a problem appearing in many engineering applications. The current methods for moment transformation are mostly based on the classical quadrature rules which cannot account for the approximation errors. Our aim is to design a method for moment transformation for …
In this paper, we define a new family of curves and call it a {\it family of similar curves with variable transformation} or briefly {\it SA-curves}. Also we introduce some characterizations of this family and we give some theorems. This definition introduces a new classification of a space curve. Also, we use this def…
Improves data normality with robust transformations.
problem Skewed data distribution.
method Modified Box-Cox and Yeo-Johnson transformations with robust parameter estimation.
result Transformed data approximates normality in the center with outliers.
The paper constructs a complex for the Dirac operator in 4 dimensions.
problem Constructing a complex for the Dirac operator in 4 dimensions.
method Using the Penrose transform, the paper constructs a relative BGG complex and its direct image.
result An explicit construction of a complex starting with the Dirac operator in any number of variables.
Training of discrete latent variable models remains challenging because passing gradient information through discrete units is difficult. We propose a new class of smoothing transformations based on a mixture of two overlapping distributions, and show that the proposed transformation can be used for training binary lat…
In this article, we propose a new algorithm for supervised learning methods, by which one can both capture the non-linearity in data and also find the best subset model. To produce an enhanced subset of the original variables, an ideal selection method should have the potential of adding a supplementary level of regres…
The paper develops a framework for abstracting causal models using category theory.
problem Difficulties in changing the variables used to describe a system, especially from fine-grained to coarse-grained.
method Introduces a category of interventional causal models and uses enriched category theory to prove compositionality properties.
result Compositionality of model transformations is established, with bounded errors for each step.
Study shows how transformers learn to combine simple tasks into complex ones.
problem Understanding how transformers learn to perform complex tasks not seen during training.
method Controlled setting involving variable assignment and modular addition; partitioned training data analysis.
result Small transformers can generalize to unseen combinations of variables and numbers.
New model learns symmetry transformations from complex data.
problem Learning symmetry transformations in complex domains like chemical space.
method Two latent subspaces, deep information bottleneck, continuous mutual information regularizer.
result Model outperforms state-of-the-art methods on artificial and molecular datasets.
Transformers improve solving mixed-integer programs, especially CLSP.
problem Solving Capacitated Lot Sizing Problem (CLSP) with mixed-integer programming.
method Employing transformer models to predict binary variables in CLSP.
result Transformer model outperforms CPLEX and LSTM in solving CLSP.
Simplified calculus for semimartingales makes complex transformations easier.
problem Complex transformations of semimartingales.
method Unified treatment of transformations for real and complex semimartingales.
result Unified calculus for semimartingales simplifies various transformations.
Backlund transformations are used to search for solutions, particularly soliton solutions, of non-linear differential equations. In this paper we present an invariant geometrical theory of Backlund transformations for second order evolution equations with one space variable. The main concept is that of connection defin…
Transformers can handle endogeneity in linear regression using IV methods.
problem Endogeneity in in-context linear regression models.
method Transformer architecture with gradient-based bi-level optimization and in-context pretraining.
result Transformers provide more robust predictions and estimates than 2SLS in endogenous scenarios.
Paper recovers latent causal structure and linear transformation from indirect observations.
problem Recovering latent causal structure and linear transformation from indirect observations.
method Established sufficient conditions for DAG recovery, leveraged score function properties, and used soft/hard interventions.
result Perfect recovery of latent DAG structure and linear transformation up to scaling using soft interventions, hard interventions with additional hypothesis testing.
Spacetimeformer learns spatiotemporal relationships from data alone.
problem Forecasting multivariate time series with distinct spatial relationships.
method Transformers with dynamic graph connections learning interactions between space, time, and value.
result Competitive results on various time series prediction benchmarks.
Random variables of the generalized Pareto distribution, can be transformed to that of the Pareto distribution. Explicit expressions exist for the maximum likelihood estimators of the parameters of the Pareto distribution. The performance of the estimation of the shape parameter of generalized Pareto distributed using …
This paper tackles causal representation learning with linear and general transformations.
problem Identify and recover latent causal variables and graphs under unknown transformations.
method Score-based algorithms that use gradients of log-density functions for identifiability and achievability.
result Two stochastic hard interventions per node are sufficient for identifiability of general transformations.
Study excess risk in statistical inference with transformations.
problem Excess risk in estimating random variables from feature vectors and transformations.
method Characterize lossless transformations, develop test statistics, and information-theoretic bounds.
result Strongly consistent partitioning test statistic for lossless transformations.
In this study, we define a family of ruled surfaces in the Euclidean 3-space E^3 and called similar ruled surfaces. We obtain some properties of these special surfaces and we show that developable ruled surfaces form a family of similar ruled surfaces if and only if the striction curves of the surfaces are similar curv…
ALT transforms time series data for better classification.
problem Efficiently classifying time series data with varying temporal scales.
method ALT algorithm using variable-length shifted time windows.
result State-of-the-art performance with minimal computational overhead.
Whitening, or sphering, is a common preprocessing step in statistical analysis to transform random variables to orthogonality. However, due to rotational freedom there are infinitely many possible whitening procedures. Consequently, there is a diverse range of sphering methods in use, for example based on principal com…
Unified framework for learning function representations using INRs and Transformers.
problem Scalability and efficiency limitations in existing generative models.
method Integrates INRs and Transformer-based hypernetworks into latent variable models.
result Improved scalability, expressiveness, and generalization over existing models.
Noting the importance of the latent variables in inference and learning, we propose a novel framework for autoencoders based on the homeomorphic transformation of latent variables, which could reduce the distance between vectors in the transformed space, while preserving the topological properties of the original space…
In this study, we define a family of null curves in Minkowski 3-space and called null similar curves. We obtain some properties of these special curves. We show that two null curves are null similar curves if and only if these curves form a null Bertrand pair. Moreover, we obtain that the family of null geodesics and n…
The paper models musical motif transformations in Beethoven's works.
problem Understanding how motifs transform in symbolic music.
method Developed a probabilistic framework using Conditional Random Fields.
result Identified patterns of motif transformations and their co-occurrences.
It is widely believed that the prediction accuracy of decision tree models is invariant under any strictly monotone transformation of the individual predictor variables. However, this statement may be false when predicting new observations with values that were not seen in the training-set and are close to the location…
Diagonal transformations preserve independence structures in non-Gaussian distributions.
problem Preserving independence structures in non-Gaussian distributions.
method Diagonal nonlinear transformations of multivariate normal variables.
result Independence structures are preserved in non-Gaussian distributions under diagonal transformations.
ALT improves TSC by capturing complex patterns in time series data.
problem Challenges in traditional TSC methods with time series complexity and variability.
method ALT incorporates variable-length shifted time windows to enhance LLT for better feature representation.
result ALT achieves state-of-the-art performance with few hyperparameters.
ACE models allow flexible conditioning and prediction of latent variables.
problem Lack of flexibility in conditioning and prediction of latent variables in probabilistic models.
method Introduces Amortized Conditioning Engine (ACE) that explicitly represents latent variables and allows runtime conditioning and prediction.
result ACE models outperform existing methods in diverse tasks like image completion, classification, Bayesian optimization, and simulation-based inference.
Optimizes LightGBM for stock market forecasting with novel feature engineering and transformation methods.
problem Accurately forecasting stock market fluctuations to mitigate risks.
method Feature engineering and transformation methods for LightGBM optimization.
result Log Returns, Returns and EMA Difference Ratio are the most effective target variable transformations.
Method identifies latent variables from high-dimensional data with piecewise affine mixing.
problem Identifying latent variables from high-dimensional observations with dependencies and piecewise affine transformations.
method Proposes a two-stage method with sparsity and Gaussianity regularization.
result Effectively recovers ground-truth latent variables from synthetic and image data.
We study a method of reducing space dimension in multi-dimensional Black-Scholes partial differential equations as well as in multi-dimensional parabolic equations. We prove that a multiplicative transformation of space variables in the Black-Scholes partial differential equation reserves the form of Black-Scholes part…
This paper clarifies VAE's property through geometric and information-theoretic interpretations.
problem The transparency of VAE model is an underlying issue.
method Quantitative understanding of VAE through differential geometry and information theory.
result VAE can be mapped to an implicit isometric embedding with a scale factor derived from the posterior parameter.
We introduce Network Maximal Correlation (NMC) as a multivariate measure of nonlinear association among random variables. NMC is defined via an optimization that infers transformations of variables by maximizing aggregate inner products between transformed variables. For finite discrete and jointly Gaussian random vari…
The paper solves pentagon equations using triangulations and edge transformations.
problem Solving pentagon equations with triangulations and edge transformations.
method General data and transformation rule method applied to triangulations.
result Recovery of initial data after transformations.
A novel graph spectral method for mixed categorical and numerical data.
problem Feature learning for mixed data types (numerical and categorical).
method Graph spectral decomposition of the graph Laplacian to model probabilistic dependence structure.
result Increased separability and clusterability of observations in the transformed feature space.
Novel deep learning model for multivariate time series prediction.
problem Challenges in multivariate time series prediction with correlations and complex temporal patterns.
method Temporal Tensor Transformation Network (TTNT) that transforms multivariate time series into tensors for improved feature extraction.
result TTNT outperforms state-of-the-art methods in window-based predictions across various tasks.
The paper uses transformed ANOVA to identify important fire detection variables.
problem Identifying key variables for forest fire detection.
method Developed a complete orthonormal system for standard normal distribution, applied Z-score transformation, and used ANOVA approximation.
result Attribute ranking reveals important variables for fire detection.
The paper tackles causal disentanglement with linear models and interventions.
problem Identify latent variables in a causal model from observed data.
method Use linear transformations and interventions to uniquely identify latent variables.
result A single intervention on each latent variable is sufficient for identifying the latent causal model.
Researchers relax the CVF's smoothness requirement to create more flexible flow models.
problem Challenges in constructing flexible density models due to the CVF's smoothness requirement.
method Introduce L-diffeomorphisms as generalized transformations that may violate smoothness on zero Lebesgue-measure sets. result The relaxation allows for the use of non-smooth activation functions like ReLU in residual flows.
Geometric variations of objects, which do not modify the object class, pose a major challenge for object recognition. These variations could be rigid as well as non-rigid transformations. In this paper, we design a framework for training deformable classifiers, where latent transformation variables are introduced, and …
In this study, we consider the notion of similar ruled surface for timelike and spacelike ruled surfaces in Minkowski 3-space. We obtain some properties of these special surfaces in E_1^3 and we show that developable ruled surfaces in E_1^3 form a family of similar ruled surfaces if and only if the striction curves of …
In this paper we analyze American style of floating strike Asian call options belonging to the class of financial derivatives whose payoff diagram depends not only on the underlying asset price but also on the path average of underlying asset prices over some predetermined time interval. The mathematical model for the …
Deep equilibrium models estimate latent variables from data.
problem Estimating latent variables from data.
method Generalized exponential family models, deep equilibrium networks.
result Deep equilibrium models solve MAP estimates for latent and transformation parameters.
A censored transformed model for proportional outcomes with boundary mass and an application to loss given default modeling.
problem Modeling proportional outcomes with boundary mass in loss given default (LGD) modeling.
method Zero-one censored transformed normal (ZOC-TN) model.
result Captures a wider range of qualitative density shapes than benchmark models while being parsimonious, computationally efficient, and numerically stable.