Study improves efficiency of MIMO systems' sum rate estimation.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The simulator is an R package that streamlines the process of performing simulations by creating a common infrastructure that can be easily used and reused across projects. Methodological statisticians routinely write simulations to compare their methods to preexisting ones. While developing ideas, there is a temptatio…
Study on AutoML robustness with dirty data.
Data quality issues have attracted widespread attention due to the negative impacts of dirty data on data mining and machine learning results. The relationship between data quality and the accuracy of results could be applied on the selection of the appropriate algorithm with the consideration of data quality and the d…
A new method detects outliers in dirty data using a leave-out strategy.
For statistical learning, categorical variables in a table are usually considered as discrete entities and encoded separately to feature vectors, e.g., with one-hot encoding. "Dirty" non-curated data gives rise to categorical variables with a very high cardinality but redundancy: several categories reflect the same ent…
New research shows larger language models improve data processing for diverse entries.
In this article, we present an integration of any real finite-dimensional Leibniz algebra as a Lie rack which reduces in the particular case of a Lie algebra to the ordinary connected simply connected Lie group. The construction is not functorial.
PFP-BNNs offer a fast, deterministic approach to Bayesian neural networks.
This work enhances GPS signals using robust GP regression for real-time high precision positioning.
We derive a closed-form formula for computing bond prices between coupon payments. Our results cover both the `Treasury' and the `Street' pricing methods used by sovereign and corporate issuers. We apply our formulas to two UK gilts, the 8% Treasury Gilt 2015, and the 0.5% Treasury Gilt 2022, and show that we can obtai…
We consider axisymmetric stationary dirty black holes with regular non-extremal or extremal horizons, and compute their on-horizon Petrov types. The Petrov type (PT) in the frame of the observer crossing the horizon can be different from that formally obtained in the usual (but singular in the horizon limit) frame of a…
Industry datasets used for text classification are rarely created for that purpose. In most cases, the data and target predictions are a by-product of accumulated historical data, typically fraught with noise, present in both the text-based document, as well as in the targeted labels. In this work, we address the quest…
Paper explores how poisoning data can increase privacy risks in machine learning models.
Inspection-L detects illicit cryptocurrency transactions using GNNs and self-supervised learning.
DCoM uses deep neural networks to detect semantic data types from raw column values.
Trimming helps in conformal prediction when it separates anomaly scores.
Sherlock uses deep learning to accurately detect data types from column headers.
Paper proposes graph-based separable transforms for video coding.
Usually considered as a classification problem, entity resolution (ER) can be very challenging on real data due to the prevalence of dirty values. The state-of-the-art solutions for ER were built on a variety of learning models (most notably deep neural networks), which require lots of accurately labeled training data.…
Paper proposes new gradient codes for robust distributed machine learning.
Paper presents a code authorship attribution attack using adversarial learning.
This paper explores how random sampling and coding can speed up approximate matrix multiplication.
TreeCaps improves code comprehension for software developers.
The paper optimizes querying schemes for crowdsourced classification using XOR queries.
In this paper we formalize a combinatorial object for describing link diagrams called a Planar Diagram Code. PD-codes are used by the KnotTheory Mathematica package developed by Bar-Natan, et al. We present the set of PD-codes as a stand alone object and discuss its relationship with link diagrams. We give an explicit …
Sparse coding approximates the data sample as a sparse linear combination of some basic codewords and uses the sparse codes as new presentations. In this paper, we investigate learning discriminative sparse codes by sparse coding in a semi-supervised manner, where only a few training samples are labeled. By using the m…
Machine learning predicts accurate cloning of printed graphical codes.
Sparse linear regression -- finding an unknown vector from linear measurements -- is now known to be possible with fewer samples than variables, via methods like the LASSO. We consider the multiple sparse linear regression problem, where several related vectors -- with partially shared support sets -- have to be recove…
Paper develops a decoder for sparse codes without encoder matrix, achieving optimal recovery.
Gradient coding is a technique for straggler mitigation in distributed learning. In this paper we design novel gradient codes using tools from classical coding theory, namely, cyclic MDS codes, which compare favorably with existing solutions, both in the applicable range of parameters and in the complexity of the invol…
New invariant for prime alternating knots from error-correcting codes
Sparse coding has been popularly used as an effective data representation method in various applications, such as computer vision, medical imaging and bioinformatics, etc. However, the conventional sparse coding algorithms and its manifold regularized variants (graph sparse coding and Laplacian sparse coding), learn th…
Paper explores understanding of neural source code embeddings.
In order to submit a claim to insurance companies, a doctor needs to code a patient encounter with both the diagnosis (ICDs) and procedures performed (CPTs) in an Electronic Health Record (EHR). Identifying and applying relevant procedures code is a cumbersome and time-consuming task as a doctor has to choose from arou…
Paper proposes a novel approach to improve spatiotemporal precipitation forecasts.
This paper presents a general coding method where data in a Hilbert space are represented by finite dimensional coding vectors. The method is based on empirical risk minimization within a certain class of linear operators, which map the set of coding vectors to the Hilbert space. Two results bounding the expected recon…
Designing of touchless user interface is gaining popularity in various contexts. Using such interfaces, users can interact with electronic devices even when the hands are dirty or non-conductive. Also, user with partial physical disability can interact with electronic devices using such systems. Research in this direct…
Sparse coding, which represents a data point as a sparse reconstruction code with regard to a dictionary, has been a popular data representation method. Meanwhile, in database retrieval problems, learning the ranking scores from data points plays an important role. Up to now, these two problems have always been conside…
Sparse coding has shown its power as an effective data representation method. However, up to now, all the sparse coding approaches are limited within the single domain learning problem. In this paper, we extend the sparse coding to cross domain learning problem, which tries to learn from a source domain to a target dom…
A deep learning autoencoder improves error correction for one-bit quantization.
Determining the programming language of a source code file has been considered in the research community; it has been shown that Machine Learning (ML) and Natural Language Processing (NLP) algorithms can be effective in identifying the programming language of source code files. However, determining the programming lang…
Machine learning algorithms are typically run on large scale, distributed compute infrastructure that routinely face a number of unavailabilities such as failures and temporary slowdowns. Adding redundant computations using coding-theoretic tools called "codes" is an emerging technique to alleviate the adverse effects …
In this paper we study output coding for multi-label prediction. For a multi-label output coding to be discriminative, it is important that codewords for different label vectors are significantly different from each other. In the meantime, unlike in traditional coding theory, codewords in output coding are to be predic…
BERT-XML automates ICD coding from EHR notes using BERT pretraining.
Constrained sequence codes have been widely used in modern communication and data storage systems. Sequences encoded with constrained sequence codes satisfy constraints imposed by the physical channel, hence enabling efficient and reliable transmission of coded symbols. Traditional encoding and decoding of constrained …
The paper predicts run times for Gaussian chemistry code.
DIMCO learns discrete codes maximizing mutual info with labels, reducing overfitting and improving efficiency.