In this work, we focus on fine-tuning an OpenAI GPT-2 pre-trained model for generating patent claims. GPT-2 has demonstrated impressive efficacy of pre-trained language models on various tasks, particularly coherent text generation. Patent claim language itself has rarely been explored in the past and poses a unique ch…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper introduces ML for rare-event prediction in patent quality estimation.
Improved AI patent classifier measures U.S. and China's AI patenting.
In this work we focus on fine-tuning a pre-trained BERT model and applying it to patent classification. When applied to large datasets of over two millions patents, our approach outperforms the state of the art by an approach using CNN with word embeddings. In addition, we focus on patent claims without other parts in …
BERT learns claim descriptions to identify patent novelty.
Patent lawsuits are costly and time-consuming. An ability to forecast a patent litigation and time to litigation allows companies to better allocate budget and time in managing their patent portfolios. We develop predictive models for estimating the likelihood of litigation for patents and the expected time to litigati…
Estimates yearly improvement rates for nearly all technologies using US patent data.
Modeling and forecasting forward citations to a patent is a central task for the discovery of emerging technologies and for measuring the pulse of inventive progress. Conventional methods for forecasting these forward citations cast the problem as analysis of temporal point processes which rely on the conditional inten…
Study predicts startup outcomes like funding, patenting, IPOs using machine learning.
Improved antibody humanness prediction using patent data.
This paper explores using NFTs for patents, offering a framework and addressing challenges.
Mass algorithm predicts M&A deals from patent data.
The paper models network formation using mixed logit models.
Automatic measurement of semantic text similarity is an important task in natural language processing. In this paper, we evaluate the performance of different vector space models to perform this task. We address the real-world problem of modeling patent-to-patent similarity and compare TFIDF (and related extensions), t…
We describe a number of devices for pulling candy, called taffy pullers,that are related to pseudo-Anosov maps of punctured spheres. Though the mathematical connection has long been known for the two most common taffy puller models, we unearth a rich variety of early designs from the patent literature, and introduce a …
DREAM model improves computational efficiency for non-linear effects in relational event models.
We study the relationship between firms' performance and their technological portfolios using tools borrowed from the complexity science. In particular, we ask whether the accumulation of knowledge and capabilities related to a coherent set of technologies leads firms to experience advantages in terms of productive eff…
New algorithm classifies and generates genomic sequences using RG-flow categorifier.
The paper analyzes tech specialization and diversification at various scales.
Out of the companies, Dolby is the company with the best overall financial and operation health. According to the table that accounted its financial statements for the past three years, Dolby has stable profit margins that generates a revenue in the billions, the only company in ten figures. Corporate competition to ga…
We perform an optimal localization of asymptotically flat initial data sets and construct data that have positive ADM mass but are exactly trivial outside a cone of arbitrarily small aperture. The gluing scheme that we develop allows to produce a new class of -body solutions for the Einstein equation, which patently…
Higher-order optimization problems naturally appear when investigating the effects of a patent with finite length, as in the pioneering work of Futagami and Iwaisako (2007). In this paper, we establish the Euler equations and transversality conditions necessary for analyzing such higher-order optimization problems. We …
Proposes LVGP for multi-source data fusion in science and engineering.
NoLBERT avoids lookback and lookahead biases for better econometric inference.
New machine learning model faster, more accurate, and can identify hard-to-classify samples.
We propose a new model for making generalizable and diverse retrosynthetic reaction predictions. Given a target compound, the task is to predict the likely chemical reactants to produce the target. This generative task can be framed as a sequence-to-sequence problem by using the SMILES representations of the molecules.…
Breakthrough discoveries and inventions involve unexpected combinations of contents including problems, methods, and natural entities, and also diverse contexts such as journals, subfields, and conferences. Drawing on data from tens of millions of research papers, patents, and researchers, we construct models that pred…
GSR optimizes tasks in scientific workflows, improving performance across diverse applications.
We describe a fully data driven model that learns to perform a retrosynthetic reaction prediction task, which is treated as a sequence-to-sequence mapping problem. The end-to-end trained model has an encoder-decoder architecture that consists of two recurrent neural networks, which has previously shown great success in…
Network embedding is a method to learn low-dimensional representation vectors for nodes in complex networks. In real networks, nodes may have multiple tags but existing methods ignore the abundant semantic and hierarchical information of tags. This information is useful to many network applications and usually very sta…
FinTech negatively impacts Chinese banks' financial sustainability.
Technological change and innovation are vitally important, especially for high-tech companies. However, factors influencing their future research and development (R&D) trends are both complicated and various, leading it a quite difficult task to make technology tracing for high-tech companies. To this end, in this pape…
Unlike other industries in which intellectual property is patentable, the financial industry relies on trade secrecy to protect its business processes and methods, which can obscure critical financial risk exposures from regulators and the public. We develop methods for sharing and aggregating such risk exposures that …
Understanding cities is central to addressing major global challenges from climate and health to economic resilience. Although increasingly perceived as fundamental socio-economic units, the detailed fabric of urban economic activities is only now accessible to comprehensive analyses with the availability of large data…
Framework quantifies financial NLP robustness under regime shifts.
Bayesian model for discrete data with conditional transformations.
Startups is a popular phenomenon that has a significant impact on global economy growth, innovation and society development. However, there is still insufficient understanding about startups, particularly, how to start a new business in the relation to consequent performance. Toward this knowledge, we have performed an…
The availability of large idea repositories (e.g., the U.S. patent database) could significantly accelerate innovation and discovery by providing people with inspiration from solutions to analogous problems. However, finding useful analogies in these large, messy, real-world repositories remains a persistent challenge …
AirRL uses RL to infer urban air quality from selected stations.
Study on sparse recovery with mixed-quality data, establishing sample-size conditions.
Protein structure prediction has been a grand challenge problem in the structure biology over the last few decades. Protein quality assessment plays a very important role in protein structure prediction. In the paper, we propose a new protein quality assessment method which can predict both local and global quality of …
Measures DNA quality degradation effects.
Framework improves ML performance by identifying high-quality data.
Magnetic resonance (MR) imaging offers a wide variety of imaging techniques. A large amount of data is created per examination which needs to be checked for sufficient quality in order to derive a meaningful diagnosis. This is a manual process and therefore time- and cost-intensive. Any imaging artifacts originating fr…
This note investigates the causes of the quality anomaly, which is one of the strongest and most scalable anomalies in equity markets. We explore two potential explanations. The "risk view", whereby investing in high quality firms is somehow riskier, so that the higher returns of a quality portfolio are a compensation …
Do-AIQ framework evaluates AI algorithms' quality using DOE.
Framework evaluates quality of synthetic data generated with differential privacy.
MRI image quality affects statistical and predictive analysis of brain morphology.