Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

316292123 · Oct 202519922001200920172026
48 results for Korean financial texts

Study introduces KorFinMTEB for Korean financial texts, revealing model limitations.

problem Limited evaluation benchmarks for low-resource domains, especially Korean.
method Developed KorFinMTEB, a tailored benchmark for Korean financial texts.
result Models perform better on translated benchmarks than on domain-specific ones.

Proposes RDASS for better Korean text summarization evaluation.

problem ROUGE scores fail to capture semantic meaning in Korean text summarization.
method Introduces RDASS metrics and a method to improve their correlation with human judgment.
result RDASS metrics correlate better with human judgment than ROUGE scores.

Study finds financial YouTube channel 3PROTV predicts stock market performance and sentiment changes.

problem Determining the informational value of financial YouTube channels.
method Analyzing 3PROTV's content and its impact on stock market performance and sentiment.
result 3PROTV's content, particularly negative sentiment, predicts stock market performance and sentiment changes.

We study the effect of globalization on the Korean market, one of the emerging markets. Some characteristics of the Korean market are different from those of the mature market according to the latest market data, and this is due to the influence of foreign markets or investors. We concentrate on the market network stru…

2005-09-13abs ↗pdf ↗

Improved sentiment analysis in Korean finance using masked PLMs.

problem Fine-grained sentiment analysis in non-English finance literature is lacking due to limited annotated data.
method Developed KorFinASC dataset and applied TGT-Masking to PLMs to remove non-stationary knowledge.
result Improved classification accuracy by 22.63% on KorFinASC.

We study the dynamical behavior of high-frequency data from the Korean Stock Price Index (KOSPI) using the movement of returns in Korean financial markets. The dynamical behavior for a binarized series of our models is not completely random. The conditional probability is numerically estimated from a return series of K…

2005-12-23abs ↗pdf ↗

In this paper, we studied the dynamics of the log-return distribution of the Korean Composition Stock Price Index (KOSPI) from 1992 to 2004. Based on the microscopic spin model, we found that while the index during the late 1990s showed a power-law distribution, the distribution in the early 2000s was exponential. This…

2005-11-14abs ↗pdf ↗

We apply the formalism of the continuous time random walk (CTRW) theory to financial tick data of the bond futures transacted in Korean Futures Exchange (KOFEX) market. For our case, the tick dynamical behaviors of the returns and volatility for bond futures are treated particularly at the long-time limit. The volatili…

2003-11-07abs ↗pdf ↗

In this study, we establish a network structure of the Korean stock market, one of the emerging markets, with its minimum spanning tree through the correlation matrix. Base on this analysis, it is found that the Korean stock market doesn't form the clusters of the business sectors or of the industry categories. When th…

2005-04-01abs ↗pdf ↗

The multifractal behavior for tick data of prices is investigated in Korean financial market. Using the rescaled range analysis(R/S analysis), we show the multifractal nature of returns for the won-dollar exchange rate and the KOSPI. We also estimate the Hurst exponent and the generalized qqth-order Hurst exponent in …

2003-05-13abs ↗pdf ↗

We study the evolution of probability distribution functions of returns, from the tick data of the Korean treasury bond (KTB) futures and the S$&$P 500 stock index, which can be described by means of the Fokker-Planck equation. We show that the Fokker-Planck equation and the Langevin equation from the estimated Kramers…

2005-12-22abs ↗pdf ↗

ClovaCall introduces a new Korean call speech corpus for contact centers.

problem Lack of large-scale call-based speech corpora for Korean dialog scenarios.
method Development of a new large-scale Korean call-based speech corpus (ClovaCall) in a restaurant reservation domain.
result Validation of the dataset with ASR models shows its effectiveness.

The herd behaviors of returns for the won-dollar exchange rate and the KOSPI are analyzed in Korean financial markets. It is shown that the probability distribution P(R)P(R) of price returns RR for three values of the herding parameter tends to a power-law behavior P(R)RβP(R) \simeq R^{-β} with the exponents β=2.2 β=2.2(the wo…

2003-04-21abs ↗pdf ↗

We introduce the minority game theory for two kinds of the Korean treasury bond (KTB) in Korean futures exchange markets. Since we discuss numerically the standard deviation and the global efficiency for an arbitrary strategy, our case is found to be approximate to the majority game. Our result presented will be compar…

2005-03-01abs ↗pdf ↗

We investigate the distribution function and the cumulative probability for Korean household incomes, i.e., the current, labor, and property incomes. For our case, the distribution functions are consistent with a power law. It is also showed that the probability density of income growth rates almost has the form of a e…

2004-03-05abs ↗pdf ↗

We consider returns of two Korean stock market indices, KOSPI and KOSDAQ index. Central parts of the probability distribution function of returns are well fitted by the Lorentzian distribution function. However, tail parts of the probability distribution function follow a power law behavior well. We found that the prob…

2004-07-16abs ↗pdf ↗

New distress dictionary improves bankruptcy prediction from disclosure text.

problem Bankruptcy prediction from financial disclosures.
method Proposes a distress dictionary based on managers' sentences, quantifies linguistic features, and builds predictive models.
result Predictive models based on the distress dictionary outperform existing methods.

Anonymization reduces economic signal extraction from financial texts.

problem Reducing meaningful economic signals from financial texts due to anonymization.
method Analyzed the impact of anonymization on textual understanding and economic signal extraction.
result Information loss due to anonymization is severe and pervasive, outweighing its benefits in certain financial applications.

Investor flows in Korean equity market transmit shared information, not private signals.

problem Whether investor flows transmit private information or only public signals.
method Transfer Entropy networks constructed from investor-type flows over umNDates{} trading days.
result Investor flows transmit shared information, not private signals.

System detects relevant financial news and predictions from unstructured text.

problem Manual extraction of relevant financial information from news is cumbersome and error-prone.
method Topic modeling with LDA, co-reference resolution, multi-paragraph segmentation, and temporal analysis.
result ROUGE-L values for relevant text and predictions/forecasts were 0.662 and 0.982, respectively.

FINCH dataset enables financial Text-to-SQL tasks, improving model evaluation.

problem Lack of a large-scale financial dataset for Text-to-SQL tasks.
method Curated financial dataset (FINCH), benchmarking reasoning and language models.
result Proposed FINCH Score for more accurate financial model evaluation.

BioFinBERT analyzes sentiment of biotech press releases and financial text around inflection points.

problem Analyzing sentiment of biotech press releases and financial text around inflection points.
method Finetuning BioBERT on financial datasets to create BioFinBERT for sentiment analysis.
result BioFinBERT accurately analyzes sentiment of biotech press releases and financial text around inflection points.

Enhanced regime shifts detection using unstructured text and financial data.

problem Detecting regime shifts in financial markets is challenging due to noisy and multicollinear data.
method Combines LLM reasoning on unstructured text and statistical validation on financial time series.
result Framework achieves F1 score of 0.82, outperforming pure data-driven methods.

Study evaluates if LLMs have company-specific biases in financial sentiment analysis.

problem Evaluating if large language models exhibit company-specific biases in financial sentiment analysis.
method Comparing sentiment scores with and without company names, constructing economic models, and empirical analysis.
result LLMs show company-specific biases in sentiment analysis, impacting investor behavior and stock prices.

Paper introduces NumLLM for better financial text understanding with numeric variables.

problem Poor performance of existing financial large language models in numeric financial text.
method Constructed financial corpus, fine-tuned with LoRA modules, merged into foundation model.
result NumLLM achieves best performance on financial question-answering benchmark, especially with numeric questions.

Simple feature engineering beats complex models in financial prediction.

problem Understanding when complex models outperform simple alternatives in financial prediction.
method Independent Component Analysis (ICA), Wavelet Coherence, Long Short-Term Memory (LSTM) networks with attention mechanisms.
result A simple linear model using normalized flows achieves superior returns compared to complex models.

We investigated the temporally evolving network structures of the Japanese and Korean stock markets through the minimum spanning trees composed of listed stocks. We tested the validity of conventional grouping by industrial categories, and found a common trend of decrease for Japan and Korea. This phenomenon supports t…

2005-11-27abs ↗pdf ↗

UniFinEval benchmarks financial models across text, images, and videos.

problem Challenges in evaluating financial multimodal models across text, images, and videos.
method Proposes UniFinEval, a unified multimodal benchmark for financial scenarios.
result Gemini-3-pro-preview achieves best performance but still lags behind experts.

BERTopic improves financial text analysis with FinTextSim's contextual embeddings.

problem Analyzing financial text data for insights and predictions.
method Integrates BERTopic with FinTextSim for topic modeling and clustering.
result BERTopic performs better with FinTextSim's embeddings, improving topic clarity and reducing misclassification.

FinGPT democratizes financial data for LLMs, enabling innovation.

problem Limited financial text datasets and disparities between general and financial text data.
method Automates collection and curation of real-time financial data from diverse Internet sources, fine-tuning with RLSP and LoRA.
result Democratizes access to financial data for LLMs, enabling innovation.

InvestLM is a financial domain LLM tuned on LLaMA-65B for investment advice.

problem Improving financial text understanding and advice generation for investment.
method Curated financial instruction dataset, LLaMA-65B, less-is-more-for-alignment approach.
result InvestLM provides comparable responses to state-of-the-art commercial models.

This study uses deep learning to analyze stock market sentiment from financial forums.

problem Improving stock market prediction accuracy through emotional analysis.
method Crawling financial forum data, training Bert model on financial corpus, using MIC for comparison.
result BERT model's emotional analysis of financial texts correlates with stock market fluctuations.

We study the tick dynamical behavior of the bond futures in Korean Futures Exchange(KOFEX) market. Since the survival probability in the continuous-time random walk theory is applied to the bond futures transaction, the form of the decay function in our bond futures model is discussed from two kinds of Korean Treasury …

2002-12-17abs ↗pdf ↗

Qwen3-8B outperforms classical models in financial text classification.

problem Financial text classification for trading systems and sentiment analysis.
method Noisy Embedding Instruction Finetuning and Rank-stabilized Low-Rank Adaptation.
result Qwen3-8B achieves better classification accuracy and fewer training epochs.