Papers by Kumiko Tanaka-Ishii
Stock Embeddings Acquired from News Articles and Price History, and an Application to Portfolio Optimization (2020.acl-main)
Copied to clipboard
| Challenge: | Recent studies have shown that news articles can be leveraged to improve price prediction. |
| Approach: | They propose a method to encode the influence of news articles through a vector representation of stocks . they use a deep learning framework to acquire the vector representation using news articles and price history . |
| Outcome: | The proposed method can be applied to other financial problems besides price prediction. |
A New Formulation of Zipf’s Meaning-Frequency Law through Contextual Diversity (2025.acl-long)
Copied to clipboard
| Challenge: | Existing studies have examined Zipf's meaning-frequency law as a relationship between word frequency and the number of meanings based on contextualized word vectors . |
| Approach: | They propose to use word frequency as a relationship between word frequency and contextual diversity to examine Zipf's meaning-frequency law for a wider variety of words and corpora than previous studies have shown. |
| Outcome: | The proposed formulation gives a new interpretation of Zipf's meaning-frequency law and enables us to examine it for a wider variety of words and corpora than previous studies have shown. |
Taylor’s law for Human Linguistic Sequences (P18-1)
Copied to clipboard
| Challenge: | Taylor's law characterizes how the variance of the number of events for a given time and space grows with respect to the mean, forming a power law. |
| Approach: | They propose a method to quantify Taylor's law in natural language and conduct Taylor analysis of over 1100 texts across 14 languages. |
| Outcome: | The proposed method is able to quantify the complexity of linguistic time series and evaluate language models. |
Repeated Sequences Reveal Gaps between Large Language Models and Natural Language (2026.acl-long)
Copied to clipboard
| Challenge: | Existing evaluation methods provide limited insight into the long-range organization of generated text. |
| Approach: | They propose a framework for evaluation based on repeatedsubsequences . they compare their distribution across scales and their results to Rényi entropies . |
| Outcome: | The proposed framework relates distribution of results to higher-order Rényi entropies on human-written and length-matched GPT-generated texts. |