Challenge: Prior research on the representational capabilities of LLMs evaluates whether they show human-level performance.
Approach: They ask how well popular LLMs capture the magnitudes of numbers from a behavioral lens.
Outcome: The proposed model captures the magnitudes of numbers from a behavioral lens.

Similar Papers

Language Models Learn Universal Representations of Numbers and Here’s Why You Should Care (2026.acl-long)

Copied to clipboard

Challenge: Prior work has shown that large language models (LLMs) often converge to accurate input embedding for numbers, based on sinusoidal representations.
Approach: They show that large language models often converge to accurate input embedding for numbers, based on sinusoidal representations.
Outcome: The proposed representations are strikingly systematic, and are interchangeable in a large swathe of experimental setups.
Do Large Language Models Mirror Cognitive Language Processing? (2025.coling-main)

Copied to clipboard

Challenge: Large language models have demonstrated remarkable abilities in text comprehension and logical reasoning.
Approach: They employ Representational Similarity Analysis to measure alignment between 23 LLMs and fMRI signals of the brain.
Outcome: The results show that training strategies affect the LLM-brain alignment.
LLMs Know More About Numbers than They Can Say (2026.eacl-short)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly used in mathematical, scientific, financial and engineering domains.
Approach: They probe the hidden states of several smaller open-source LLMs to find out how big they are .
Outcome: The proposed model improves verbalized accuracy by 3.22% over base models.
Language Models Encode Numbers Using Digit Representations in Base 10 (2025.naacl-short)

Copied to clipboard

Challenge: Large language models (LLMs) often make errors when handling simple numerical tasks . a natural hypothesis is that these errors stem from how LLMs represent numbers .
Approach: They propose to examine how LLMs represent numbers with circular representations per digit . they propose to use digit-wise representations to shed light on errors on numerical tasks .
Outcome: The proposed model is internally represented with individual circular representations per-digit in base 10 . the proposed model could be used to analyze numerical mechanisms in large language models .
Large Language Models: The Need for Nuance in Current Debates and a Pragmatic Perspective on Understanding (2023.emnlp-main)

Copied to clipboard

Challenge: Current Large Language Models (LLMs) are unparalleled in their ability to generate grammatically correct, fluent text.
Approach: They argue that LLMs only parrot statistical patterns in training data and that language learning in LLM cannot inform human language learning.
Outcome: The proposed model can generate grammatically correct, fluent text without requiring human intervention.
NUMCoT: Numerals and Units of Measurement in Chain-of-Thought Reasoning using Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Existing LLMs are not able to handle numerals and units of measurement, but they can be improved by introducing perturbations.
Approach: They propose to analyze existing LLMs on processing numerals and units of measurement by perturbing their datasets.
Outcome: The proposed model improves on ancient Chinese arithmetic problems and can handle numeral and measurement conversions.
Scaling Behavior for Large Language Models regarding Numeral Systems: An Example using Pythia (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are struggling with performing numeric operations accurately.
Approach: They propose to use different numeral systems to scale different numerates in transformer-based large language models.
Outcome: The proposed model is more data-efficient than base 10 and base 10 3 . the model is also more efficient on addition and multiplication .
1,729 vs. 1729: The Effect of Scripts and Formats on LLM Numeracy (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have impressive proficiency in basic arithmetic, but little attention has been given to how they perform when numerical expressions deviate from the prevailing conventions present in their training corpora.
Approach: They investigate numerical reasoning across a wide range of numeral scripts and formats . they show that LLM accuracy drops substantially when numerical inputs are rendered in underrepresented scripts or formats despite the underlying mathematical reasoning being identical .
Outcome: The proposed methods can narrow the gap between LLMs and human models when they deviate from prevailing numerical conventions.
Measuring scalar constructs in social science with LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Valid scalar measurement of skalar constructs is a fundamental task in text analysis.
Approach: They evaluate four approaches to measuring scalar constructs using large language models . pairwise comparisons produced better measurements than prompting LLMs, they say . validation of skalar measurement enables wide range of substantive applications in social science research .
Outcome: The proposed methods improve on pairwise comparisons and finetuning . the proposed methods can be used in social science research .
Do Large Language Models Know How Much They Know? (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models are highly capable systems, but their capabilities and limitations are unclear.
Approach: They develop a benchmark that challenges LLMs to recall all information they possess on specific topics.
Outcome: The proposed model can recall excessive, insufficient, or the precise amount of information they possess on a given topic, indicating their awareness of how much they know about the given topic.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations