Back to papers
March 19, 2026cs.CLcs.AIAdvanced

What Really Controls Temporal Reasoning in Large Language Models: Tokenisation or Representation of Time?

AI-Generated Summary

This paper introduces MultiTempBench, a benchmark for testing how well large language models handle temporal reasoning (like date arithmetic and timezone conversion) across five languages and different calendar systems. The researchers discovered that how dates are broken into tokens (small text pieces) is crucial for low-resource languages, while high-resource languages rely more on learning temporal patterns, revealing that tokenization quality is a key bottleneck limiting AI's ability to reason about time.

HF Upvotes

1

Difficulty
Advanced
Categories

cs.CL, cs.AI

AI Tags
temporal reasoningmultilingual NLPtokenizationbenchmarkinglarge language modelslow-resource languagesinterpretability