Do LLMs Understand Sequential Structure? A Controlled Study of Inference and Generation
Abstract
Large language models (LLMs) are increasingly used as interactive agents and simulators, yet it remains unclear whether they can recover latent sequential structure beyond surface action frequencies. This distinction is critical for behavioral simulation, where actions are often shaped by prior context rather than marginal frequencies alone. We study this question using controlled two-player Rock--Paper--Scissors interactions and a one-player stochastic n-gram continuation task. Across these experiments, we test whether LLMs can identify latent strategies, follow simple Markov rules, and sustain higher-order conditional dependencies. Our framework separates distribution matching from conditional rule following. Results show that longer context does not improve identification, correct recognition does not ensure faithful simulation, and higher-order dependencies substantially degrade rule recovery. Apparent behavioral fidelity can therefore mask incorrect generative mechanisms.
Community
Do LLMs truly understand sequential behavior, or do they simply reproduce surface-level patterns?
We investigate this question through controlled experiments on strategy identification, Markov rule following, and higher-order sequential dependencies.
Our findings reveal a critical gap between recognizing a strategy and faithfully simulating it: longer context does not necessarily help, and matching action distributions can hide fundamentally incorrect generative mechanisms.
These results challenge how we evaluate LLMs as behavioral simulators and interactive agents.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Are LLMs Good Financial User Simulators? Multi-view Investor Logic Alignment (MILA) (2026)
- SIMLIFE: Pattern Understanding for Long-Horizon Human-Agent Partnership (2026)
- HeadEdit: Calibrating Language Model Behavior Through the Frozen Unembedding Matrix (2026)
- CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes (2026)
- LLMs Learn Better In-Context from Rules than from Examples (2026)
- Lexical Perturbations Disrupt LLM Reasoning: An Empirical Study of Attention Diversion (2026)
- Playing social deduction games with reinforcement fine-tuned large language models (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.04977 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper