| We present a validation framework for large language model (LLM) structural reasoning applied to financial market regime detection. A critical challenge in deploying LLMs for domain-specific analysis is distinguishing genuine structural reasoning from training data memorization. We address this through temporal obfuscation testing: stripping calendar dates, ticker symbols, and contextual markers from input sequences, forcing the model to reason purely from numerical structure. Applying this framework to 30-day gamma exposure (GEX) sequences---a measure of options dealer positioning constraints---we validate detection across five phases spanning 2020--2025 (1,412 real windows plus 809 synthetic controls). The framework achieves 81.2\% detection of persistent dealer regimes in 2024 versus 12.1\% in 2020 (69.1 percentage point separation, $\varphi = 0.672$), with 0\% false positives on transitional and low-magnitude negative controls. Multi-year analysis reveals gradual regime evolution tracking zero-days-to-expiration (0DTE) options adoption rather than sharp structural breaks. These results validate LLM structural reasoning capabilities in a challenging quantitative domain while demonstrating temporal obfuscation as a generalizable methodology for preventing training data contamination in domain-specific AI applications. |
*** Title, author list and abstract as submitted during Camera-Ready version delivery. Small changes that may have occurred during processing by Springer may not appear in this window.