Open-weight LLMs encode and steer physical laws
Markus J. Buehler demonstrates that open-weight language models internally represent physical mechanisms of materials science rather than relying solely on surface text patterns. Using Jacobian lenses and causal activation patching, the study shows these internal representations can be read directly and dynamically steered to control physical reasoning outputs.
Advancing mechanistic interpretability into complex domain-specific physics shows that LLMs internalize structured world models rather than relying purely on shallow statistical association.
- –Identifies concept readability, constitutive orientation, and causal control over materials science mechanisms within model hidden states.
- –Utilizes Jacobian lenses and direct readouts to decode physical mechanism families without requiring predefined target word sets.
- –Tests internal physical reasoning against a 60-law counterfactual benchmark to evaluate how state transformations track inverted physical laws.
- –Demonstrates that activation patching can steer internal representations to shift model predictions toward physically accurate outcomes.
DISCOVERED
2h ago
2026-07-23
PUBLISHED
3h ago
2026-07-23
RELEVANCE
AUTHOR
_akhaliq