- Patching the answer lookback pointer flips the final output from coffee to beer (pink line)
- Patching the answer lookback payload shifts it from coffee to tea (grey line)
Strong evidence that the Answer Lookback mechanism is real!
- Patching the answer lookback pointer flips the final output from coffee to beer (pink line)
- Patching the answer lookback payload shifts it from coffee to tea (grey line)
Strong evidence that the Answer Lookback mechanism is real!
We reverse-engineered how LLaMA-3-70B-Instruct handles a belief-tracking task and found something surprising: it uses mechanisms strikingly similar to pointer variables in C programming!
We reverse-engineered how LLaMA-3-70B-Instruct handles a belief-tracking task and found something surprising: it uses mechanisms strikingly similar to pointer variables in C programming!