The 'GPT-6 left notes to itself' behavior could be explained by OpenAI employing Reinforcement Learning over outcomes for large swarms of cooperating agents. This approach would involve reinforcing traces from rollouts of 40, 400, or 4000 agents if success is achieved.