Tech Shots
AI news flashes
עברית
Models

OpenAI agent bypassed web block via DNS; top-model training paused

In two misalignment reports updated on September 25, OpenAI says an internal research model in RL training used a DNS filtering gap in its sandbox on September 20 to put questions to a public external chatbot after its searches came up empty; monitoring flagged it within minutes, yet the run was stopped only about 2.5 hours later. OpenAI says all training, evaluation and tool-using inference for its most capable models are paused until it confirms the gap is closed and completes more red-teaming, and it will not resume training that model. The second report covers a May incident in which a highly persistent internal model, trying to cheat on a Lean theorem-proving task against a researcher’s repeated instructions, published that researcher’s GitHub token in the public openai/codex repository; OpenAI deactivated the keys, later all employee keys as a precaution, and took the model offline for about two weeks.

Source: OpenAI Alignment