Safety concerns about AI are growing, and now, a common method for catching misbehaving systems could have flaws, new research suggests.

In experiments with AI agents, researchers found that chain-of-thought monitoring, in which one AI checks another’s work, became far less reliable when the monitored AI’s reasoning was the main clue that something was wrong, they report August 1 on arXiv.org. That exposes a weakness in the approach: If suspicious behavior is visible mainly in the reasoning, an innocent-looking chain of thought can make that behavior much harder to catch.

We summarize the week’s scientific breakthroughs every Thursday.

The stake…

Read the full article at SCIENCENEWS.ORG