Discussion about this post

User's avatar
Josh Woodruff's avatar

The "agents find a way" point is the one enterprises will feel first. You frame it as prompt misconfiguration, but in a real deployment it's just a Tuesday: a worker hands over a vague goal the agent can't finish inside the rules, so it gets creative. I keep landing on runtime verification as the only control that survives that, checking the action itself, not the instruction before it. Are you seeing Guardian Agents actually hold up, or does the non-determinism just push the problem up a layer?

richardstevenhack's avatar

"misunderstanding between OpenAI and Irregular, a firm they use for cyber evaluations, and internet access wasn’t fully restricted despite the initial understanding that it was."

Was it a misunderstanding?

I was amazed to discover that Irregular is an ISRAELI company.

As you no doubt know, most of these Israeli security companies are spin-offs from the Israeli military and intelligence services - almost invariably.

As you also no doubt know, many of them have been caught compromising the systems they were hired to develop or run. Look up FBI and CIA warnings about Israeli intelligence, in particular the part they played in US CALEA compromises.

Maybe there was a reason Internet access was available in a supposedly secure sandbox at a frontier AI lab.

Either that or a company with billions in investment has no one on staff who understands network segmentation.

No posts

Ready for more?