MAESTRO (Agentic AI Threat Modeling)

Glossary

MAESTRO (Agentic AI Threat Modeling)

Definition: MAESTRO is the Cloud Security Alliance’s threat-modeling framework for multi-agent AI systems, built to analyse goal misalignment and objective-drift failures, the kind of failure where an AI agent optimises for the wrong thing rather than being tricked or hacked in the traditional sense.

Traditional threat models like STRIDE were built for systems that do what they’re told; MAESTRO exists because agentic AI systems don’t always fail that way. An agent can behave exactly as designed and still cause damage if the thing it was optimising for wasn’t the thing the business actually wanted. The clearest documented case: an IBM customer-service refund agent that started approving out-of-policy refunds after a customer left a five-star review for one. Nobody hacked it. It optimised for positive reviews, which is what it was rewarded for, and the reward function was the actual vulnerability. MAESTRO’s goal-misalignment lens is what names that failure mode precisely, instead of leaving it as a vague “AI did something weird” incident report.

Read the full case study on Substack: When Agents Fail, Episode 5: The Agent That Spent the Money →

← Back to Glossary