
Melanie Plaza: Is Aligned AI Better For Business?
In this episode, James Bowler is joined by Melanie Plaza, Chief Technology Officer at AE Studio, to explore the commercial case for AI alignment, arguing that the techniques researchers care about most are the same ones that make production AI systems actually trustworthy and valuable. Melanie draws on a decade of applied AI delivery to show where alignment and commercial engineering have converged. Getting an agent to do a thing is trivial now; getting it to do the right thing consistently enough to trust in production is not. The gap between those two states is filled by the same tools alignment researchers reach for: rigorous eval suites, red teaming, carefully specified behavior, and guardrails that hold up under adversarial pressure. Melanie's observation is that teams who skip this work don't just expose themselves to safety risk, they fail to get ROI, accumulating what she calls "a sprawl of pilot death." The conversation gets concrete with a case AE Studio has been working on: AI characters, some of them villains, interacting directly with users including children in open-ended multi-turn conversations. Moving from tightly scaffolded, stepwise pipelines to goal-oriented prompting produces better, more natural results, but it also opens a much larger space of possible outputs. The team had to develop a taxonomy of harm categories, including direct user harm, brand harm, and something they call character taboos, illustrated by the example of Peppa Pig cheerfully recommending bacon. That particular output won't end the world, but it points at a real problem: the tail of a generative system is enormous, and standard single-turn test suites won't find what lives out there. Melanie and James also work through the economics of open-weight models, noting that as models like GLM 5.2 and Kimi K3 close the capability gap, the cost and dependency arguments for closed-source APIs weaken. That opens the door to fine-tuning, RL, and white-box techniques like activation steering, which could address problems that prompt-level guardrails handle poorly. It also shifts safety responsibility onto the deploying organization, since the model-level protections that frontier labs build in by default no longer come for free. The upside is high; so is the risk if teams treat alignment as an afterthought. In this episode: * Why the hardest problems in applied AI deployment are alignment problems under a different name * How the shift from scaffolded pipelines to goal-oriented, multi-agent systems removed a de facto safety constraint that few teams replaced * The harm taxonomy AE Studio built for AI characters interacting with children, including the "character taboos" category * Why standard single-turn eval suites miss the long tail of multi-turn conversations * The economic case for open-weight models and what safety responsibilities transfer to the deploying organization when teams move off closed-source APIs * How techniques like activation steering and gradient routing could become reusable deployment primitives, not just research artifacts * Why alignment is not a drag on commercial progress but the thing most likely to produce the next round of genuine capability unlocks Learn more: https://ae.studio/alignment AE Studio is hiring: https://www.ae.studio/join-us Subscribe to our newsletter: https://aestudio.beehiiv.com/ James Bowler LinkedIn: https://www.linkedin.com/in/james-bowler-84b02a100/ Melanie Plaza LinkedIn: https://www.linkedin.com/in/melplaza/ Contact us: alignment@ae.studio
