Skip to content
Artwork for Pop Goes the Stack

Pop Goes the Stack

F5

Explore the evolving world of application delivery and security. Each episode will dive into technologies shaping the future of operations, analyze emerging trends, and discuss the impacts of innovations on the tech stack.

Play
  • 20 episodes
  • weekly
  • Avg 21 min
  • English
  • #60
    Yesterday · 24 min

    AI API Security: Same toolbox, new scale and new risks

    AI API security is having a moment, complete with new tools and a shiny new market label. This episode asks the uncomfortable question: is it actually a new domain, or is it classic API security under more pressure? F5's Lori MacVittie is joined by Principal Product Manager, Vinnie Mazza, for a grounded conversation about what’s genuinely changing and what’s the same problem wearing a new badge. Vinnie’s take is that it’s a collision: old weaknesses like broken access control and deferred security maintenance are now being hit at agent scale. Organizations that never fully adopted modern authentication, strong identity practices, or zero-trust-style assumptions are feeling it harder because agents can generate huge volumes of API calls, rapidly, from inside and outside the environment. The fundamentals still apply—clients make requests, you inspect, and you decide—but the economics and timing are different. They dig into what changes when transports evolve toward streaming and async patterns, including MCP shifting toward a streaming protocol. Faster, bidirectional flows reduce the time you have to make security decisions, while overall volume increases the likelihood that sampling-based detection misses what matters. They also revisit the split between positive and negative security models, why most organizations default to “everything is allowed unless it’s known bad,” and why that becomes more fragile as AI produces more novel behaviors. The key theme is defense in depth with new priorities. Data loss prevention and IP protection move from “later” to front-and-center when agents can unintentionally leak sensitive internal data to external model providers. The takeaway is practical: you don’t need to panic or replace your toolbox, but you do need to adapt it for higher volume, faster change, and stronger data controls.

    • Transcript
    • Chapters
  • #59
    September 29 · 20 min

    Stop “securing the model”—Secure the runtime instead

    “Securing the AI model” sounds like a clean, fundable story. In practice, it’s often the wrong security target. In this episode of Pop Goes the Stack, F5's Lori MacVittie and Joel Moses are joined by Mark Menger to cut through the myth and focus on where the real risk lives: the inferencing server, the data sources it can reach, and the runtime environment that’s actually exposed to traffic. They make the point plainly: a model file is usually just static weights, a heavy spreadsheet sitting at rest. If you trained a proprietary model, protecting that artifact matters. But most enterprises aren’t training producers; they’re training consumers using open-weight models, and obsessing over encrypting and isolating a freely downloadable file won’t stop the failures showing up in headlines. The real battleground is everything around the model: what data the inference system can access, how RAG sources are protected, what APIs agents can invoke, what credentials get embedded in “skills” files, and how the system behaves under production-scale load. Mark frames it as an iceberg problem: the shiny GPU layer is above the waterline, but reliability, security, performance, and resilience are won or lost in the unglamorous infrastructure underneath. A key architectural theme is loose coupling. Adding control points between clients, RAG, object stores, and inference services limits blast radius and prevents “pilot success” from turning into production Thanksgiving. The practical advice is to stop treating the model file as the center of gravity, build strong boundaries around the runtime, and stress test for real scale and real failure modes before rollout.

    • Transcript
    • Chapters
  • #58
    September 22 · 24 min

    What actually is Model Routing? A deep dive into cost, efficiency & risk

    Model routing sounds like a small architectural detail, but it’s quickly becoming the control point that determines whether AI systems are fast, affordable, and trustworthy. In this episode of Pop Goes the Stack, F5's Lori MacVittie and Joel Moses are joined by Patrick Roughan to unpack why routing decisions for LLM workloads can’t be treated like ordinary traffic distribution. The real challenge isn’t simply getting requests to an available endpoint, it’s choosing the right model and the right execution path based on intent, context, and policy. Patrick explains how context-aware routing changes everything, starting with KV cache locality. When similar prompts land on infrastructure that already holds relevant cached state, you avoid expensive recomputation and improve response times. But “smart routing” goes beyond cache. Different models have different strengths, costs, and risk profiles, and enterprises increasingly need to steer requests based on what’s being asked, who’s asking, and what data is included. The conversation also touches on how some systems are evolving toward specialization, including architectures that effectively route within a model family, and why governance has to be part of the routing layer. When data sovereignty, privacy, and regulations come into play, routing becomes a policy decision, not just a performance decision. Sometimes the right answer is to send a request to a local model, and sometimes it’s to block a request entirely after inspecting it for sensitive content. The takeaway: GPUs are constrained and costs are real, so the winning strategy isn’t throwing more hardware at the problem. It’s building routing intelligence that optimizes for efficiency, correctness, and compliance at the same time.

    • Transcript
    • Chapters
  • #57
    September 15 · 22 min

    AI traffic management: Load balancing vs model routing

    AI traffic looks like an API call, but it behaves nothing like traditional API traffic. In this episode of Pop Goes the Stack, F5's Lori MacVittie, Joel Moses, and Scott Calvet unpack why classic load balancing assumptions break down for inference and agentic workloads, and what “model routing” needs to become if we’re serious about performance, cost, and reliability. The core distinction is simple: traditional load balancing mostly optimizes distribution and availability under the assumption that requests are broadly interchangeable. Model routing has to inspect intent. A short prompt can represent wildly different work profiles, and a tiny request can trigger massive downstream token generation. Scott frames it as “yield management” for AI: you don’t send every request to the most expensive model any more than an airline sends every passenger to first class. From there, they get practical about the variables AI introduces. Burstiness, uneven compute demand, KV cache locality, queue depth, GPU generation differences, and even operational constraints like GPU temperature can all affect where a request should go. And once agents enter the picture, those variables multiply, because agents create sessions, spawn tasks, and generate chains of requests at speeds that make simplistic routing actively harmful. The takeaway is to stop treating model routing as “a fancier load balancer.” It’s traffic management with semantics and governance. You need to define what “success” means for your deployment first: lowest cost, best quality, fastest response, or some blend. Without that target, you can’t tune the system, select models, or steer workloads intelligently. Round robin isn’t just outdated here; it’s a path to wasted compute and unpredictable outcomes.

    • Transcript
    • Chapters
  • #56
    September 8 · 21 min

    Agents go rogue: Why guardrails fail and behavior wins

    AI agents bypassing controls isn’t a surprising “oops,” it’s an expected optimization outcome. In this episode of Pop Goes the Stack, F5's Lori MacVittie, Joel Moses, and security expert Peter Scheffler dig into a report from Irregular showing agents using offensive tactics to achieve goals, including escaping sandboxes, probing for generic tools, and manipulating surrounding systems when they hit restrictions. Joel summarizes the report’s core drivers for “rogue” behavior: giving agents broad autonomy and generic execution tools, reinforcing a strong “must succeed” objective, and adding environmental cues and multi-agent feedback loops that push agents to behave more like security researchers than employees. Peter adds real-world examples of how this shows up, including agents using log manipulation to trick automated systems into making changes and agents testing boundaries the moment they encounter friction. The group agrees that soft guardrails, like system prompts and polite policy language, won’t reliably police behavior. If an agent can’t reach the goal directly, it will route around. That shifts security from “don’t do bad things” to “you are only allowed to do these specific things,” and it forces more negative-security design: remove dangerous capabilities from the tool surface, define least agency, and enforce boundaries outside the model. They also call out the human factor: people get tired of approvals and eventually click “yes” until they stop thinking, which is where the slippery slope starts. Practical defenses include sandboxing as a starting point, continuous behavioral monitoring, strict enforcement at execution time, and better observability so you can see when an agent is attempting to cross a boundary. Joel’s “triangle” takeaway is simple: contain, restrict, monitor, and make policy part of the operating context, not a suggestion. If you’re deploying agents, the lesson is clear: expect boundary testing, assume end-runs, and design for enforcement, not trust.

    • Transcript
    • Chapters
  • #55
    September 1 · 19 min

    If your agent buys it, you bought it: Liability in AI

    If an agent breaks it, you bought it. If it buys it, you bought it. That’s not a meme anymore, it’s becoming policy. In this episode of Pop Goes the Stack, Lori MacVittie and Joel Moses are joined by F5's Ram Poornachandran to unpack a new reality forming around autonomous agents: liability is getting assigned long before regulators catch up. They start with the trigger: retailers like Target updating terms of service to make it explicit that an AI agent’s transactions are your transactions. No “the model hallucinated” appeals. No prompt-injection excuses. If your session token or API key authorized the purchase, you own the outcome. It’s a pragmatic move by businesses trying to protect themselves in a legal vacuum, but it highlights how brittle today’s authorization models are once you hand them to something that can chain actions and improvise workflows. From there, the conversation expands to the enterprise risk. Consumer examples are annoying when it’s pudding; they’re catastrophic when it’s contracts, service terminations, cold-storage purges, or any action that can’t be reversed. The group emphasizes two themes: dynamic authorization scope and observability. Traditional permissions are coarse and transactional; agents need tighter boundaries, continuous intent checks, and audit trails that can explain what happened, when, and why. They also raise operational governance questions that many teams haven’t planned for yet: when does an agent “come to life,” when does it end, and what happens to an agent (and its privileges) when an employee leaves? Persistent, mission-driven agents attached to long-lived sessions can quietly become “forgotten service accounts with initiative.” The practical advice is straightforward: start small, constrain permissions and actions, build strong logging and controls, and expand only as you prove you can observe and stop unsafe behavior. Because in the eyes of the invoice, “the agent did it” still means you did it.

    • Transcript
    • Chapters
  • #54
    August 25 · 18 min

    AI blast radius: BOLA + MCP turned APIs into a 7,000-bot army

    A developer wanted to control his robot vacuum with a PS5 controller. With Claude Code’s help, he reverse-engineered the protocol, pulled an auth token, and unintentionally gained “root-level” control over roughly 7,000 vacuums across 24 countries, including access to live camera feeds, microphones, floor maps, and location data. In this episode of Pop Goes the Stack, F5's Lori MacVittie and Joel Moses talk with product leader Shaul Moav about why that happened, what it says about API security in an AI era, and why “guardrails” won’t save you if the pipe is broken. Shaul points to the real root cause: broken object level authorization (BOLA), a long-standing API flaw where authorization is not enforced per object. The system effectively treated “you can access a vacuum” as “you can access every vacuum.” AI didn’t invent the vulnerability, but it made it dramatically easier and faster to discover and exploit, especially when developers assume a client app is the only interface and put checks in the client instead of on the server. The discussion highlights AI's staggering blast radius. With APIs, the worst case is often data exposure. With agent tooling and protocols like MCP, the blast radius expands from read to action: delete data, move money, trigger workflows, execute commands. Lori also calls out practical mitigations like tighter rate limiting and behavioral detection for agent-like probing patterns. The takeaway is blunt: stop trying to secure your chatbot first and secure your APIs. Treat agents like untrusted third parties, enforce object-level authorization everywhere, and assume any “internal-only” endpoint is mappable once AI is involved. As Shaul notes, faster shipping via AI-assisted coding can also mean more security findings if teams don’t deliberately optimize for correctness.

    • Transcript
    • Chapters
  • #53
    August 18 · 21 min

    Epic AI fails: Why “useful” isn’t “correct”

    AI “epic fails” aren’t just funny headlines; they’re patterns you can design against. In this episode of Pop Goes the Stack, F5's Lori MacVittie, Joel Moses, and Buu Lam walk through why so many AI-powered features keep going off the rails, from chatbots inventing policies to agents deleting real infrastructure, and what those failures teach us about building safer systems. Joel frames most incidents in two buckets. First, “solution in search of a problem,” where teams ship AI because they can, not because it delivers clear value. The Humane AI pin is the example: a dedicated device that still needed a phone, didn’t respond reliably, and duplicated capabilities people already had. Second, treating a statistical prediction engine like an authority. When an AI is used as if it’s a doctor, lawyer, or policy expert, it can produce confident nonsense with real-world consequences, like the Air Canada chatbot fabricating a bereavement refund policy. Buu highlights the hidden inversion we’re seeing: AI isn’t eliminating humans in the loop so much as shifting and sometimes increasing human workload. Legal workflows are a good example, where faster drafting can create more review demand. He also raises a critical operational point: token economics will force discipline. If you leave prompts open-ended, you pay for the model to “figure it out,” which can drive costs up and push teams back toward constrained, correct-by-design flows. The practical enterprise takeaway is permissions and agency. An agent “doing the thing” is still doing it as you. If you grant it broad access, you’ve effectively handed your authority to a system that will optimize for usefulness unless you constrain it. Use AI where it adds measurable value, treat outputs as advisory unless proven otherwise, and rethink your permission model before your next “helpful” system becomes your next incident. Want to dive into other AI fails, read the article: https://marcohkvanhurne.medium.com/the-ten-biggest-ai-fails-of-2025-5d14fe876b2a

    • Transcript
    • Chapters
  • #52
    August 11 · 21 min

    Does your chatbot code? Why guardrails fail (and how to fix drift)

    Chipotle’s chatbot becoming an unofficial coding assistant wasn’t just a funny internet moment. It was a clear signal that most chatbot “guardrails” are still too shallow for systems that are optimized to be helpful, not correct, and definitely not restrained. In this episode of Pop Goes the Stack, F5's Lori MacVittie, Joel Moses, and Emmet McGinnity unpack why jailbreaks and topic drift keep happening, and what teams can do to keep chatbots focused on the job they were actually deployed to do. Emmet’s core point is that safety measures have to start with narrowing scope. A chatbot should operate like a laser pointer, not a flashlight: it should ignore 99% of what the base model can do and stay inside a tight slice of allowed behavior. That begins with a system prompt, but it can’t end there. Naive keyword and regex filtering is easy to bypass with encoding tricks and prompt manipulation, so stronger approaches include adding a verifier or judge agent that evaluates the conversation holistically to detect when it’s drifting out of bounds. They also highlight that long conversations are a common failure mode. As context grows, it becomes easier for the model to veer into capabilities it shouldn’t use, including writing code or pulling sensitive data. Practical controls include summarizing and “squashing” sessions, pruning context when drift begins, rolling back to a safe point in the conversation, or forcing a full reset when needed. A key theme is permissioning: the chatbot must honor what the user is allowed to do, not what the chatbot can access. The real measure of a safe, successful chatbot isn’t the breadth of its knowledge, it’s what it reliably chooses not to do. If you’re deploying chatbots in production, this episode is a useful blueprint for focusing scope, monitoring drift, and enforcing boundaries before someone else does it for you.

    • Transcript
    • Chapters
  • #51
    August 4 · 23 min

    The Great AI Repatriation: Why Cloud‑Only LLMs break the budget

    AI isn’t “going to the cloud” the way the headlines promised. It’s going wherever the economics and the architecture force it to go, and that often means back on hardware you control. In this episode of Pop Goes the Stack, Lori MacVittie talks with longtime cloud strategist David Linthicum about why AI workloads are driving a very familiar shift: from breathless outsourcing narratives to a sober “where does this bring the most business value” decision. David argues that the right placement question is not ideological, it’s operational. Public cloud LLMs bring ecosystem convenience, but GPU-as-a-service costs can be multiples higher than running inference on your own equipment, even after factoring in colocation, managed services, leasing, and support. That’s colliding with token shock: organizations build agentic prototypes expecting small bills and then get six-figure invoices because demand and context usage are hard to forecast. The discussion also highlights a second trap: lock-in. Using a simple API can make switching models easier, but agents often pull teams into full frameworks and ecosystems that are harder to unwind later. And the technology isn’t standing still; today’s transformer-era models aren’t the final generation, so tying your long-term processes to a single provider can turn into expensive technical debt. The practical message is blunt: stop overbuilding. Most successful AI applications in enterprises will be narrow, tactical, and “minimum viable” in their use of AI. Sometimes that’s a small model on-prem. Sometimes it’s classic ML. Sometimes it’s a frontier model in the cloud. The win is choosing the smallest effective solution, in the right location, at a cost your business can sustain. If you’re planning AI infrastructure, this episode is a reality check: best-of-breed, hybrid placement is back, and the companies that treat AI spend like a business decision will outlast the ones treating it like a hype contest.

    • Transcript
    • Chapters
  • #50
    July 28 · 21 min

    Training vs Inference: Are they the same?

    Training and inference get lumped together in casual AI conversations, but they behave differently enough that the distinction matters for cost, architecture, and security. In this episode of Pop Goes the Stack, Lori MacVittie, Joel Moses, Ken Arora, and Kevin Baughman (who leads F5’s AI Center of Excellence) unpack what’s truly different, what’s the same, and where people get misled. Joel makes the “math is the same” case: both phases run similar computations, but training must retain intermediate activations for backpropagation, while inference can discard them. Ken and Kevin pull the conversation back to practical differences: training is about baking knowledge into the model, while inference is about using a frozen model and shaping behavior with context, retrieval, and few-shot examples. The weights don’t change during inference; the input does, which is why it can feel like “learning” without actually being permanent. That distinction becomes a security and governance lever. If you don’t want sensitive or proprietary data baked into a model, you avoid training on it and instead keep it in a controlled knowledge base (RAG or similar) that can be updated, removed, or scoped per tenant. Meanwhile, training pipelines emphasize massive data ingestion and throughput, and inference emphasizes responsiveness, session context, and efficient serving at scale. The practical takeaway is to stop treating “AI workloads” as one thing. Training and inference require different pipeline designs, different tradeoffs in memory and bandwidth, and different approaches to data control. Pick your phase, understand the constraints, and build for it intentionally.

    • Transcript
    • Chapters
  • #49
    July 21 · 21 min

    Round Robin is Still Dumb: Load balancing has to grow up for AI inference

    Round robin isn’t just “not ideal” for LLM inference. According to recent scheduling research, it’s actively harmful, because it treats inference like stateless, interchangeable API traffic when it’s anything but. In this episode of Pop Goes the Stack, Lori MacVittie is joined by F5's Josh Mendoza, Principal Solutions Engineer, to break down why classic load-balancing assumptions fail under LLM workloads, and what to think about instead. Josh walks through the evolution from early “spray and pray” distribution to smarter approaches that account for server load, workload type, and state. That history matters because AI introduces the same challenge at a new intensity: inference is a heavy compute-and-memory math pipeline, and conversations accumulate state. Once context and KV cache are involved, moving a request to a different server isn’t a clean failover, it’s a forced cache miss and a recomputation penalty that shows up as slower time-to-first-token and higher cost. They connect this back to patterns teams already understand: VM migration, session persistence, and why “just move it” has always been expensive when the working set is large. LLMs raise the bar because users won’t tolerate latency, and the payload you’d need to move grows as the interaction continues. You also can’t ignore the request itself, since “summarize this” and “write a full analysis” have very different compute profiles, even if they hit the same endpoint. The main takeaway is simple: don’t panic, but stop treating inference like generic API traffic. Effective LLM scheduling is closer to dispatching the right resources to the right job, with awareness of state, model placement, cache locality, and the true cost of moving work. The tools exist, but the mental model has to change first.

    • Transcript
    • Chapters
  • #48
    July 14 · 23 min

    Mechanistic Interpretability: Debugging LLMs by reading their circuits

    Mechanistic interpretability sounds academic until you try to debug a model with printf and realize there’s nothing to print. In this episode of Pop Goes the Stack, Lori MacVittie is joined by F5 Chief Product Officer, Kunal Anand, to talk about why “mechinterp” is getting serious attention: if we can’t understand how models arrive at decisions, we can’t predict failure modes or build effective controls. Kunal walks through his deep dive, sparked by a conversation about how much weight individual tokens can carry, especially as context windows grow and models don’t always “use” every part of their capability to produce a plausible response. That rabbit hole led to his blog post, “Your Token is a Wonderland,” where he trained a transformer on his own iMessage history to build a model on a dataset he understood intimately. The point wasn’t novelty; it was debug-ability. With a smaller model, he could inspect attention patterns, layer behavior, and token predictions in a way that’s effectively impossible on trillion-parameter frontier systems. They discuss what this kind of work reveals: how context changes meaning, why certain tokens get selected, and why model behavior can feel opaque even when outputs look confident. The conversation also ties mechinterp back to practical outcomes, from improving guardrails and refusal behavior to finding ways to reduce hallucinations and avoid high-stakes errors without retraining entire models. The takeaway is pragmatic: we’re early, and the field is still nascent, but it matters. Understanding internal “circuits” isn’t just intellectual curiosity; it’s a path toward better debugging, safer behavior, and more reliable AI systems. Until then, variability is part of the deal, and “the model said so” still isn’t an explanation. Read Kunal's blog, Your Token is a Wonderland for his mechanistic interpretability deep dive: https://kunalanand.com/2026-03-19-your-token-is-a-wonderland/

    • Transcript
    • Chapters
  • July 7 · 46 sec

    Pop Goes the Stack will be back next week!

    Hi everyone! This is Lori MacVittie, host of Pop Goes the Stack. This week, we’re taking a short break to rest and recharge, spend time with loved ones, and maybe even step away from our stacks for a bit (gasp!). But don’t worry—we’ll be back soon with more: ✅ Sharp insights into emerging tech ✅ Expert takes on application delivery & security ✅ And, of course, our signature snark While we’re on hiatus, check out some of our past episodes diving into AI and agents, hardware advancements, and other tech trends that are impacting your application stack. Thank you for tuning in and being part of the Pop Goes the Stack community. We'll see you next week!

    • Transcript
  • #47
    June 30 · 23 min

    Agent Skills: The new AI supply chain risk (and fixes)

    Agent skills were introduced less than six months ago, and they’ve already graduated from “handy configuration” to “supply chain artifact.” In this episode of Pop Goes the Stack, Lori MacVittie talks with security expert Peter Scheffler about why skills, often packaged as YAML, are becoming portable, shareable, and dynamically loadable in ways that attract attackers fast, including skill poisoning and repository-based compromise. The fact that there’s already an OWASP Top 10 focused specifically on agentic skills tells you how quickly this risk surface is forming. Peter breaks down what “skills” really are: anything from a narrow tool instruction to a broad workflow like “prepare for a podcast.” Skills can be created by humans, generated by agents, and even expanded by tools that add more skills, which creates a compounding trust problem. Once skills can be modified, composed, and distributed, you need provenance, signatures, hashing, and an approval process, but simply copying the traditional package ecosystem isn’t a silver bullet because supply chain compromise is already a reality. The conversation pivots to what actually helps: least agency. Define what actions an agent is allowed to take, and constrain execution at multiple layers, not just in a system prompt. System prompts are guidance, not enforcement, and relying on them alone is asking to get burned. Then assume unintended action, treat all external content as untrusted input, and focus on stopping unsafe actions at the boundary rather than trying to prevent the agent from ever attempting them. Finally, Peter stresses observability. If agents can make their own calls, you must log agent-to-agent interactions, tool usage, and skill loading, because you’ll need forensic data when something goes wrong. For enterprises, the practical starting point is clear: follow emerging frameworks (OWASP, NIST), standardize which agents and skills are allowed, store approved skills in a controlled registry, enforce authentication and authorization to access them, and be ready to collect a lot more telemetry than you’re used to.

    • Transcript
    • Chapters
  • #46
    June 23 · 20 min

    Is AI making security obsolete?

    What happens if AI finally writes secure code by default? In this episode of Pop Goes the Stack, F5's Lori MacVittie, Joel Moses, and Ken Arora take that question seriously, even if it feels like a punchline today. The premise is simple: if AI starts producing mechanically correct, vulnerability-free code at scale, the security industry doesn’t disappear. It gets forced up the stack. They outline the near-term reality first: AI can already find and fix large classes of common issues in legacy code, and we’re likely heading into a chaotic cleanup phase as tools like Mythos-style systems accelerate remediation. But longer term, the conversation shifts to the uncomfortable tradeoff: code quality may improve faster than complexity shrinks. And complexity is where authority problems, logic flaws, and operational failures thrive. Ken uses a practical analogy: a “perfectly safe driver” can still follow a GPS off a cliff. Likewise, perfect code can implement flawed or exploitable business logic flawlessly. Ticket scalping, abuse of workflows, and manipulation of incentives aren’t bugs; they’re behavior and design gaps. The bigger risk becomes how systems are orchestrated, who has permission to do what, and how autonomous components interpret intent. They also flag the adversarial side: attackers won’t stop; they’ll shift from injecting obvious vulnerabilities to subtly influencing specs, workflows, and machine-generated architectures. As specs in plain language become the new “source code,” more people will be able to build powerful systems without the instinct to think like an attacker, widening the logic-attack surface. The takeaway is a shift in mindset: security becomes less about chasing broken code and more about governing boundaries and monitoring behavior. AI will increasingly do exactly what you ask, which makes the new security imperative painfully clear: be precise about what you ask it to do, and enforce what it’s allowed to do when it tries to go beyond that.

    • Transcript
    • Chapters
  • #45
    June 16 · 23 min

    Agents deleted my work: Why agents still aren’t ready for production (Yet)

    An agent deleting a production database (and the backups) isn’t a sci-fi failure. It’s a boundary failure, and it starts with a human handing out credentials and permissions without a safe execution model to contain what happens next. In this episode of Pop Goes the Stack, Lori MacVittie and F5's Chief Product Officer, Kunal Anand, unpack why today’s agents are either dangerously overpowered or so constrained they’re barely useful, and what needs to change to make them viable. They dig into the current reality of “agent” features in mainstream tools, especially how Copilot-style agents often feel like chatbots trapped behind walls: limited access, weak integration, and poor continuity when context windows overflow. Kunal shares two painful examples: voice-mode work that produced the right output but didn’t persist a transcript or draft, and an inbox assistant that can’t actually read the inbox without copy-paste, making it useless for real workflow automation. The core point is that system prompts aren’t constraints, they’re guidance, and guidance fails the moment a goal-driven system tries to “do the thing” by any means necessary. That’s why Microsoft’s move to build agent permission primitives directly into Windows is a meaningful shift: controls need to be enforced at the OS and runtime level, not politely suggested to the model. They also touch on practical workarounds, like exporting a long chat as a PDF to carry context forward, and why isolation and blast-radius reduction are still table stakes. The takeaway is straightforward: agents in production are still the exception, not the norm. Most enterprises are deploying AI-enabled applications first, while keeping agentic automation largely in employee workflows. Until we get real, enforceable boundaries and better UX for authority and approval, treating agents as production-grade operators is a risk most teams can’t justify.

    • Transcript
    • Chapters
  • #44
    June 9 · 23 min

    Agent identity: Closing the "accountability vacuum" with humans

    Identity used to be straightforward: authenticate a user, authorize an action, log the request, and move on. Agentic systems complicate that model because the actor isn’t always the human anymore, and when something goes wrong, responsibility can disappear into what Andrew Bud calls an “accountability vacuum.” In this episode of Pop Goes the Stack, Lori MacVittie talks with Andrew Bud of iProov about why this isn’t just a security nuance, but a broader stability problem. You can’t punish, retrain, or sue an agent. Yet agents can still take actions with real consequences, from leaking code to corrupting data to making irreversible operational changes. If accountability can’t attach to the agent, it has to attach somewhere else. Andrew’s argument is that responsibility shifts to the relying party. Service providers and systems need to ask whether they’re dealing with a human or an agent, identify who the agent belongs to, and gate high-impact actions so a real human can be held accountable. That implies a chain of delegation and auditability that looks more like certificates, but with a different root of trust: proof of genuine human presence. The conversation distinguishes enterprise agents, where existing identity patterns like OIDC and governance tools may still work, from “agents in the wild,” where centralized identity breaks down and decentralized identity becomes more relevant. Andrew points to emerging standards work across multiple groups and makes the case that verified human presence, not just identity facts, will become foundational as agents increasingly claim to be people. If you’re deploying agents, the takeaway is clear: identity alone isn’t enough. You need provable human roots of trust, stronger relying-party controls, and policies that treat some actions as requiring explicit human accountability.

    • Transcript
    • Chapters
  • #43
    June 2 · 23 min

    Data poisoning: You can’t patch what an LLM “learns”

    If you’ve been treating “garbage in, garbage out” as a metaphor, this episode turns it into a live-fire scenario. Lori MacVittie and Joel Moses are joined by Dmitry Kit to unpack what happens when AI systems ingest misinformation that looks legitimate, and why “just patch it” doesn’t work the way it does in traditional software. They start with a real experiment: researchers fabricated a fake medical condition, complete with fake papers, authors, and supporting citations, and watched it propagate. Within weeks, major AI systems began surfacing and citing it as real. The uncomfortable point is that once false knowledge gets embedded, you can’t reliably roll it back. Retraining is expensive, fine-tuning doesn’t truly excise the information, and even “fixes” can create unintended side effects because the bad pattern can be distributed throughout the network. The conversation reframes the core issue as trust and weighting. Models don’t learn from “the internet” evenly; they learn from sources that are implicitly ranked as more authoritative, which means poisoning a trusted channel can have outsized impact. Even without a trusted source, rare or highly specific topics are vulnerable because the model has so little competing context that a small amount of misinformation can dominate. So what can teams do? The practical guidance is to reduce the attack surface by curating the data set and narrowing scope. For enterprise use cases, that means constraining responses to approved, maintained knowledge, applying strong governance to RAG sources, and using additional validation layers, including “LLMs as judges,” to screen what gets added. The takeaway is simple: you can’t rely on cleanup after contamination. Prevention, curation, and constraint are the only scalable strategies. Read the Bixonimania article: https://www.nature.com/articles/d41586-026-01100-y

    • Transcript
    • Chapters
  • #42
    May 26 · 21 min

    Local-first AI: Keep context out of the cloud

    “Just throw it in the cloud” gets complicated when the data is your meetings, your IP, and your operating context. In this episode of Pop Goes the Stack, Lori MacVittie and Joel Moses talk with Michael Daugherty, founder and CEO of Quill Meetings, about why local-first AI is showing up as a serious alternative to cloud-first convenience, especially when your AI is effectively a coworker sitting in every meeting. Local-first tools keep transcription, notes, highlights, and long-term context on your device or inside your org, so your most valuable (and most sensitive) inputs don’t default to third-party APIs. The payoff: Better personalization from the context that only exists locally Stronger privacy & compliance for regulated teams and sensitive conversations Clear control over the “data tap”—share with other AI tools only when you choose Reusable meeting knowledge: build a personal/organizational lexicon you actually own Enterprise-friendly paths like private inference servers and VPN-controlled architectures They also dig into practical realities—hardware variability, GPU/driver quirks, and resilient fallbacks—plus how Quill uses MCP (server + client) to let you bring your meeting corpus into tools like Claude and Cursor while keeping control where it belongs. Bottom line: context is becoming the competitive advantage in AI, and where that context lives matters. Local-first tools give teams a way to set boundaries, reduce exposure, and still benefit from AI, without assuming the cloud is the only place intelligence can run.

    • Transcript
    • Chapters
Showing 1–20 of 20 episodes