Skip to content
Artwork for EDGE AI POD

EDGE AI POD

EDGE AI FOUNDATION

Discover the cutting-edge world of energy-efficient machine learning, edge AI, hardware accelerators, software algorithms, and real-world use cases with this podcast feed from all things in the world's largest EDGE AI community. 

These are shows like EDGE AI Talks, EDGE AI Blueprints as well as EDGE AI FOUNDATION event talks on a range of research, product and business topics.

Join us to stay informed and inspired!

Play
  • 20 episodes
  • weekly
  • Avg 21 min
  • English

Support the show

Goes straight to the publisher. podnod takes nothing.

  • Thursday · 50 min

    Bridging the Research-Reality Divide in Edge AI

    Ever wondered why so many groundbreaking AI innovations never make it to market? The answer lies in the treacherous gap between research and reality – a challenge that's costing companies millions and delaying critical technologies from reaching consumers. This riveting panel discussion brings together seasoned experts from Intel, Wind River, Advantech, EmbedUR, and The Things Industries who've accumulated plenty of "scar tissue" trying to bridge this divide. Their conversation cuts through the hype to reveal the practical obstacles that prevent brilliant AI concepts from becoming commercial products. The panelists don't hold back as they address the hard truths: safety certification requirements that can derail deployment in mission-critical industries, the dangers of incorporating AI technology without clear use cases, and the lack of standardization that forces developers to reinvent the wheel with each implementation. One panelist shares how aerospace customers peppered them with certification and explainability questions for 45 minutes when presented with new edge AI capabilities – revealing how regulatory requirements can completely block adoption in certain sectors. You'll gain invaluable insights into the four pillars needed for successful edge AI deployment: standardization, traceability, explainability, and certification. The discussion also explores the surprising disconnect between technology maturity and business processes, revealing why even the simplest IoT implementations fail when organizations aren't digitally ready. Whether you're a researcher, developer, product manager, or business leader, this conversation provides the roadmap for turning your AI innovations into market-ready solutions. Because as one panelist bluntly puts it, "At the end of the day, the KPI is cash." Subscribe now to hear the strategies that can help your next AI project cross the finish line from laboratory to real-world deployment. Send us Fan Mail Support the show Learn more about the EDGE AI FOUNDATION - edgeaifoundation.org

    • Transcript
    • Chapters
  • August 20 · 15 min

    Synthetic Data For Smarter Vision

    You shouldn’t need a warehouse full of broken products to build a reliable vision model. We dig into how synthetic data flips the script for manufacturing and quality control, showing why controlled scenes, perfect labels, and rapid iteration can outrun old pipelines of manual collection and error-prone annotation. With Sherry List and Goran from Synthetic AI Data, we walk through the strategy behind a no-code engine that empowers fusion teams—engineers, developers, and operators—to generate the exact scenarios their models need. The story centers on a high-stakes, high-volume domain: metal and aluminum cans. From standard to sleek, slim to stubby, tabs in different materials and colors, and both beverage and food lids, we map the defect landscape that actually matters on the line. Bent, broken, missing, and lifted tabs are recreated with photoreal materials, realistic reflections, and varied lighting, then captured from top-down and side angles to match real inspection setups. The result is control—over class balance, severity, camera pose, and environment—so teams can stress-test models and discover what improves accuracy before hitting production. We also reveal the scale required to make a difference: roughly three million synthetic images, high and low resolution, fully annotated in COCO, YOLO, TensorFlow, and more. Releasing the dataset under CC0 via the EDGE AI Labs removes friction for researchers and practitioners to explore defect detection, domain randomization, and multi-view training. Along the way, we share a cautionary tale of “screwdriver datasets” and explain why simulation delivers safer, faster, and more reproducible results. If you care about computer vision performance, cost, and time-to-value, this conversation offers a practical blueprint you can use today. Subscribe for more deep dives, share this with a teammate who wrangles labels, and leave a review telling us what you’ll build with the dataset. Send us Fan Mail Support the show Learn more about the EDGE AI FOUNDATION - edgeaifoundation.org

    • Transcript
    • Chapters
  • August 13 · 15 min

    Cloud Can’t Keep Up, So Your Toaster Gets A Brain

    The cloud can’t carry the weight of billions of sensors forever, and we’re proving why. We walk through a new class of ultra‑low‑power, heterogeneous neuromorphic microcontrollers that bring real intelligence to the edge, where timing, latency, and privacy matter most. From raw IMU streams to on‑device actions, you’ll hear how spiking neural networks, tiny CNNs, and a RISC‑V core team up to decode the world in real time without draining a battery. We dig into the full signal path: encoding analog magnitude and velocity into spikes, pushing temporal patterns through an SNN accelerator, and decoding results for decisions on the spot. Our Talamo SDK lets you train in a familiar PyTorch‑like workflow, visualize progress with TensorBoard, and then hand everything to a system compiler that maps your pipeline across hardware and software, generating deployable binaries. No guesswork, no fragile glue code. To keep iteration fast, our cycle‑approximate SoC simulator mirrors the chip’s timing behavior so closely that functional results match hardware one‑to‑one, enabling confident development even before silicon lands on your desk. We also showcase a live wearable gesture demo built on accelerometer and gyroscope data, using integrate‑and‑fire and temporal‑contrast encoders to capture amplitude and motion dynamics. You’ll get candid results: auto‑generated code trades a bit of size and power for big gains in developer speed, while the simulator runs near real time. To cap it off, we announce a commercial, award‑winning ultra‑low‑power neuromorphic chip designed for consumer electronics, smart home, industrial monitoring, and wearables. Ready to build products that sense, understand, and act at the source? Follow, share, and leave a review to tell us what you want to create at the edge. Send us Fan Mail Support the show Learn more about the EDGE AI FOUNDATION - edgeaifoundation.org

    • Transcript
    • Chapters
  • August 6 · 17 min

    Edge AI That Cuts Chemical Waste

    What if a lab test that takes 12 to 24 hours could be replaced by a live estimate that guides dosing in real time? We walk through a high-stakes water story where boron control in desalination demanded more than a clever model—it needed a secure, local-first AI system that works across wildly different plants. Our journey with Acciona started with a simple idea: a virtual sensor to predict boron and avoid overusing caustic soda or risking fines. The reality was messy. Membranes, sensors, and SCADA setups varied from site to site. Cybersecurity kept data locked on-prem, and lab workflows produced sparse, noisy labels. A single global model wasn’t resilient enough. So we flipped the playbook and orchestrated many models at the edge—one per rack when needed—packaged in Docker, deployed with a click, and monitored locally with InfluxDB and Grafana. We break down the full stack: MQTT brokers to standardize data, connectors for heterogeneous OT systems, TensorFlow for inference, and JupyterLab plus MLflow for on-device training and versioning. This architecture kept raw data inside the plant while a cloud console managed applications securely. The payoff was immediate: accurate boron estimates tightened dosing, cut chemical spend, reduced penalties, and built operator trust by showing predictions alongside lab results. One site saved over $200,000 in a year; scaled across the fleet, the impact reaches well into the millions, with healthier water as a bonus. Beyond boron, the same edge AI approach unlocks energy optimization for high-pressure pumps, membrane fouling detection, and even computer vision tasks—without compromising critical infrastructure security. If you care about industrial AI that actually ships, this is a practical blueprint: local models, secure orchestration, and a path from pilot to fleet. Enjoy the episode, share it with a teammate who wrestles with on-prem constraints, and leave a review. Want to see it live? Ask us for the free trial and we’ll set up a demo. Send us Fan Mail Support the show Learn more about the EDGE AI FOUNDATION - edgeaifoundation.org

    • Transcript
    • Chapters
  • July 30 · 42 min

    The Skinny Transformer: Squeezing Gen AI into Tiny Devices

    The future of artificial intelligence isn't just in massive cloud data centers—it's happening right now on the devices all around us. This insightful panel discussion brings together leading experts from major semiconductor companies and academia to explore how Generative AI is transforming edge computing. What makes this conversation particularly valuable is the panelists' emphasis on practical reality versus future potential. While many assume GenAI requires enormous computing resources, the experts reveal that today's edge hardware—from smartphones to IoT devices—already supports numerous generative applications. The key isn't waiting for more powerful chips but rethinking how we approach model design, quantization, and specialization. Danilo Pau from STMicroelectronics shares a fascinating vision of natural language interaction with everyday objects like thermostats, while Qualcomm's Evgeny Kuznetsov highlights how real-time translation and synthetic data generation deliver immediate productivity benefits. ARM's John Mark Yodis emphasizes that education and framework selection are more significant barriers than hardware limitations. The technical discussion delves into cutting-edge compression techniques, with quantization advancing from Int8 to Int4, Int2, and even Int1 representations. Professor Huanrui Yang explains how foundation models can be specialized and pruned to maintain performance only in domains relevant to specific edge applications. This targeted approach enables capabilities previously thought impossible on resource-constrained devices. Perhaps most exciting is the panel's exploration of unique edge advantages—proximity to data, sensor integration, and specialized hardware—that enable entirely new GenAI applications impossible in the cloud. Through orchestration across heterogeneous computing resources and domain-specific adaptation, the next wave of intelligent systems will distribute AI processing across the compute spectrum. Whether you're a developer looking to deploy GenAI on current hardware, a researcher exploring new compression techniques, or a product manager planning your AI roadmap, this discussion provides crucial insights into what's possible today and where the technology is heading tomorrow. Don't wait for the future—generative AI at the edge is already here. Send us Fan Mail Support the show Learn more about the EDGE AI FOUNDATION - edgeaifoundation.org

    • Transcript
    • Chapters
  • July 23 · 12 min

    How Microsecond AI Control Transforms Power Systems And Cuts Errors

    What if the control loop could think ahead and correct itself before errors take hold? We dive into a practical leap for motors, inverters, and energy storage: ultra-low-latency edge AI that predicts error trajectories at startup and intervenes inside the loop in about 100 microseconds. Instead of piling on sensors and pushing raw signals to the cloud, we work directly from existing operational data, chart the most efficient path, and act locally—then pass only meaningful transients upstream for fleet analytics and predictive maintenance. We start by grounding the challenge: linear systems tolerate classic PID, but nonlinear dynamics create overshoot, oscillation, and costly performance tradeoffs. Throwing bigger processors at the problem hits limits on cost, memory, and thermals. The solution mirrors a lesson from the smartphone era—where dynamic voltage and frequency scaling transformed performance-per-watt—by bringing adaptive optimization to the plant itself. Our Ultra-Edge technology extends PID behavior into nonlinear territory, shrinking speed error during torque steps and tightening control, even on modest 32 MHz platforms, with further gains as faster silicon comes online. From factory floors to the power grid, the implications are big. In motor drives, torque transitions smooth out with fewer current spikes. In utilities and data centers, grid-forming converters coordinate with renewables and battery energy storage to deliver synthetic inertia, riding through disturbances and supporting stability rather than tripping offline. By acting in microseconds, converters offer a stabilizing boost, enabling higher renewable penetration and a more credible path to net zero. Meanwhile, microcontroller-level filtering trims a million samples per second down to high-value events so teams get signal without noise. If you care about real-time control, nonlinear systems, and scaling stability with clean energy, this conversation brings clear examples, measured results, and a roadmap for adoption—from pilots and soft IP to demo platforms and a growing model library. Subscribe, share with a teammate who owns drives or converters, and leave a review with your biggest control pain point so we can tackle it next. Send us Fan Mail Support the show Learn more about the EDGE AI FOUNDATION - edgeaifoundation.org

    • Transcript
    • Chapters
  • July 16 · 28 min

    Edge of Tomorrow: How NXP is Revolutionizing On-Device AI

    The AI landscape is transforming rapidly, and NXP Semiconductors is at the forefront of bringing these capabilities where they matter most—directly to edge devices. Alberto Alvarez delivers a compelling overview of how NXP is enabling sophisticated generative AI to run locally on microprocessors, without relying on cloud connectivity. Unlike companies focused on massive cloud-based AI training, NXP targets the critical deployment phase, where privacy, security, and efficiency are paramount. Their approach empowers developers to create AI-enhanced solutions for industrial automation, healthcare, automotive systems, and smart environments that keep sensitive data completely local. The presentation unveils the EAQ GenAI flow—a comprehensive software pipeline that allows developers to fine-tune and optimize large language models for specific applications without exposing proprietary data to third-party servers. This pipeline includes automatic speech recognition (ASR) based on the Whisper architecture, LLM reasoning with LLAMA3, retrieval-augmented generation (RAG) for domain-specific knowledge, and natural text-to-speech synthesis—all running efficiently on NXP's hardware. Most impressively, through a partnership with Kinara, NXP demonstrates a fully edge-based multimodal AI implementation running on their iMX810 Plus platform. This system combines an 8-billion parameter language model with computer vision capabilities, allowing it to analyze images, reason about visual content, and respond to questions—all without sending any data to the cloud. The implementation achieves remarkable performance metrics, generating 6.5 tokens per second with response latency as low as 1.5 seconds for follow-up questions about images. From robots with enhanced reasoning capabilities to medical assistants that can analyze diagnostic imagery, the possibilities for this technology are vast and expanding daily. As NXP continues pushing the boundaries of what's possible at the edge, they're laying the groundwork for the next frontier: agentic AI systems that can perceive, reason, and act autonomously across multiple modalities. Ready to build secure, private AI applications that don't compromise on capability? Explore NXP's resources and start creating tomorrow's intelligent edge solutions today. Send us Fan Mail Support the show Learn more about the EDGE AI FOUNDATION - edgeaifoundation.org

    • Transcript
    • Chapters
  • July 9 · 13 min

    Hardware-Aware AI, Not Just Bigger Models

    What if the obstacle to fast, reliable AI isn’t your dataset or your optimizer—but the silicon under your model? We dig into why performance collapses when architecture and hardware don’t align, and we lay out a clear path to ship models that actually fly on the devices your users own. Starting with the Ferrari-and-hummingbird metaphor, we show how theoretical efficiency—FLOPs, parameters, even TOPS—often fails to predict real-world latency, power, and user experience. We walk through a surprising benchmark: MobileNet V2, small and “efficient,” runs slower than an older ResNet18 on GPUs because depthwise, sequential kernels underutilize parallel hardware. Then we zoom out to hardware selection itself, where NPUs can outperform GPUs despite lower TOPS due to operator support, kernel fusion, and memory behavior. The takeaway is simple: architecture matters only in context, and context means the execution engine, compiler stack, and memory hierarchy that will carry your model in production. From there, we share a four-step framework to become hardware aware: profile on real devices from day one, verify operator compatibility early, automate bottleneck discovery and model selection in CI, and optimize with context using hardware-aware pruning and mixed precision. To show how this works in practice, we unpack our Llama 3.2-1B project on Snapdragon Gen 3, where targeted pruning and precision tuning delivered 31% faster token generation, 25% faster prompt processing, and a 126% faster initialization—all with under 1% accuracy loss. If you build models for the edge, mobile, GPUs, or NPUs, this conversation will help you avoid dead-ends and design for the hardware you actually ship on. Subscribe for more deep dives, share this episode with your team, and leave a review to tell us which hardware you’re targeting next. Send us Fan Mail Support the show Learn more about the EDGE AI FOUNDATION - edgeaifoundation.org

    • Transcript
    • Chapters
  • July 2 · 15 min

    What If A Pair Of Glasses Could Read Intent?

    Imagine steering a game with nothing but a blink and a glance. That’s the spark behind our latest build: a noninvasive brain-computer interface that runs entirely on a tiny edge microcontroller, translating eye movements into reliable, real-time commands without a laptop or cloud. We start with the human why. Millions live with neurological conditions that constrain movement but preserve eye control—a narrow channel with huge potential. We compare the promises and trade-offs of invasive BCIs like Neuralink, BrainGate, and Synchron against accessible wearables from Emotiv, Muse, and OpenBCI. The big gap is obvious: people need precise, low-latency control without surgery, high cost, or a desktop tether. Our approach uses electrostatic charge sensing with a glasses-ready electrode layout at the nose bridge and a reference behind the ear, capturing strong ocular signals that are practical for daily wear. From there, we break down the full on-device pipeline. A high-pass filter removes drift, a 50 Hz notch kills power-line noise, and a low-pass smooths the signal so a smaller model can focus on meaningful features. A lightweight Z-score event detector stays always-on and wakes the classifier only when something happens, buffering a 300-sample window at 240 Hz across two channels. The classifier is a tiny 1D CNN—convolution, ReLU, pooling, softmax—clocking about 0.76 ms inference with roughly 18 KB flash and 6 KB RAM. With K-fold cross-validation on nine participants, we see around 90% accuracy for four classes: discard involuntary blinks, map voluntary blinks to “click,” and detect left and right glances. We showcase it with a playful demo: blink to jump over obstacles, glance right to change lanes and collect coins. Beyond the fun, the implications are serious—restoring agency with affordable hardware that works off-grid in real time. We close by outlining what’s next: integrating the sensors into everyday glasses, testing across more users and environments, and adding quick calibration for personalization. If accessible control matters to you—whether for assistive tech, gaming, or new hands-free interfaces—this is a glimpse of what near-future wearables can do. Enjoy the episode? Follow the show, share it with a friend, and leave a quick review to help more listeners discover these conversations. Send us Fan Mail Support the show Learn more about the EDGE AI FOUNDATION - edgeaifoundation.org

    • Transcript
    • Chapters
  • June 25 · 9 min

    Got Fake Chips? Our AI Doesn't Fall For That

    Semiconductor counterfeiting has grown into a $200 billion annual problem threatening the integrity of global electronics supply chains. As both chip shortages and sophisticated counterfeiting techniques persist, traditional detection methods fall short—requiring complex setups, hardware modifications, or extensive data labeling. Two machine learning engineers from Analog Devices' advanced R&D team unveil their elegant solution: an unsupervised learning approach that captures the unique "fingerprints" of authentic chips by analyzing power signatures during memory operations. What makes their method revolutionary is its lightweight footprint (under 60KB) and ability to run directly on standard Cortex-M4 microcontrollers at the edge, requiring no cloud connectivity or specialized equipment. The team shares their methodology for creating a robust dataset of 1,000 secure authenticator chips and developing a convolutional autoencoder architecture that achieved 100% accuracy in distinguishing authentic components from close counterparts. Their model learns the normal reconstruction patterns of legitimate chips, then flags anomalies when encountering counterfeits with distinctly different power signatures. Beyond secure authenticators, this approach proves universally applicable to any semiconductor from which analog fingerprints can be collected. Rather than replacing traditional cryptographic methods, it serves as an additional security layer that remains effective even when encryption keys might be compromised through side-channel attacks. Ready to strengthen your supply chain against increasingly sophisticated counterfeits? Discover how this scalable, software-based solution could be integrated with your existing security infrastructure to provide an additional layer of protection for critical semiconductor components. Send us Fan Mail Support the show Learn more about the EDGE AI FOUNDATION - edgeaifoundation.org

    • Transcript
    • Chapters
  • June 18 · 11 min

    Smarter AI, Faster Hardware

    Your phone, watch, and even your fridge want real-time intelligence—but power and latency won’t tolerate bloated models or generic compute. We walk through a practical path from Python to custom hardware using high-level synthesis, then invite you to prove it in our Efficient Inferencing Hackathon. With a ready-to-run RISC‑V Rocket Core baseline for MNIST, a full Siemens EDA toolchain, and on-demand training, you’ll learn how to cut latency and power while protecting accuracy through precision mapping, parallelism, and smarter dataflow. We start by mapping the compute landscape—CPUs for flexibility, GPUs for throughput, TPUs/NPUs for tensors, and custom FPGA/ASIC designs for peak power-performance-area. From there, we get tactical: use quantization to right-size bit-widths; apply loop pipelining and unrolling to unlock throughput; partition memories and stream between layers to eliminate round-trips; and iterate quickly with HLS directives instead of rewriting RTL. You’ll see how a baseline inference in the millisecond range can be driven far lower with disciplined co-design, and how Catapult HLS, Questa, and PowerPro provide the feedback loop—latency, area, and power—to make confident trade-offs. Participants receive a virtual machine, C kernels for convolution and dense layers, and a step-by-step path from Keras to synthesizable RTL. The goal is simple and demanding: deliver the fastest MNIST implementation that meets accuracy, area, and energy targets. Along the way, the HLS Academy community offers guidance from experts and peers, and winners will be announced at the Edge AI Foundation event in Taipei, with prizes including a 3D printer, an FPGA board, and Bose earbuds. Ready to turn models into efficient silicon? Join the workshop series, claim your VM via the QR code at hls.academy, and use the promo code with two underscores to unlock full access. If this resonates, subscribe, share with a teammate who ships edge AI, and leave a review to help others find the show. Send us Fan Mail Support the show Learn more about the EDGE AI FOUNDATION - edgeaifoundation.org

    • Transcript
    • Chapters
  • June 11 · 1 hr 1 min

    Village OS: AI For Sustainable Living

    What if a neighborhood could think, heal, and feed itself? We sit down with James Ehrlich of Stanford to unpack Village OS, a generative AI platform that designs resilient communities by starting with a simple question: what does the land want? From the urban edge of Riyadh to peri-urban sites worldwide, James shows how geospatial data, climate histories, hydrology, and cultural patterns come together to shape housing, farms, energy, and mobility as one living system. We trace James’s path from early game design and digital effects into the world of eco-villages and permaculture, where taste, health, and connection inspired a research agenda: use technology to serve nature and people. The demo moves from contour maps and fluid dynamics to soil restoration, aquaponics, and agrovoltaics that grow shade crops under solar. Real-time modeling toggles apartments, townhomes, and single-family mixes while projecting costs, returns, and service loads for water, energy, and waste. The punchline is elegant: at the neighborhood scale, waste becomes an asset, powering heat, cooling, and purification while closing loops for food and energy security. Funding and measurement get equal attention. Village OS projects ESG and SDG outcomes and carbon sequestration across decades, offering a transparent view for sovereign wealth funds, pensions, and institutional capital. After groundbreak, the operating layer shifts to edge AI: tinyML sensors and small language models form a digital mycelial network with low latency, low energy, and high autonomy, connected by a thin, privacy-safe cloud channel for cross-site learning. It’s resilience defined by human well-being—lower stress, safer streets, access to fresh food, and spaces for elders and children—backed by systems that can ride out disruption. If you care about sustainable housing, regenerative agriculture, microgrids, and the future of edge AI, this conversation offers a practical, hopeful blueprint. Subscribe, share with a friend who’s into systems thinking, and leave a review with the one feature you’d want in your ideal resilient neighborhood. Send us Fan Mail Support the show Learn more about the EDGE AI FOUNDATION - edgeaifoundation.org

    • Transcript
    • Chapters
  • June 4 · 14 min

    When Edge AI Meets Hearing Loss, Access Gets Real

    Crowded cafés, clinking plates, and echoey halls make conversations exhausting. We set out to change that by fitting real deep learning into an ear-sized device and proving it can separate speech from noise with almost no delay or battery hit. The result isn’t louder sound; it’s clearer lives and less fatigue. We walk through the full Clara enhancement path: transforming raw mic input into log-mel features, stabilizing for gain shifts, and feeding a 40-layer temporal convolutional recurrent network that predicts a mask to preserve voice and suppress noise. Then we show how a light touch of the original signal brings back space and warmth, avoiding the hollow, underwater audio that turns people off. Along the way, we tackle painful transients—the cutlery and clatter that spike hearing aids—and explain how wide dynamic range compression keeps everything comfortable and intelligible. The heart of the story is edge AI done right. Our SPU001 chip uses unstructured sparsity to skip zero multiplies in hardware, shrinking memory needs and power draw by orders of magnitude. That lets a pruned model with effective 10 MB scale run from just one MB of SRAM while holding algorithmic latency near eight milliseconds and total path time under ten. Metrics back it up: higher scale-invariant signal-to-distortion ratios, better hearing aid speech quality scores, and strong user reports. A rapid partnership with New Sound brought this to market in about three months, and audiologists on a noisy show floor heard the difference immediately. If you care about hearing tech, edge computing, or just making conversations effortless again, this one is for you. Hear how small silicon and smart modeling turn “AI” from a buzzword into a daily benefit. Subscribe for more deep dives on practical edge AI, share with someone who struggles in noisy rooms, and leave a review with your toughest audio environment—we might feature it next. Send us Fan Mail Support the show Learn more about the EDGE AI FOUNDATION - edgeaifoundation.org

    • Transcript
    • Chapters
  • May 28 · 22 min

    Cows Chewed Our Sensors And Still Taught Us About Edge AI

    A failed 5G rollout in a legendary forest forced us to rethink everything we knew about AI infrastructure. Instead of pushing data to distant servers, we turned wearables, sensors, and tiny controllers into a cooperative network that can sense, decide, and act without the cloud. The result is a hands-on tour of decentralized AI: how to split models across devices, why feature fusion matters more than raw horsepower, and what it takes to make ad hoc networks reliable in the wild. We walk through practical patterns for collaboration at the edge, from complementary sensing in search-and-rescue to pooled compute in crowded venues. You’ll hear how we orchestrate parallel processing on microcontrollers, assign inference to one core and radio handling to another, and compress features to keep bandwidth low. We also dig into continual learning and federated averaging, outlining strategies to adapt models locally while protecting privacy and avoiding catastrophic forgetting. Along the way, we share early results from agriculture and public safety pilots, plus the gritty realities of hardware constraints, scarce datasets, and the challenge of testing at scale. If you’re curious about TinyML, edge AI, and how generative models might run collaboratively across many small devices, this conversation lays out a practical path forward. You’ll come away with a clearer picture of when decentralization beats centralized cloud systems, which protocols survive in noisy environments, and why the future of AI may look less like a monolith and more like a swarm. Subscribe, share this episode with a builder who loves constraints, and leave a review to tell us where you’d deploy a swarm of tiny models next. Send us Fan Mail Support the show Learn more about the EDGE AI FOUNDATION - edgeaifoundation.org

    • Transcript
    • Chapters
  • May 21 · 16 min

    How AI Compensates for PID Controller Limitations in Electric Vehicles with STMicroelectronics

    How can artificial intelligence transform electric vehicle performance? Discover the groundbreaking application of neural networks to motor control challenges that even Formula 1 legend Michael Schumacher helped identify. The automotive industry's electrification demands increasingly sophisticated silicon solutions, particularly for traction inverters controlling electric motors. Traditional control systems face a fundamental challenge: they must operate at extraordinary speeds (currently 20kHz, trending toward 100kHz) while managing rapid transitions between states. When drivers make sudden accelerator changes, conventional PID controllers produce energy-wasting overshoots that drain precious battery power. Our research presents a novel approach using neural networks to compensate for these limitations. By generating time-varying correction factors, our AI solution reduces maximum overshoots by up to 70% in demanding scenarios. This innovation represents a critical advancement for electric vehicle efficiency, potentially extending range and improving performance. What makes this application particularly fascinating is the extreme time constraints. While most AI applications process data at relatively leisurely rates (think 30 frames per second for vision systems), motor controllers must complete their calculations within microseconds. Our current implementation achieves 70-microsecond inference times on automotive-grade microcontrollers, with further optimizations planned through hardware acceleration. The collaboration between academic researchers and industry partners (MathWorks and STMicroelectronics) demonstrates the power of combining simulation expertise with real-world deployment capabilities. Using Simulink as the development platform and ST's developer cloud for automatic deployment to physical microcontrollers, we've created a streamlined methodology for applying AI to automotive control systems. Want to dive deeper into the technical details? Check out our published research paper on arXiv and discover how neural networks are transforming the heart of electric vehicle propulsion systems. Share your thoughts on how AI might further revolutionize automotive technology! Send us Fan Mail Support the show Learn more about the EDGE AI FOUNDATION - edgeaifoundation.org

    • Transcript
    • Chapters
  • May 14 · 20 min

    How to simplify and securely maintain up-to-date AI Models in the Edge

    Ever shipped a smart device and worried what happens after it leaves the lab? We dig into the hard parts of edge security—where models live on-device, firmware updates are routine, and attackers treat your fleet as a supply chain—then break them down into moves any team can adopt. From secure boot that blocks untrusted code at power-on to verified boot with discrete secure elements, we show how to anchor trust in hardware so software can prove itself before it runs. We walk through the real risks teams face—model theft, OTA hijacking, plaintext credentials in flash, and silent downgrades—and map them to practices that actually scale across mixed hardware. You’ll hear why encrypting data at rest frustrates drive cloning, how end-to-end encrypted and signed updates prevent tampering, and why automatic rollback turns “bricks” into recoverable hiccups. Updating AI models becomes a strength when you ship small, signed artifacts instead of full images, with logs that satisfy CRA and NIS2 audits while giving operators the visibility they need. We also tackle the build-versus-buy dilemma with clear-eyed math. Building a secure update stack across Qualcomm, NXP, PSoC, and diverse compute modules takes specialists and months; a platform approach spreads cost, speeds delivery, and still lets you own your keys so you can switch later without stranding devices. That key ownership underpins true end-to-end trust: you sign, devices verify, and the infrastructure moves at your pace. If you care about safeguarding IP, maintaining uptime, and earning customer trust, this is your blueprint. If this deep dive helps, follow the show, share it with your hardware and firmware teams, and leave a quick review—what part of your edge stack needs the strongest lock? Send us Fan Mail Support the show Learn more about the EDGE AI FOUNDATION - edgeaifoundation.org

    • Transcript
    • Chapters
  • May 7 · 15 min

    AI-Driven Brain-Computer Interface (BCI) Unlocking the Minds Potential

    Imagine steering a game or selecting a letter with nothing but a blink or a glance. We set out to make that feel normal, not magical, by building a non-invasive brain–computer interface that runs entirely on a low-power microcontroller and fits into everyday wearables like glasses. No surgery, no cloud dependency—just smart sensing, tight signal processing, and a tiny neural net that turns eye movements into reliable commands. We start with the “why”: millions live with motor impairments yet can still move their eyes, leaving a powerful window for communication and control. From there, we map the BCI landscape—high-precision invasive implants like Neuralink, BrainGate, and Synchron on one side; accessible non-invasive tools like Emotiv, Muse, and OpenBCI on the other—and unpack the trade-offs across accuracy, latency, cost, and ethics. Our approach uses electrostatic charge sensing to read subtle changes around the eyes, with electrodes positioned for comfort and signal quality. A lean pipeline cleans the data with high-pass, notch, and low-pass filters; a Z-score event detector wakes the model only when something meaningful happens. The model is a compact 1D CNN that classifies four classes—discard involuntary blinks, trigger with a voluntary blink, and detect left or right glances—achieving about 90% accuracy on a small multi-participant dataset. Running on an STM32H7, it uses roughly 18 KB flash and 6 KB RAM, with sub-millisecond inference; the overall response is driven by the short data window at 240 Hz, delivering real-time control for basic tasks. We demo blink-to-jump and look-to-steer gameplay to prove responsiveness and highlight how the same system could power communication aids and smart-home control. Looking ahead, we focus on integrating the electrodes into comfortable glasses, adding quick calibration for personal variability, and expanding the command set without sacrificing simplicity. If this mix of accessibility, edge AI, and practical human–machine interaction resonates with you, follow the show, share it with a friend, and leave a review so we can reach more builders and caregivers working on assistive tech. What would you control first with a glance? Send us Fan Mail Support the show Learn more about the EDGE AI FOUNDATION - edgeaifoundation.org

    • Transcript
    • Chapters
  • April 30 · 11 min

    An Embedded Transformer- base face recognition system in the STM32N6

    What if transformer-level face recognition could run on a microcontroller without giving up speed or accuracy? We set out to make that real on the STM32N6 by pairing its neural processing unit with a hybrid model that blends convolutional efficiency and attention-like global context. Along the way, we rewired core assumptions about attention, reworked unsupported operators, and delivered a full on-device pipeline that actually feels instant. We start with the hardware edge: ARM Cortex M55, 4 MB of continuous RAM, and an NPU pushing up to 600 GOPS at remarkable power efficiency. That lets us chain models—RetinaFace-style detection with landmarks, alignment for a stable canonical view, MobileNetV2 anti-spoofing to block print and replay attacks, and a final recognizer that outputs a 512‑dimensional embedding. The recognizer is built on EdgeFace, itself based on EdgeNext, chosen for its sweet spot between parameter count and accuracy. It behaves like a transformer where it matters—capturing long-range relationships—yet fits into the tight compute envelope of a microcontroller. The turning point is attention without the dot product. Because the ST toolchain doesn’t support batch matmul, we replaced it with a convolutional self-attention mechanism. Depthwise and pointwise convolutions encode relationships across pixels and channels, a sigmoid stands in for softmax, and element-wise products reconstruct attention’s weighting behavior. This maps cleanly to the NPU, avoids quadratic costs, and preserves the ability to stabilize identities across pose, lighting, and occlusion. Benchmarks show roughly 40 ms per frame end to end—about 25 FPS—plus substantial speedups over STM32H7 and higher accuracy than MobileFaceNet across validation sets. That opens doors for privacy-first access control, frictionless enrollment on-device, and personalized experiences where latency matters and data should never leave the edge. If you’re exploring embedded AI, this walkthrough shows how to align model design with silicon capabilities and deliver results that feel both fast and trustworthy. Enjoy the deep dive? Subscribe, share this episode with a fellow edge AI builder, and leave a quick review to help others find the show. Send us Fan Mail Support the show Learn more about the EDGE AI FOUNDATION - edgeaifoundation.org

    • Transcript
    • Chapters
  • April 23 · 18 min

    Verification, Validation & Certification of AI in Safety-Critical Applications

    A cyclist disappears to the model, not to your eyes—and that mismatch is the heart of safety-critical AI. We open with the “vanishing cyclist” to show how tiny, imperceptible perturbations can flip life-or-death decisions, then walk through a practical path to trust that spans data, verification, and deployment. Along the way, we share real stories from BMW, Airbus, and Madrid Metro to ground the engineering in results, not hype. We break down how to build a resilient pipeline: domain-specific data labeling, realistic synthetic generation for rare and risky scenarios, and tight interoperability across MATLAB, Python, PyTorch, TensorFlow, and ONNX. We dig into explainability beyond classification with D-RISE for object detectors and semantic segmentation, helping you see what the network actually uses to decide. Then we raise the bar with formal verification for robustness—mathematical guarantees within defined perturbation sets—so you aren’t mistaking the absence of found attacks for true safety. Finally, we get practical about the edge. Model compression and projection recover accuracy with fewer parameters, enabling fast, power-efficient deployment to CPUs, GPUs, and FPGAs, backed by code generation for the entire application. We also cover runtime safeguards like out-of-distribution detection to catch smog-on-the-runway moments and escalate safely. Throughout, we connect the work to evolving standards, the EU AI Act, and updated workflows that adapt the V-model for learning systems, so your process and artifacts are ready for audits and certification. If you care about trustworthy AI for cars, planes, rail, and medical devices—and want tools and habits that survive contact with reality—this one’s for you. Listen, subscribe, and leave a review with your biggest trust gap or the safeguard you’d ship first. Send us Fan Mail Support the show Learn more about the EDGE AI FOUNDATION - edgeaifoundation.org

    • Transcript
    • Chapters
  • April 16 · 20 min

    Aptos: Creating ML models that fit your edge device like a glove

    Shipping edge AI shouldn’t feel like a marathon through model zoos, missing ops, and latency ceilings. We lay out a practical path to get from your data and constraints to a hardware-ready model—measured on real boards—without the endless back-and-forth between data science and firmware teams. If you’ve wrestled with quantization loss, unsupported kernels, or picking the “right” NPU, this walkthrough will feel like oxygen. We start by naming the pain: quick demos that collapse under real device limits, foundation models that fail after export, and feedback loops that burn months. From there, we unpack Aptos, our automation engine that turns edge AI into a data in, model out process. The system explores parameterized architecture recipes and neural architecture search, trains promising candidates, and deploys them to a hardware farm packed with evaluation kits. Every candidate returns hard numbers—latency, per-layer timing, memory, on-device accuracy, and power—so tradeoffs are grounded in measurements, not wishful thinking. What makes it fast is the learning layer. As Aptos accumulates results, meta models predict runtime, memory fit, and stable hyperparameter ranges before committing compute. That means less time wasted on dead ends and more time converging on models that satisfy your KPIs, whether you care about sub-5 ms inference on an i.MX 8 Plus, battery life in the field, or non-square inputs that match your camera feed. We also fold in research-backed techniques—pruning, quantization, distillation—so you benefit from the latest without chasing papers. If your team is eyeing a chip migration or evaluating new NPUs, a dropdown swap in Aptos triggers a fresh search tuned to the new hardware, minimizing lock-in and keeping options open. The result is timeline compression: where projects used to take 12–18 months with large teams, we aim to surface strong, deployable candidates in one to two weeks. Subscribe for more deep dives into edge AI deployment, share this episode with your team, and leave a review telling us which device you want to target next. Send us Fan Mail Support the show Learn more about the EDGE AI FOUNDATION - edgeaifoundation.org

    • Transcript
    • Chapters
Showing 1–20 of 20 episodes