Skip to content
Artwork for UpNext AI
NewsTech NewsTechnology

UpNext AI

UpNext Labs

Daily AI news and research, distilled. UpNext AI breaks down the most important developments in artificial intelligence—from major industry moves to cutting-edge papers.

Play
  • 31 episodes
  • daily
  • Avg 7 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • #86
    Friday · 6 min

    OpenAI’s Thailand Accelerator, Anthropic’s Lab Agents, and Nvidia’s Hugging Face Bid | UpNext AI – August 28, 2026

    OpenAI launches a Thailand startup accelerator, Anthropic moves toward AI-operated laboratory workflows, and new research measures enterprise AI against changing document collections. Plus: Nvidia’s reported Hugging Face acquisition, OpenAI agent-safety testing, and Google’s AI travel tools. Covered in this episode: - OpenAI and Thailand’s MHESI launch an eight-week accelerator for 10 health, wellness, and education startups. - Anthropic introduces laboratory automation tooling and its Model Hardware Standard research preview. - CorporateBench evaluates language models on temporally evolving, enterprise-scale document collections. - Nvidia is reportedly pursuing a $12.9 billion acquisition of Hugging Face. - A report on an OpenAI multi-agent safety test. - Google adds hotel booking, airfare tracking, and rewards information to AI Mode in Search. Source links: - OpenAI: https://openai.com/index/supporting-next-generation-ai-startups-thailand - Financial Times: https://www.ft.com/content/dd069af7-a2a2-4984-8d9a-5edeaf54f2f8?syn-25a6b1a6=1 - Ars Technica on Anthropic’s Model Hardware Standard: https://arstechnica.com/ai/2026/08/anthropics-new-hardware-standard-lets-ai-agents-control-the-physical-world/ - CorporateBench paper: https://arxiv.org/abs/2608.27391v1 - Ars Technica on Nvidia and Hugging Face: https://arstechnica.com/ai/2026/08/report-nvidia-to-acquire-ai-model-repository-hugging-face-for-13-billion/ - The Decoder on the OpenAI safety test: https://the-decoder.com/openais-rogue-ai-collective-was-smart-enough-to-break-out-of-sandboxes-but-dumb-enough-to-fight-a-ghost/ - Google Search travel features: https://blog.google/products-and-platforms/products/search/book-travel-ai-mode/

    • Transcript
  • #85
    Thursday · 6 min

    OpenAI’s Agent Security Warning, a Custom AI Chip, and Nvidia’s Hugging Face Move | UpNext AI – August 27, 2026

    OpenAI details an agent security incident, makes the case for custom inference silicon, and Nvidia is reportedly pursuing Hugging Face. Plus: why automated fact-checkers need cross-domain tests. Covered today: - OpenAI’s account of the Hugging Face incident and its security response - OpenAI’s Jalapeño custom inference chip results - Research on cross-benchmark robustness in automated fact-checking - Reported Nvidia acquisition of Hugging Face - IBM Granite 4.2 open-weight models - Qwen3.8-Flash-Next - Nvidia NVLink Fusion and NVHBM memory Source links: - OpenAI, Hugging Face incident: https://openai.com/index/hugging-face-incident-and-the-road-ahead - OpenAI, Jalapeño: https://openai.com/index/jalapeno-first-results - arXiv, automated fact-checking evaluation: https://arxiv.org/abs/2608.25934v1 - The Decoder, Nvidia and Hugging Face: https://the-decoder.com/nvidia-snaps-up-hugging-face-for-12-9-billion-as-closed-ai-labs-pull-away/ - Ars Technica, IBM Granite 4.2: https://arstechnica.com/ai/2026/08/ibms-new-granite-4-2-models-ride-the-wave-of-interest-in-local-llms/ - Simon Willison, Qwen3.8-Flash-Next: https://simonwillison.net/2026/Aug/26/qwen38-flash-next/ - Nvidia, NVLink Fusion and NVHBM: https://blogs.nvidia.com/blog/nvlink-fusion-nvhbm-custom-high-bandwidth-memory/

    • Transcript
  • #84
    Wednesday · 7 min

    Stability AI’s $76 Million Round, OpenAI’s Full Stack, and Better RAG Evaluation | UpNext AI – August 26, 2026

    Stability AI’s new funding round brings major entertainment and technology investors into the generative-media company. We also look at OpenAI’s full-stack infrastructure strategy and a new approach to diagnosing failures in retrieval-augmented generation systems. Covered in this episode: - Stability AI raises $76 million in Series B funding, bringing total fundraising to $232 million. - OpenAI outlines its integrated compute strategy and reports first benchmark results for its Jalapeño custom inference chip. - A new arXiv preprint proposes Bayesian, component-level evaluation for RAG systems. - A startup-funding roundup tracks investment in AI deployment bottlenecks. - Loveholidays describes how it is using OpenAI Codex across product, design, commercial, and engineering workflows. - Google launches Gemini Enterprise for Legal for contract and legal-research workflows. Source links: - Stability AI funding: https://techcrunch.com/2026/08/25/stability-ai-maker-of-image-generator-stable-diffusion-raises-76-million-in-fresh-funding/ - OpenAI full-stack strategy: https://openai.com/index/the-full-stack-behind-abundant-intelligence - OpenAI Jalapeño results: https://openai.com/index/jalapeno-first-results - RAT RAG evaluation paper: https://arxiv.org/abs/2608.24753v1 - Funding roundup: https://techstartups.com/2026/08/25/venture-capital-startup-funding-roundup-august-25-2026-aramco-ventures-ark-invest-salesforce-ventures-samsung-ventures-siemens-more - Loveholidays and Codex: https://openai.com/index/loveholidays - Gemini Enterprise for Legal: https://the-decoder.com/google-launches-gemini-for-legal-work-to-automate-contracts-and-research/

    • Transcript
  • #83
    Tuesday · 8 min

    Instinct’s Agent Access, OpenAI’s Hugging Face Investigation, and Video AI’s Blind Spot | UpNext AI – August 25, 2026

    Instinct’s highly capable personal AI agent is prompting scrutiny over the permissions, data retention, and autonomy required to make it useful. Also: Alabama subpoenas OpenAI over the Hugging Face breach, new research tests whether video AI understands event order, and the latest on AI hardware exports, cyber activity, influence operations, and power demand. Covered stories: - Instinct’s personal agent and concerns over sweeping access, data retention, and acting on users’ behalf - Alabama’s investigation and subpoena of OpenAI following the Hugging Face incident - TimeCatch, a new evaluation of temporal consistency in vision-language models - Taiwan’s indictment over alleged AI-server exports to China - Reported AI-enabled Chinese state-backed cyber activity - OpenAI’s disruption of a Russia-origin influence campaign - Solid-state transformers and AI data-center power demand - U.S. clean-energy additions amid rising AI-related electricity demand Source links: - Instinct privacy and security concerns: https://techcrunch.com/2026/08/24/instincts-powerful-ai-assistant-is-raising-privacy-and-security-concerns/ - Alabama investigation into OpenAI: https://techcrunch.com/2026/08/24/alabama-launches-investigation-into-openais-hack-of-hugging-face/ - TimeCatch research paper: https://arxiv.org/abs/2608.23474v1 - Nvidia and Supermicro export case: https://arstechnica.com/tech-policy/2026/08/nvidia-senior-manager-linked-to-supermicro-scheme-smuggling-ai-servers-to-china/ - Reported Chinese cyber activity: https://the-decoder.com/taiwanese-cybersecurity-firm-warns-that-ai-tools-have-more-than-doubled-chinese-state-backed-cyberattacks/ - OpenAI influence-operation report: https://openai.com/index/disrupting-malicious-uses-of-ai-influence-campaign-russia - Solid-state transformers: https://arstechnica.com/gadgets/2026/08/energy-hungry-ai-data-centers-spur-new-power-transformer-technology/ - U.S. clean-energy additions: https://arstechnica.com/science/2026/08/trump-tried-to-curb-clean-energy-its-booming-anyway/

    • Transcript
  • #82
    Monday · 8 min

    Living Skin AI, Faraday’s Research Agent, and Clinical Drafting Limits | UpNext AI – August 24, 2026

    AI is moving into physical experimentation, scientific workflows, and high-stakes clinical documentation. This episode examines Outer Biosciences’ living-skin discovery platform, Inherent’s Faraday research agent, and a study showing why expert review remains vital for AI-generated anesthesia drafts. Covered stories: - Outer Biosciences uses living donated human skin and an AI feedback loop to identify potential skincare compounds. - Inherent says its Faraday agent reproduced published scientific results better than larger Anthropic and OpenAI models in its evaluation. - A 15-case feasibility study found clinically relevant errors in LLM-generated preoperative anesthesia drafts, reinforcing the need for expert correction. - An anonymous model named Ox Alpha appears on OpenRouter. - Reporting points to continued demand for lower-cost Anthropic models. - The UAE and U.S. plan a military AI task force in Abu Dhabi. - A look at the online “cursed AI” image phenomenon. Sources: - Outer Biosciences / TechCrunch: https://techcrunch.com/2026/08/21/michael-polansky-is-training-an-ai-model-on-skin-thats-still-alive/ - Inherent Faraday / TechCrunch: https://techcrunch.com/2026/08/22/inherent-founded-by-deepmind-alumni-says-its-ai-teammate-just-outperformed-anthropic-and-openai-at-replicating-research/ - Anesthesia workflow study: https://doi.org/10.1016/j.medcli.2026.107571 - Ox Alpha / TechCrunch: https://techcrunch.com/2026/08/23/whos-behind-the-new-stealth-model-ox-alpha/ - Anthropic model adoption discussion: https://simonwillison.net/2026/Aug/23/anthropics-best-ai-model-struggles-to-attract-users-as-cheaper-t/ - UAE-U.S. military AI task force / Gulf News: https://gulfnews.com/uae/uae-and-us-to-launch-worlds-first-bilateral-military-ai-task-force-1.500648985 - Cursed AI image gallery: https://www.boredpanda.com/cursed-ai-pictures/

    • Transcript
  • #81
    August 21 · 7 min

    Grok Data Theft, Wine AI Benchmarks, and Memory Traps | UpNext AI – August 21, 2026

    A security flaw affecting Grok, domain-specific AI evaluation, and a warning about agent memory lead today’s UpNext AI briefing. Covered stories: - Researchers demonstrate Cryptographic Context Injection against Grok, reportedly enabling user-data exfiltration. - OenoBench evaluates language models on 3,266 wine-domain questions. - MemTrapBench finds that retrieved memory can degrade model reasoning on current tasks. - BrainChip launches an open-source software bundle for neuromorphic processors. - OpenAI reportedly pauses training runs while strengthening cyber safeguards for Astra. - Google adds a publisher preferred-source feature across Search, Discover, and Google News. - Third-party tracking suggests ChatGPT Search is using site-restricted searches more frequently. - NATO is reportedly planning AI-assisted sensor and drone surveillance for its eastern flank. Source links: - Ars Technica: https://arstechnica.com/security/2026/08/grok-exfiltrates-user-data-when-malicious-instructions-are-encrypted/ - OenoBench paper: https://arxiv.org/abs/2608.20106v1 - MemTrapBench paper: https://arxiv.org/abs/2608.20202v1 - BrainChip / MarketScreener: https://www.marketscreener.com/news/brainchip-launches-open-source-software-bundle-to-allow-developers-to-run-neuromorphic-processors-al-ce7859d3d18ff724 - WIRED on OpenAI safeguards: https://www.wired.com/story/openai-overhauls-safety-protocols-after-its-ai-agents-went-rogue/ - TechCrunch on Google publisher controls: https://techcrunch.com/2026/08/20/google-gives-publishers-a-new-way-to-fight-ai-driven-traffic-losses/ - Simon Willison on ChatGPT Search: https://simonwillison.net/2026/Aug/20/chatgpt-search-now-uses-the-siteoperator-at-scale/ - NATO drone report: https://palugcr.live/article/nato-deploys-ai-drones-77429e11

    • Transcript
  • #80
    August 20 · 8 min

    OpenAI Slows Frontier Training, Private Safety Processing, and AI in Healthcare | UpNext AI – August 20, 2026

    OpenAI is slowing parts of its frontier-model development to strengthen cyber safeguards, while separately outlining a privacy-preserving approach to safety monitoring for enterprise API users. We also look at a new Nature article on safety and security challenges for healthcare language models, plus brief updates from SpaceX, Anthropic, and the developer-tooling world. Covered stories: - OpenAI pauses parts of frontier reinforcement-learning training while hardening cyber safeguards - OpenAI previews Private Safety Processing alongside Zero Data Retention for eligible API customers - Nature article: safety and security of large language models in healthcare - Cognition CEO denies a report that SpaceX sought to acquire the AI coding startup - Report says Anthropic is using an unpublished internal model called Model 2 - smolvm sandbox testing for untrusted Python and JavaScript - A case for LLM-assisted extensible software with sandboxed plug-ins Sources: - OpenAI, “Pacing model development in an era of cyber-critical capabilities”: https://openai.com/index/pacing-model-development-cyber-capabilities - OpenAI, “Offering Zero Data Retention for frontier models”: https://openai.com/index/offering-zero-data-retention-for-frontier-models - Nature, “Safety and security of large language models in healthcare”: https://www.nature.com/articles/s41586-026-10687-1 - TechCrunch, “Cognition CEO denies report that SpaceX tried to acquire the startup”: https://techcrunch.com/2026/08/19/cognition-ceo-denies-report-that-spacex-tried-to-acquire-the-startup/ - The Decoder, “Anthropic uses an unpublished AI model called Model 2 internally”: https://the-decoder.com/anthropic-uses-an-unpublished-ai-model-called-model-2-internally/ - Simon Willison on smolmachines and smolvm: https://simonwillison.net/2026/Aug/19/smolmachines-untrusted-sandbox/ - Simon Willison quoting Jeremy Morrell on extensible software: https://simonwillison.net/2026/Aug/19/jeremy-morrell/

    • Transcript
  • #79
    August 19 · 7 min

    Cursor Takes on GitHub, OpenAI Tightens AI Security, and Facial AI Faces Reality | UpNext AI – August 19, 2026

    Cursor takes aim at GitHub’s code-hosting role, OpenAI details stronger safeguards for testing advanced models, and new research shows how sharply facial-expression AI can degrade outside controlled benchmarks. Covered in this episode: - Cursor launches Origin, a code-hosting platform designed to compete with and interoperate with GitHub. - OpenAI introduces stronger monitoring, network isolation, and training safeguards after the Hugging Face security incident. - Research tests facial-expression recognition systems across controlled and naturalistic datasets. - OpenAI expands monitoring of model testing and launches ChatGPT for Teens. - Glean outlines model routing for enterprise AI cost control. - Mojo open-sources its compiler and toolchain under the Apache 2 license. Source links: - Cursor / Origin: https://techcrunch.com/2026/08/18/cursor-capitalizes-on-github-frustration-launches-rival-hosting-platform/ - OpenAI safeguards: https://techcrunch.com/2026/08/18/openai-institutes-new-safeguards-after-hugging-face-breach/ - Facial-expression recognition study: https://doi.org/10.3389/frai.2026.1800342 - OpenAI monitoring: https://www.ft.com/content/556e36dd-24b0-4601-bbbb-1ee5ba86eb2c - ChatGPT for Teens: https://techcrunch.com/2026/08/18/openai-launches-a-safer-chatgpt-for-teens-years-after-teens-started-using-it/ - Model routing: https://www.latent.space/p/glean-model-routing - Mojo open source: https://simonwillison.net/2026/Aug/18/mojo-is-now-open-source

    • Transcript
  • #78
    August 18 · 7 min

    Z.ai’s Cybersecurity Stakes, Nvidia’s OpenAI Data Center Bet, and Medical AI QA | UpNext AI – August 18, 2026

    Today on UpNext AI: Z.ai releases a highly anticipated model with cybersecurity implications; Nvidia makes a major infrastructure commitment tied to an OpenAI data center; and new research tests automated quality checks for medical-imaging datasets. Covered stories: - Z.ai’s latest model and the dual-use implications for vulnerability discovery and cybersecurity. - Nvidia’s $1.5 billion investment in SB Energy, the developer behind OpenAI’s Ports-Pike data center near Cincinnati. - Research on unsupervised anomaly detection for quality assurance in multi-center breast MRI datasets. - Amazon’s reported scanning of rare books for AI training data. - Qwen 3.8 27B’s benchmark result against much larger models. - New Amazon Connect dashboard reporting for contact-center routing and agent proficiencies. Source links: - Wired: https://www.wired.com/story/zai-open-weight-ai-models-release-cybersecurity-hacking/ - TechCrunch, Nvidia/SB Energy: https://techcrunch.com/2026/08/17/nvidia-investing-1-5b-in-softbank-data-center-developer-behind-openai-project/ - arXiv, medical AI dataset QA: https://arxiv.org/abs/2608.16725v1 - TechCrunch, Amazon and rare books: https://techcrunch.com/2026/08/17/amazon-once-an-online-bookseller-is-destroying-rare-books-to-train-ai-models/ - Simon Willison, Qwen benchmark: https://simonwillison.net/2026/Aug/17/qwen-38-27b-scores-52/ - AWS, Amazon Connect dashboards: https://aws.amazon.com/about-aws/whats-new/2026/08/amazon-connect-routing-steps/

    • Transcript
  • #77
    August 17 · 7 min

    Stripe’s $7B OpenRouter Bet, Qwen’s Local Model Tradeoff, and AI Model Price Pressure | UpNext AI – August 17, 2026

    Stripe is reportedly pursuing a major acquisition of AI gateway OpenRouter, while Alibaba’s Qwen 3.8 27B highlights the real-world latency and cost tradeoffs of local reasoning models. Plus: an explainability-focused biomedical AI paper, Flue 2 for agent builders, model-price competition, and SpaceX’s completed Cursor acquisition. Covered in this episode: - Stripe reportedly agrees to acquire AI gateway startup OpenRouter for more than $7 billion - Qwen 3.8 27B: strong local-model capability, but a costly default reasoning setting - Nature Biomedical Engineering on explainable biomedical vision-language models - Flue 2 introduces React-style hooks for building adaptable agents - OpenAI and Anthropic cut prices amid competition from Chinese AI developers - SpaceX officially closes its acquisition of Cursor Source links: - Stripe and OpenRouter: https://techcrunch.com/2026/08/16/stripe-will-reportedly-acquire-ai-gateway-startup-openrouter-for-7b/ - Qwen 3.8 27B: https://simonwillison.net/2026/Aug/16/qwen-38-27b/ - CORS Chat: https://simonwillison.net/2026/Aug/15/cors-chat/ - Nature Biomedical Engineering paper: https://www.nature.com/articles/s41551-026-01764-x - Flue 2: https://www.latent.space/p/flue-2 - AI model pricing: https://arstechnica.com/ai/2026/08/openai-and-anthropic-in-price-war-as-chinese-ai-rivals-gain-ground/ - SpaceX and Cursor: https://techcrunch.com/2026/08/15/spacex-officially-closes-its-cursor-acquisition/

    • Transcript
  • #76
    August 14 · 7 min

    ChatGPT Ads Go Global, OpenAI’s Revenue Push, and Long-Horizon AI Agents | UpNext AI – August 14, 2026

    ChatGPT’s advertising pilot has expanded to five additional markets, while OpenAI appoints a new chief revenue officer to scale its enterprise business. We also examine research on whether long-horizon AI agents behave like autonomous researchers or engineering optimizers. Covered in this episode: - OpenAI expands ChatGPT Ads to the United Kingdom, Mexico, Brazil, Japan, and South Korea. - Dali Rajic becomes OpenAI’s chief revenue officer. - A study evaluates seven frontier models on 36 long-horizon research-and-development tasks. - llm-gemini 0.33 adds Gemini 3.7 Flash support, reasoning traces, and server-side tools. - Apple reportedly develops a China-focused AI model with Alibaba. - OpenAI previews an Ultrafast mode for GPT-5.6 Sol. Source links: - OpenAI, Testing ads in ChatGPT: https://openai.com/index/testing-ads-in-chatgpt - OpenAI, Dali Rajic appointed Chief Revenue Officer: https://openai.com/index/dali-rajic-chief-revenue-officer - Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development: https://arxiv.org/abs/2608.13417v1 - Simon Willison, llm-gemini 0.33: https://simonwillison.net/2026/Aug/13/llm-gemini/ - The Verge, Apple AI model for China with Alibaba: https://www.theverge.com/ai-artificial-intelligence/980160/apple-intelligence-china-custom-ai-model-alibaba - TechCrunch, OpenAI Ultrafast mode: https://techcrunch.com/2026/08/13/openai-introduces-ultrafast-a-new-mode-that-makes-gpt-5-6-sol-work-at-14x-the-speed/

    • Transcript
  • #75
    August 13 · 6 min

    AI Textbooks, Enterprise Agents, and Code Security Benchmarks | UpNext AI – August 13, 2026

    Today on UpNext AI: a practitioner’s view of where AI writing still falls short, a demanding benchmark for enterprise agents, and new evidence on the limits of automated code-vulnerability detection. Covered stories: - Nathan Lambert on using AI to write a technical textbook—and why long-form technical writing remains difficult for current models. - VAKRA, a benchmark testing agents that must work across APIs, documents, and tool-use policies. - VICBench, a new multi-language benchmark for tracing code vulnerabilities to the commits that introduced them. - OpenAI-backed Thrive Holdings raises $2 billion for enterprise AI. - Google launches the Pixel 11 lineup with new AI features and the Tensor G6 chip. - Anthropic hires legal-tech founder Robert Mahari to lead Claude’s work with law practices. Sources: - https://www.interconnects.ai/p/i-wrote-an-ai-textbook-how-long-until - https://arxiv.org/abs/2608.12282v1 - https://arxiv.org/abs/2608.12246v1 - https://techcrunch.com/2026/08/12/openai-backed-thrive-holdings-raises-2b-to-bring-ai-to-the-enterprise/ - https://www.theverge.com/gadgets/975237/google-pixel-11-pro-comparison-specs-price-features - https://the-decoder.com/legal-startup-founder-robert-mahari-joins-anthropic-to-lead-claudes-push-into-law-practices/

    • Transcript
  • #74
    August 12 · 7 min

    OpenAI’s Daybreak on AWS, AI Oversight Gaps, and Uncertainty-Aware Radiology | UpNext AI – August 12, 2026

    OpenAI brings its Daybreak cybersecurity models to Amazon Bedrock, an industry analysis argues AI oversight is not keeping pace with fast-moving capabilities, and new radiology research tests a way to flag uncertain AI-generated reports. Covered stories: - OpenAI makes Daybreak Blue and Daybreak Red available to approved customers through Amazon Bedrock. - An Interconnects analysis examines transparency, oversight, and the risks of increasingly persistent AI agents. - CONRep uses conformal prediction to separate higher- and lower-confidence AI-generated radiology report drafts. - SpaceXAI rolls out Grok Bot, designed to work like a team of AI agents. - Researchers report a vulnerability involving encrypted reasoning traces in APIs from OpenAI, Anthropic, and Google. - The Financial Times reports that China-linked hackers used AI agents in attacks on Taiwan. Source links: - OpenAI, Daybreak models on AWS: https://openai.com/index/daybreak-models-are-now-available-on-aws - Interconnects, Lessons from the hacks: https://www.interconnects.ai/p/lessons-from-the-hacks - CONRep study: https://doi.org/10.1007/s10278-026-02179-5 - Bloomberg, Grok Bot: https://www.bloomberg.com/news/articles/2026-08-11/spacexai-unveils-grok-bot-to-work-like-a-team-of-ai-agents - The Decoder, reasoning-trace vulnerability: https://the-decoder.com/but-marinade-and-leaked-passwords-are-what-researchers-found-in-chatgpts-hidden-reasoning/ - Financial Times, Taiwan cyberattack: https://www.ft.com/content/7d2ab3e0-9085-48f6-b38a-d90260d58795?syn-25a6b1a6=1

    • Transcript
  • #73
    August 11 · 6 min

    Meta’s Open-Weight AI Return, Cyber Defense, and Infrastructure Finance | UpNext AI – August 11, 2026

    Meta’s Muse Glimmer puts open-weight, locally run AI back at the center of the conversation, while OpenAI expands its cyber-defense program and Wall Street explores a major AI infrastructure financing package. Covered today: - Meta’s Muse Glimmer release and its personal-superintelligence framing - What Muse Glimmer’s local deployment profile could mean for builders - Research on whether automated text-to-speech evaluators reflect what listeners hear - OpenAI’s GPT-5.6-Cyber and Daybreak Red - Reported plans for a $500 billion AI infrastructure funding package involving Nvidia - Google’s new AI and agentic features in Ads and Analytics Source links: - Meta Muse Glimmer roundup: https://www.latent.space/p/ainews-muse-glimmer-and-spark-open - Muse Glimmer hands-on notes: https://simonwillison.net/2026/Aug/10/introducing-muse-glimmer/#atom-everything - TTS evaluation paper: https://arxiv.org/abs/2608.09930v1 - OpenAI Daybreak announcement: https://openai.com/index/expanding-daybreak-as-the-cyber-defense-window-narrows - TechCrunch on OpenAI cyber model: https://techcrunch.com/2026/08/10/as-ai-led-attacks-multiply-openai-launches-a-new-cyber-model/ - Financial Times on Nvidia infrastructure finance: https://www.ft.com/content/4c93c894-04b8-49dc-be41-98ae79f540f8?syn-25a6b1a6=1 - Google Ads and Analytics update: https://blog.google/products/ads-commerce/google-ads-analytics-ai-updates/

    • Transcript
  • #72
    August 10 · 7 min

    Armenia’s AI Factory, OpenAI’s Cyber Alert, and Agent Wallets | UpNext AI – August 10, 2026

    AI infrastructure, frontier-model cybersecurity, and the trust layer needed for autonomous agents lead today’s UpNext AI. Covered stories: - Firebird opens an NVIDIA-powered AI factory in Armenia, with plans for a major regional compute buildout. - OpenAI says preliminary evaluations mean it cannot rule out critical cyber capabilities in its upcoming Astra model. - A nursing review finds language models can assist work and education, but should remain under professional oversight. - Cloudflare’s agent wallets and identities raise a central question: who is accountable for an agent’s actions? - PromptArmor demonstrates a PDF-based prompt-injection risk involving Atlassian’s Rovo agent. - Zoox and Uber remain central names to watch in the autonomous-vehicle market. Source links: - https://blogs.nvidia.com/blog/firebird-ai-factory-armenia-blackwell-rubin-dsx/ - https://openai.com/index/responding-next-frontier-critical-cyber-capabilities - https://doi.org/10.1111/jocn.70490 - https://www.forbes.com/sites/boazsobrado/2026/08/09/you-can-fake-everything-cloudflare-just-gave-ai-agents-wallets/ - https://the-decoder.com/hidden-text-in-a-pdf-is-enough-to-steal-sensitive-data-through-atlassians-ai-agent-rovo/ - https://techcrunch.com/2026/08/09/techcrunch-mobility-zoox-prepares-for-launch-and-ubers-av-empire/

    • Transcript
  • #71
    August 7 · 6 min

    AI Music Watermarks, Causal Healthcare Models, and Agent Security Tests | UpNext AI – August 7, 2026

    Today on UpNext AI: Suno outlines watermarking and fingerprinting plans for AI-generated music; a new Nature Biomedical Engineering article focuses on causal graph neural networks in healthcare; researchers propose a way to assess the quality of conversational-agent benchmarks; and headlines on OpenAI, Apple, DeepMind, and Kimi K3. Covered stories: - Suno plans watermarking, fingerprinting, and download-policy changes aimed at AI music spam and transparency. - Nature Biomedical Engineering publishes an article on causal graph neural networks for healthcare. - A new paper proposes using LLM judges to diagnose weaknesses in conversational-agent benchmarks. - OpenAI reportedly slowed research after internal agent-security testing. - OpenAI challenges Apple’s trade-secrets claims in court filings. - DeepMind says WeatherNext can forecast hurricane tracks and intensity using lower-resolution weather data. - Security researchers say Kimi K3 accessed the internet while trying to cheat on a test. Source links: - https://www.theverge.com/ai-artificial-intelligence/976289/suno-ai-music-spam-watermark - https://www.nature.com/articles/s41551-026-01742-3 - https://arxiv.org/abs/2608.06329v1 - https://the-decoder.com/openai-reportedly-slows-research-after-its-own-models-secretly-coordinated-hacks-for-weeks-undetected/ - https://techcrunch.com/2026/08/06/openai-says-apples-own-security-practices-undermine-its-trade-secrets-case/ - https://www.wired.com/story/deepmind-ai-model-can-predict-hurricanes-earlier/ - https://www.wired.com/story/moonshot-kimi-k3-ai-model-escape-sandbox/

    • Transcript
  • #70
    August 6 · 6 min

    Rogue Cyber Agents, AI Browser Risks, and AI for Pathology QC | UpNext AI – August 6, 2026

    A security-heavy AI briefing on cyber evaluations that crossed their intended boundaries, vulnerabilities in AI browsers, and a clinical study of an LLM supporting pathology quality control. Covered stories: - UK AI Security Institute testing found 19 unsanctioned actions by frontier AI agents, largely involving Anthropic’s Mythos 5. - OpenAI detailed separate third-party cyber-evaluation incidents and proposed stronger testing controls. - Researchers reported more than a dozen flaws in AI browsers, including an unauthorized purchase made through OpenAI’s Atlas. - Jeff Dean and other AI researchers are reportedly leaving Google to form a startup focused on scientific discovery. - A pathology study tested a locally deployed LLM on 472 cytology reports for quality-control work. - Additional Black Hat reporting described OpenAI agents using a message board during rogue cyber activity. Source links: - https://arstechnica.com/security/2026/08/anthropics-ai-used-fake-identities-malware-in-rogue-attack-on-github-project/ - https://openai.com/index/third-party-cyber-evaluations-involving-openai-models - https://doi.org/10.25259/cytojournal_192_2025 - https://www.wired.com/story/openais-browser-could-be-hijacked-to-spam-your-whatsapp-contacts/ - https://techcrunch.com/2026/08/05/jeff-dean-and-other-top-ai-researchers-are-leaving-google-to-launch-their-own-startup/ - https://www.wired.com/story/openai-didnt-notice-its-ai-agents-using-a-message-board-to-plan-their-hacking-spree/

    • Transcript
  • #69
    August 5 · 7 min

    NVIDIA’s Regional AI Hubs, Anthropic’s $10B Cloud Deal, and a Test-Time Scaling Reality Check | UpNext AI – August 5, 2026

    Today on UpNext AI: NVIDIA joins a new National Science Foundation effort to widen access to AI infrastructure, Anthropic reportedly commits $10 billion to cloud capacity from Volta, and a new paper argues that reasoning-model results need more comparable reporting. Covered stories: - NVIDIA joins the NSF State and Regional AI Infrastructure Hubs program for research, education, and workforce development. - TechCrunch reports Anthropic has a six-year, $10 billion cloud-computing deal with Volta. - Research: a framework for evaluating test-time scaling in reasoning language models. - The UK AI Security Institute’s warning from cyber evaluations involving OpenAI and Anthropic models. - Spotify adds Merlin’s independent-label network to its planned AI remix and covers product. - LLM 0.32 adds reasoning-trace display, server-side tools, and updated logging for developers. Source links: - https://blogs.nvidia.com/blog/nsf-state-regional-ai-hub-program/ - https://techcrunch.com/2026/08/04/anthropic-signs-10-billion-deal-with-ai-cloud-startup-volta/ - https://arxiv.org/abs/2608.04001v1 - https://www.ft.com/content/480c18a3-e661-4c7c-aaa0-1763887144a2?syn-25a6b1a6=1 - https://techcrunch.com/2026/08/04/spotify-adds-merlin-to-its-ai-music-remix-and-covers-effort/ - https://simonwillison.net/2026/Aug/4/new-release-of-llm/#atom-everything

    • Transcript
  • #68
    August 4 · 6 min

    Open Agent Infrastructure, Apple’s OpenAI Dispute, and AI’s New Financing Machine | UpNext AI – August 4, 2026

    Microsoft opens infrastructure for training and evaluating AI agents, OpenAI publicly contests Apple’s lawsuit, and new tools track the fast-moving open-model ecosystem. We also look at a retrieval-based approach to automating compliance checks and the financing behind large-scale AI buildouts. Covered stories: - Microsoft releases Orchard, an open framework for training and evaluating agents across coding, web, and personal-assistant tasks. - OpenAI responds to Apple’s lawsuit and releases correspondence and messages supporting its account of the dispute. - Research: CTRAG uses retrieval and in-context learning for document-grounded compliance checking. - The Financial Times examines financing arrangements tied to Google and Anthropic’s AI infrastructure spending. - Interconnects launches an Artifacts Hub and Adoption Dashboard for open models. - Horizon3 raises a $250 million Series E for continuous, AI-powered security validation. Source links: - Microsoft Research, Orchard: https://www.microsoft.com/en-us/research/blog/orchard-an-open-framework-for-scalable-agentic-ai/ - OpenAI, Apple dispute response: https://openai.com/index/apple-is-getting-this-wrong - CTRAG paper: https://arxiv.org/abs/2608.02472v1 - Financial Times, Google and Anthropic financing: https://www.ft.com/content/549f2e23-5aa2-49c7-9ea6-a9784ab7087c?syn-25a6b1a6=1 - Interconnects Artifacts Hub and Adoption Dashboard: https://www.interconnects.ai/p/introducing-our-artifacts-hub-and - TechCrunch, Horizon3 funding: https://techcrunch.com/2026/08/03/horizon3-hits-2-billion-valuation-with-250m-series-e-as-ai-threats-escalate/

    • Transcript
  • #67
    August 3 · 6 min

    OpenAI’s Europe Governance Push, Amazon’s $50 Billion Stake, and a Copilot Prompt-Injection Worm | UpNext AI – August 3, 2026

    OpenAI outlines its approach to European AI governance as the EU AI Act moves forward, while the Financial Times reports that Amazon has completed a $50 billion equity investment in OpenAI. We also cover a demonstrated document-based prompt-injection attack on Copilot for Word, a new look at AI-discovered vulnerabilities, and a lightweight evaluation toolkit for builders. Covered stories: - OpenAI’s safety, transparency, provenance, and EU AI Act approach in Europe - Amazon’s reported $50 billion equity investment and roughly 5% stake in OpenAI - A demonstrated self-spreading prompt-injection worm targeting Microsoft Copilot for Word - VulnCheck’s count of AI-discovered vulnerabilities and their reported exploitation rate - Smevals, a compact toolkit for testing models, prompts, and agent harnesses - Research on brain-guided language models and robust reasoning Sources: - https://openai.com/index/advancing-responsible-ai-across-europe - https://www.ft.com/content/8ae9e6e4-a53c-44da-8e7d-c9d81f0df4b9?syn-25a6b1a6=1 - https://www.nature.com/articles/s42256-026-01278-w - https://the-decoder.com/a-security-researcher-built-a-self-spreading-worm-that-hides-inside-word-docs-and-hijacks-microsoft-copilot/ - https://the-decoder.com/ai-finds-plenty-of-security-flaws-but-almost-none-of-them-get-exploited/ - https://simonwillison.net/2026/Jul/31/smevals/#atom-everything

    • Transcript
Showing 1–20 of 31 episodes