Skip to content
Artwork for The Real Python Podcast
The Real Python Podcast · July 24 · 58 min

Configuring a Versatile LLM Harness & Scraping the Web With Scrapy

Which is more important, the model or the “harness” around an LLM? What are ways to assemble an efficient agentic developer workflow? This week on the show, Ayan Pahwa joins us to discuss harnessing, web scraping, and self-hosting Python applications. Ayan is a developer advocate at Zyte and an experienced project builder. We discuss a recent article he wrote about creating an extension for the web scraping tool Scrapy. He also digs into his self-hosting setup for Python applications and tools. Our discussion extends to the complexities of developing effective harnesses. Ayan shares his setup and how he navigated shifting from prompt engineering to context and loop engineering. This episode is sponsored by HydraDB. Course Spotlight: Introduction to Web Scraping With Python In this video course, you’ll learn all about web scraping in Python. You’ll see how to parse data from websites and interact with HTML forms using tools such as Beautiful Soup and MechanicalSoup. Topics: 00:00:00 – Introduction 00:02:01 – Scrapy and building an extension 00:08:46 – Zyte and the web scraping API 00:11:19 – Sponsor: HydraDB 00:12:22 – noalgotube project 00:15:49 – Homelab & self hosting projects 00:22:19 – What goes into a harness? 00:32:33 – Where did you start exploring LLM tools? 00:36:13 – Local models & edge computing 00:39:31 – ExtractPod and discussing Apple’s AI 00:43:35 – Video Course Spotlight 00:44:54 – Managing token use and tools 00:52:37 – What are you excited about in the world of Python? 00:55:10 – What do you want to learn next? 00:56:38 – What is the best way to follow your work online? 00:56:58 – Thanks and goodbye Show Links: How to build your first Scrapy extension Web Scraping API - All-in-one Web Scraper - Zyte API Web Scraping With Scrapy and MongoDB – Real Python noalgotube: I Built My Own YouTube Feed Because the Algorithm Stopped Working for Me - CodeNSolder noalgotube: A personal content aggregator for YouTube channels and blog RSS feeds Why homelab? Building a proper self-hosted setup from scratch - CodeNSolder OPNsense: Open source, feature rich firewall and routing platform, offering cutting-edge network protection Proxmox - Powerful open-source server solutions Pi-hole – Network-wide Ad Blocking omni-tools: Self-hosted collection of powerful web-based tools for everyday tasks Frigate NVR Harness Engineering, part 1: What is an agent harness and why it matters My agentic coding setup: Claude Code, multi-agent orchestration, and how I actually work ExtractPod EP07 - AI Harnesses, our model usage and a Scottish dinner staple. - YouTube OpenCode - The open source AI coding agent OpenRouter Gemma 4 — Google DeepMind LM Studio Bionic - Agent for Open Models GLM-5.2: Built for Long-Horizon Tasks caveman: 🪨 why use many token when few token do trick — Claude Code skill ponytail: Makes your AI agent think like the laziest senior dev in the room Episode #301: Running Python Locally in a Sandbox Aillio – Bullet R2 - Coffee Roaster CodeNSolder Ayan Pahwa - Zyte HydraDB Level up your Python skills with our expert-led courses: Introduction to Web Scraping With Python Getting Started With Claude Code Testing MCP Servers With a Python MCP Client Support the podcast & join our community of Pythonistas

0:00-58:21

transcript

No transcript — this publisher did not publish one.

show notes

Which is more important, the model or the “harness” around an LLM? What are ways to assemble an efficient agentic developer workflow? This week on the show, Ayan Pahwa joins us to discuss harnessing, web scraping, and self-hosting Python applications.

Ayan is a developer advocate at Zyte and an experienced project builder. We discuss a recent article he wrote about creating an extension for the web scraping tool Scrapy. He also digs into his self-hosting setup for Python applications and tools.

Our discussion extends to the complexities of developing effective harnesses. Ayan shares his setup and how he navigated shifting from prompt engineering to context and loop engineering.

This episode is sponsored by HydraDB.

Course Spotlight: Introduction to Web Scraping With Python

In this video course, you’ll learn all about web scraping in Python. You’ll see how to parse data from websites and interact with HTML forms using tools such as Beautiful Soup and MechanicalSoup.

Topics:

  • 00:00:00 – Introduction
  • 00:02:01 – Scrapy and building an extension
  • 00:08:46 – Zyte and the web scraping API
  • 00:11:19 – Sponsor: HydraDB
  • 00:12:22 – noalgotube project
  • 00:15:49 – Homelab & self hosting projects
  • 00:22:19 – What goes into a harness?
  • 00:32:33 – Where did you start exploring LLM tools?
  • 00:36:13 – Local models & edge computing
  • 00:39:31 – ExtractPod and discussing Apple’s AI
  • 00:43:35 – Video Course Spotlight
  • 00:44:54 – Managing token use and tools
  • 00:52:37 – What are you excited about in the world of Python?
  • 00:55:10 – What do you want to learn next?
  • 00:56:38 – What is the best way to follow your work online?
  • 00:56:58 – Thanks and goodbye

Show Links:

Level up your Python skills with our expert-led courses:

Support the podcast & join our community of Pythonistas

links30