
Orchestrating local LLMs for sensitive data with Airflow
transcript
show notes
Running LLMs over sensitive data means external APIs are off the table, but open-weight models still need reproducibility, validation, and a paper trail before downstream systems can depend on their output. In this episode, Kenten is joined by Chhayank Jain, Senior ML Engineer at Visa, to talk through how Airflow orchestrates local LLM pipelines end to end: nightly inference over millions of records, silent failure handling with Pydantic and distribution checks, and using a separate DAG to manage LoRA fine-tuning cycles.
Key Takeaways:
00:00 Introduction.
01:47 From modeling to pipelines. Chhayank shares how leading a readmission risk model across hospitals with different EHR systems turned him into a "pipeline person."
03:14 Day-to-day as an LLM engineer. Serving with ONNX Runtime, failure design, and agentic tooling with LangGraph. The model is only about one-fifth of the work.
05:03 The sensitive data problem. Why hosted LLM APIs are off the table for PHI and payments data, and why the real challenge is running Llama 3 nightly with enough checks that downstream models can trust the output.
07:34 Where Airflow fits. Sensors per site, dependency ordering, retries and backfills, and run records as the "paper trail" for audits.
10:03 DAG structure. Four stages: ingest and normalize, inference on GPU pods via the KubernetesPodOperator, validation, and load to the feature store.
12:32 Loud vs silent failures. OOM kills and timeouts are easy. Silent failures where the output looks clean but is wrong are the scary category, handled with Pydantic schemas, cross-checks against structured records, and distribution monitoring.
16:45 LoRA fine-tuning in its own DAG. A separate monthly or quarterly DAG pulls labeled data, fine-tunes with Hugging Face PEFT on top of Llama 3, and only writes a new adapter if it beats the frozen holdout.
20:33 Hardest production lessons. Pin everything (model version, temperature, features, seed), match parallelism to GPU capacity, and invest in traceability so you can answer "which adapter produced this row" six months later.
23:00 Airflow wishlist. First-class GPU awareness so teams stop reinventing it through pools and queues, and treating model versions as Airflow assets.
Resources Mentioned:
- Llama 3
- Pydantic
Thanks for listening to "The Data Flowcast: Mastering Apache Airflow® for Data Engineering and AI." If you enjoyed this episode, please leave a 5-star review to help get the word out about the show. And be sure to subscribe so you never miss any of the insightful conversations.
#AI #Orchestration #Airflow





