
Building A Community And Technology Stack For Scalable Big Data Geoscience At Pangeo
Streams straight from the publisher. podnod never proxies or re-hosts episode audio.
Summary
Science is founded on the collection and analysis of data. For disciplines that rely on data about the earth the ability to simulate and generate that data has been growing faster than the tools for analysis of that data can keep up with. In order to help scale that capacity for everyone working in geosciences the Pangeo project compiled a reference stack that combines powerful tools into an out-of-the-box solution for researchers to be productive in short order. In this episode Ryan Abernathy and Joe Hamman explain what the Pangeo project really is, how they have integrated a combination of XArray, Dask, and Jupyter to power these analytical workflows, and how it has helped to accelerate research on multidimensional geospatial datasets.
Announcements
- Hello and welcome to Podcast.__init__, the podcast about Python’s role in data and science.
- When you’re ready to launch your next app or want to try a project you hear about on the show, you’ll need somewhere to deploy it, so take a look at our friends over at Linode. With the launch of their managed Kubernetes platform it’s easy to get started with the next generation of deployment and scaling, powered by the battle tested Linode platform, including simple pricing, node balancers, 40Gbit networking, dedicated CPU and GPU instances, and worldwide data centers. Go to pythonpodcast.com/linode and get a $100 credit to try out a Kubernetes cluster of your own. And don’t forget to thank them for their continued support of this show!
- So now your modern data stack is set up. How is everyone going to find the data they need, and understand it? Select Star is a data discovery platform that automatically analyzes & documents your data. For every table in Select Star, you can find out where the data originated, which dashboards are built on top of it, who’s using it in the company, and how they’re using it, all the way down to the SQL queries. Best of all, it’s simple to set up, and easy for both engineering and operations teams to use. With Select Star’s data catalog, a single source of truth for your data is built in minutes, even across thousands of datasets. Try it out for free and double the length of your free trial today at pythonpodcast.com/selectstar. You’ll also get a swag package when you continue on a paid plan.
- Your host as usual is Tobias Macey and today I’m interviewing Ryan Abernathy and Joe Hamman about Pangeo, a community platform for Big Data geoscience
Interview
- Introductions
- How did you get introduced to Python?
- Can you describe what Pangeo is and the story behind it?
- What is your role in the project/community and how did you get involved?
- What are the goals of the project and community?
- What are the areas of effort and how are they organized?
- What are the scientific domains that Pangeo is focused on supporting?
- What are the primary challenges associated with data management and analysis in these scientific communities?
- What are the forms that these data take and how have they been evolving? (e.g. formats/sources)
- What are some of the challenges introduced by the widespread adoption of cloud resources and the associated architectural patterns?
- Can you describe the technical components that fall under the Pangeo umbrella?
- How do they come together to form a functional workflow for geo sciences?
- How has the scope of the Pangeo project changed or evolved since it started?
- What are the most interesting, innovative, or unexpected ways that you have seen Pangeo used?
- What are the most interesting, unexpected, or challenging lessons that you have learned while working on Pangeo?
- When is Pangeo the wrong choice?
- What do you have planned for the future of Pangeo?
Keep In Touch
- Joe
- @HammanHydro on Twitter
- Ryan
Picks
- Tobias
- Ryan
- Klara And The Sun by Kazuo Ishiguro
- Joe
- Range by David Epstein
Closing Announcements
- Thank you for listening! Don’t forget to check out our other show, the Data Engineering Podcast for the latest on modern data management.
- Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
- If you’ve learned something or tried out a project from the show then tell us about it! Email hosts@podcastinit.com) with your story.
- To help other people find the show please leave a review on iTunes and tell your friends and co-workers
Links
- Pangeo
- Pangeo Forge
- CarbonPlan
- M2LInES
- LEAP
- Columbia University
- XArray
- MIT
- MatLab
- PHP
- Ruby
- Java
- NumPy
- SciPy
- Matplotlib
- C
- Fortran
- Perl
- Dask
- Jupyter
- IDL
- HDF5
- Unidata
- NetCDF
- CF Metadata Conventions
- Intake
- FSSpec
- Parquet
- Zarr
- Data Engineering Podcast
- Pangeo Forge
- Airbyte
- Fivetran
- Stitch
- TileDB
- Pythia
The intro and outro music is from Requiem for a Fish The Freak Fandango Orchestra / CC BY-SA
pythonpodcast.com/linode
pythonpodcast.compythonpodcast.com/selectstar
pythonpodcast.com@HammanHydro
twitter.com@rabernat
twitter.comrabernat
github.comWebsite
ocean-transport.github.ioMountain Biking
rei.comKlara And The Sun
amazon.comRange
davidepstein.comData Engineering Podcast
dataengineeringpodcast.comsite
pythonpodcast.comiTunes
itunes.apple.comPangeo
pangeo.ioPangeo Forge
pangeo-forge.orgCarbonPlan
carbonplan.orgM2LInES
m2lines.github.ioLEAP
leap.columbia.eduColumbia University
columbia.eduXArray
xarray.devMIT
web.mit.eduMatLab
mathworks.comPHP
php.netRuby
ruby-lang.orgJava
en.wikipedia.orgNumPy
numpy.orgSciPy
scipy.orgMatplotlib
matplotlib.orgC
en.wikipedia.orgFortran
fortran-lang.orgPerl
perl.orgDask
dask.orgData Engineering Podcast Episode
dataengineeringpodcast.comJupyter
jupyter.orgIDL
en.wikipedia.orgHDF5
hdfgroup.orgUnidata
unidata.ucar.eduNetCDF
unidata.ucar.eduCF Metadata Conventions
cfconventions.orgIntake
intake.readthedocs.ioPodcast Episode
pythonpodcast.comFSSpec
filesystem-spec.readthedocs.ioParquet
parquet.apache.orgData Engineering Podcast Episode
dataengineeringpodcast.comZarr
zarr.readthedocs.ioData Engineering Podcast
dataengineeringpodcast.comPangeo Forge
pangeo-forge.readthedocs.ioAirbyte
airbyte.comData Engineering Podcast Episode
dataengineeringpodcast.comData Engineering Podcast Episode
dataengineeringpodcast.comStitch
stitchdata.comTileDB
tiledb.comData Engineering Podcast Episode
dataengineeringpodcast.comPythia
projectpythia.orgThe Freak Fandango Orchestra
freemusicarchive.orgCC BY-SA
creativecommons.org
- 1:45Introduction to Ryan Abernathy and Joe Hammond
- 2:57Ryan's Journey with Python
- 4:11Joe's Introduction to Python
- 5:18Overview of the Pangio Project
- 8:21The Birth of Pangio
- 10:35Challenges and Community Building
- 14:41Core Elements and Extensions of the Pangio Stack
- 19:10Scientific Domains Using Pangio
- 24:05Managing Source Formats and Data Abstraction
- 27:29Cloud Computing and Its Impact on Geosciences
- 33:54Modern Data Stack and Scientific Data
- 36:14Exploring Cloud Native Storage Formats
- 38:37Evolution and Future Goals of Pangio
- 42:06Innovative Uses and Success Stories
- 44:00Lessons Learned in Community Building
- 45:47Future Plans for Pangio
- 48:17Community Resources and Education
- 50:03Picks and Recommendations