0:00-22:38
Streams straight from the publisher. podnod never proxies or re-hosts episode audio.
Video annotation is an expensive and time-consuming process. As a consequence, the available video datasets are useful but small. The availability of machine transcribed explainer videos offers a unique opportunity to rapidly develop a useful, if dirty, corpus of videos that are "self annotating", as hosts explain the actions they are taking on the screen.
This episode is a discussion of the HowTo100m dataset - a project which has assembled a video corpus of 136M video clips with captions covering 23k activities.
Related LinksThe paper will be presented at ICCV 2019
HowTo100m
di.ens.frICCV 2019
iccv2019.thecvf.com@antoine77340
twitter.comAntoine on Github
github.comAntoine's homepage
di.ens.fr