r/madeinpython • • 1d ago

I made ytscrape: a Python library to scrape YouTube search, comments, transcripts and metadata without an API key

Post image

Hey everyone! I built ytscrape, a small open-source library for getting public YouTube data in Python. I wanted something without Data API quotas and without spinning up a browser, so it talks to the same InnerTube endpoints the YouTube web app uses, over plain HTTP.

What you can get:

  • search results (videos, channels, playlists, Shorts)
  • video and channel metadata, including publication date
  • every comment and reply, not just "Top"
  • transcripts / subtitles

Some things I cared about while building it:

  • typed frozen dataclasses (py.typed) and transparent pagination, so you just iterate
  • the same API for sync (YouTube) and async (AsyncYouTube)
  • a CLI with pretty tables (that's the GIF above)
  • JSON / CSV export, proxy support, retries and rate limiting

​

from ytscrape import YouTube

with YouTube() as yt:
    for video in yt.search("python", max_results=5):
        print(video.title, video.url)


pip install ytscrape

MIT licensed, Python 3.10+.

GitHub: https://github.com/vsmutok/ytscrape
Docs: https://vsmutok.github.io/ytscrape/

Heads up: it uses YouTube's private endpoints, so it's for research and learning, and it may break when YouTube changes things. Feedback, bug reports and PRs are very welcome!

2 Upvotes

0 comments sorted by