Extract public web data without bypassing access controls.
web scraping
Guides authorized web scraping with fallback extraction, access-failure handling, and safeguards for untrusted content.
When to use it
Use for social-media scraping, yt-dlp workflows, and handling CAPTCHA or HTTP 403 blocks.
It changes how the agent performs authorized web scraping, validates destinations, selects fallbacks, and handles denied access.
This skill
public web pages
Reads public web pages
robots.txt
Reads sites' robots.txt rules
browser network requests
Inspects browser network requests
YouTube
Reads YouTube metadata and media
The supplied scraping implementations are Python code.
The HTTP extraction strategies import and use requests.
The extraction strategies import BeautifulSoup from bs4.
The first extraction strategy uses trafilatura.