crawler module · apify/crawlee-python
BeautifulSoupCrawler
HTTP-based crawler that parses HTML with Beautiful Soup for efficient non-browser extraction.
Location
Repository path
README.mdInvocation
`from crawlee.crawlers import BeautifulSoupCrawler` then `await crawler.run([...])`
Setup
Installation / activation
Install the `beautifulsoup` extra and run a configured crawler.
Keywords
beautifulsouphttphtmlparsercrawler
Repository context
apify/crawlee-python
Apache-2.0 Python 3.10+ library and CLI for building asynchronous HTTP and browser-based crawlers with routing, retries, proxy and session management, persistent queues, pluggable storage, templates, and optional parsing, Playwright, database, Redis, observability, and AI-related integrations; no native Claude Code integration is evidenced.
web scrapingweb crawlingbrowser automationdata extractiondata collectionHTTP automationAI data preparationRAG ingestiondeveloper tooling