Fetcher family
librarySync/async HTTP fetching.
Python scraping framework with adaptive parsers, HTTP/browser/stealth fetchers, concurrent spiders, CLI, and MCP/agent integrations.
Python scraping framework with adaptive parsers, HTTP/browser/stealth fetchers, concurrent spiders, CLI, and MCP/agent integrations.
python - <<'PY'
from scrapling.fetchers import Fetcher
page = Fetcher.get('https://example.com')
print(page.css('h1::text').get())
PYEach card below corresponds to a component identified in the repository research. Open a component record for its path, purpose, capabilities, dependencies, risks, and relationships.
Sync/async HTTP fetching.
Anti-bot-oriented fetching.
JS/browser-backed fetching.
Concurrent crawl/session/proxy workflows.
Expose scraping tools to agents.
These capabilities are derived from the current first-party README and supporting documentation. Names follow upstream terminology so you can search the source documentation precisely.
Upstream treats this as a distinct part of D4Vinci/Scrapling. In plain language, use this area when your task concerns spiders - a full crawling framework. Begin with the documented default, test it on disposable input, and open the source section for its current options and limitations.
Upstream treats this as a distinct part of D4Vinci/Scrapling. In plain language, use this area when your task concerns adaptive scraping & ai integration. Begin with the documented default, test it on disposable input, and open the source section for its current options and limitations.
Upstream treats this as a distinct part of D4Vinci/Scrapling. In plain language, use this area when your task concerns high-performance & battle-tested architecture. Begin with the documented default, test it on disposable input, and open the source section for its current options and limitations.
Upstream treats this as a distinct part of D4Vinci/Scrapling. In plain language, use this area when your task concerns developer/web scraper friendly experience. Begin with the documented default, test it on disposable input, and open the source section for its current options and limitations.
Upstream treats this as a distinct part of D4Vinci/Scrapling. In plain language, use this area when your task concerns getting started. Begin with the documented default, test it on disposable input, and open the source section for its current options and limitations.
Upstream treats this as a distinct part of D4Vinci/Scrapling. In plain language, use this area when your task concerns basic usage. Begin with the documented default, test it on disposable input, and open the source section for its current options and limitations.
Upstream treats this as a distinct part of D4Vinci/Scrapling. In plain language, use this area when your task concerns spiders. Begin with the documented default, test it on disposable input, and open the source section for its current options and limitations.
Upstream treats this as a distinct part of D4Vinci/Scrapling. In plain language, use this area when your task concerns advanced parsing & navigation. Begin with the documented default, test it on disposable input, and open the source section for its current options and limitations.
Upstream treats this as a distinct part of D4Vinci/Scrapling. In plain language, use this area when your task concerns cli & interactive shell. Begin with the documented default, test it on disposable input, and open the source section for its current options and limitations.
Upstream treats this as a distinct part of D4Vinci/Scrapling. In plain language, use this area when your task concerns performance benchmarks. Begin with the documented default, test it on disposable input, and open the source section for its current options and limitations.
Upstream treats this as a distinct part of D4Vinci/Scrapling. In plain language, use this area when your task concerns text extraction speed test (5000 nested elements). Begin with the documented default, test it on disposable input, and open the source section for its current options and limitations.
Upstream treats this as a distinct part of D4Vinci/Scrapling. In plain language, use this area when your task concerns element similarity & text search performance. Begin with the documented default, test it on disposable input, and open the source section for its current options and limitations.
Start with one capability card instead of asking the repository to do everything at once. This makes permissions, inputs, and output easier to understand.
Use a sample URL, copied repository, test document, or sandbox account. Keep production credentials and irreplaceable files out of the first run.
Watch Terminal output or the host application's activity view. If the behavior differs from the README, stop with Control + C or cancel inside the host before retrying.
Check generated files, diffs, API responses, logs, or previews. Never assume “command finished” means the result is correct.
Record the working command and non-secret configuration in your project's README. Store secrets in the documented environment file or password manager.
These examples are selected from the repository's first-party documentation. Replace obvious placeholders, keep quotation marks intact, and run them only in the context indicated by the surrounding explanation.
from scrapling.spiders import Spider, Request, Response
from scrapling.fetchers import FetcherSession, AsyncStealthySession
class MultiSessionSpider(Spider):
name = "multi"
start_urls = ["https://example.com/"]
def configure_sessions(self, manager):
manager.add("fast", FetcherSession(impersonate="chrome"))
manager.add("stealth", AsyncStealthySession(headless=True), lazy=True)
async def parse(self, response: Response):
for link in response.css('a::attr(href)').getall():
# Route protected pages through the stealth session
if "protected" in link:
yield Request(link, sid="stealth")
else:
yield Request(link, sid="fast", callback=self.parse) # explicit callbackfrom scrapling.spiders import Spider, Response
class MySpider(Spider):
name = "demo"
start_urls = ["https://example.com/"]
async def parse(self, response: Response):
for item in response.css('.product'):
yield {"title": item.css('h2::text').get()}
MySpider().start()from scrapling.spiders import ShopifySpider
class MyStore(ShopifySpider):
target_website = "example.com"
result = MyStore().start() # Every product in the store, one item per variantNone documented
A configuration file changes behavior without changing source code. Make one change at a time and keep a backup before editing JSON, TOML, YAML, or environment files.
None documented
Paths identify where the relevant implementation or generated files live. Paths beginning with ~ are inside your home folder.
None documented
Prefer temporary, least-privileged credentials. Never commit .env, tokens, cookies, private keys, or session exports.
None documented
Check pricing, data retention, rate limits, and account permissions before enabling optional integrations.
git diff before committing.127.0.0.1 unless you intentionally secure and expose them.The manual reflects first-party files checked on August 18, 2026. It explains the indexed repository rather than promising that every optional third-party integration is available or safe.