Library · established

D4Vinci/Scrapling

Python scraping framework with adaptive parsers, HTTP/browser/stealth fetchers, concurrent spiders, CLI, and MCP/agent integrations.

web scrapingbrowser automationdata acquisition
Installation GuideInstruction Manual

What it contains

Python scraping framework with adaptive parsers, HTTP/browser/stealth fetchers, concurrent spiders, CLI, and MCP/agent integrations.

Typical uses

Build resilient scraping/crawling pipelines, including JS-heavy pages, while exposing web retrieval capabilities to agents.

Why people choose it

Inference: one Python stack spans lightweight requests, dynamic browsing, crawling, and adaptive selectors.

License
BSD 3-Clause
Maintenance
active
Latest release
Unknown
Last meaningful update
Unknown
Maturity
established
Production readiness
production

Classification

Subject domains

web scrapingbrowser automationdata acquisition

Task categories

web scrapingbrowser automationdata acquisition

Project phases

researchimplementation

Secondary types

frameworkCLIMCP server

Select when

  • Need Python scraping from HTTP through browser-backed crawling

Do not select when

  • A first-party API is preferable
  • Scraping is not authorized
Implementation complexity
medium
Setup effort
medium
Learning curve
medium
Confidence
high

Routing heuristic 88.0%

Routing specificity
4/5
Setup simplicity
3/5
Operational maturity
5/5
Composability
4/5
Local control
5/5
Security boundary clarity
5/5
First-party routing documentation
5/5

Strengths

None documented.

Poor-fit scenarios

  • A first-party API is preferable
  • Scraping is not authorized

Limitations

None documented.

README.md

Concurrent crawl/session/proxy workflows.

crawlingconcurrency
Languages
Python
Frameworks
None documented
Runtimes
Python
Operating systems
None documented
Installation methods
pip
Interfaces
library, CLI, MCP
Required credentials
None documented
External services
None documented
Hardware
None documented
Major dependencies
Playwright for browser modes

Compatibility notes

None documented.

One-sentence semantic summary

Python scraping framework with adaptive parsers, HTTP/browser/stealth fetchers, concurrent spiders, CLI, and MCP/agent integrations.

Capability keywords

adaptive parsingHTTP fetchbrowser automationstealth fetchcrawlingMCP

User intent phrases

  • Need Python scraping from HTTP through browser-backed crawling

Negative match phrases

  • A first-party API is preferable
  • Scraping is not authorized

Differentiators

None documented.

Security considerations

  • Review site authorization, terms, robots policies, proxy/cookie handling.

Privacy considerations

None documented.

Uncertainty

Unresolved items

None documented.

Inference notes

None documented.

Evidence

Repository usage guidance

installation summary
pip
basic usage summary
Build resilient scraping/crawling pipelines, including JS-heavy pages, while exposing web retrieval capabilities to agents.
documented entry points
None documented
key configuration files
None documented
important directories
None documented
documentation paths
README.md
example paths
None documented

Relationships

No explicit repository relationships indexed.

Recommended combinations

Not part of a curated combination.

Routing rules

rule_005: Need Python scraping from HTTP through browser-backed crawling P995

Rationale

Python scraping framework with adaptive parsers, HTTP/browser/stealth fetchers, concurrent spiders, CLI, and MCP/agent integrations.

Required conditions

None documented.

Preferred

D4Vinci/Scrapling

Fallback

None