698 servers ยท Data & APIs ยท Web Scraping & Crawling
Web Scraping & Crawling servers from Data & APIs, aggregated from every major registry, trust-ranked, and ready to install.
Enables web browsing capabilities through tools for content extraction, link following, and browser automation with customizable parameters for scraping, data collection, and web crawling tasks.
Public Spotify metadata, lyrics, and podcast data for LLM agents with no API key required.
Integrates with the Serper API to enable web searches and webpage content extraction, supporting research, content aggregation, and data mining tasks.
Integrates with Oxylabs web scraping services to extract, clean, and structure web content for real-time data analysis and monitoring workflows.
Extracts structured data from web pages based on natural language descriptions, converting website content into JSON format without custom scraping code.
Provides specialized investment research tools for analyzing SEC filings, earnings calls, financial data, stock market information, private company details, funding rounds, M&A transactions, and web scraping capabilities.
Provides a unified search and web scraping platform that integrates multiple search providers like SearxNG and Tavily, along with Firecrawl for advanced web content extraction, enabling flexible web data retrieval and structured information gathering.
Integrates with Octagon API to provide multi-source data aggregation, web scraping, academic research synthesis, competitive analysis, market intelligence, technical analysis, policy research, and trend analysis for professional-grade research capabilities.
AI-powered web scraping and structured data extraction through the ScrapeGraph API.
Enables LLMs to make advanced HTTP requests with realistic browser emulation, bypassing anti-bot measures while supporting all HTTP methods, authentication, and automatic response handling for web scraping and API interactions.
Integrates with Olostep's web scraping API to extract webpage content in markdown format, discover website URLs through search queries, and retrieve structured Google search results with country-specific routing and JavaScript rendering support.
Integrates with the Scrapezy API to extract structured data from websites based on user-specified prompts, enabling flexible web scraping for data collection, content aggregation, and automated research tasks.
Integrates with Cloudflare's Browser Rendering API to enable web scraping and screenshot capture using Puppeteer for dynamic content processing and automated visual testing.
Integrates web crawling with RAG functionality to enable website content retrieval, storage in vector databases, and semantic searching over crawled data for enhanced knowledge access
Fast, local-first web content extraction for LLMs. Scrape, crawl, extract structured data โ all from Rust. CLI, REST API, and MCP server.
Extracts LinkedIn profile data, company information, and connection details using advanced anti-detection web scraping techniques for recruitment automation, lead generation, and professional network analysis.
Website cloning engine with intelligent crawling, asset downloading, PDF generation, authentication support, and dynamic content rendering for website archival, offline browsing, and data extraction workflows
Provides a bridge to Dumpling AI's data extraction API for performing web searches, scraping content, extracting structured data, and processing various document formats through 20+ specialized tools.
Scrapes web pages and extracts structured data via the Thunderbit Open API with seven tools including batch processing of up to 100 URLs.
Enables comprehensive web research by leveraging Tavily's Search and Crawl APIs to aggregate information from multiple sources, extract detailed content, and structure data specifically for generating technical documentation and research reports.
Provides search access to Prisma Cloud documentation by crawling and indexing pages from both main docs and API documentation, implementing caching with TTL expiration and relevance scoring to return structured results with snippets and URLs for quick documentation access.
Enables Claude to execute XPath queries on XML and HTML content, supporting both direct parsing and web scraping through Puppeteer for structured data extraction from documents and websites.
Manages multiple Tavily API keys through intelligent load balancing and rotation to prevent rate limiting while providing reliable access to web search, content extraction, crawling, and mapping capabilities.
MCP server for OpenGraph.io API providing link unfurling, screenshots, HTML scraping, and Open Graph metadata extraction.
Access LinkedIn, Instagram, Twitter, and Reddit data through one API. No scraping. No maintenance.
Web scraping server that extracts structured data and screenshots from any site with anti-bot bypass.
Documentation crawler service that fetches and serves technical documentation from specified sources, enabling access to up-to-date information through a structured API.
Advanced search and retrieval for web crawler data. Supports WARC, wget, Katana, SiteOne, and InterroBot crawlers.
Enables AI to search and analyze content from Xiaohongshu (Little Red Book) social media platform through web scraping, providing access to comments and search results for market research and trend analysis in the Chinese consumer market.
Licensed creator content for AI agents. Discover music, video, and text catalogs as pre-indexed Pockets โ a few hundred tokens instead of a 10,000-token scrape, delivered in under 30ms. Free discovery; pulls are licensed per-use via HTTP 402 paywall โ connection is the contract, and 85% of every pull pays the rights holder instantly. **Tools:** `list_pockets` ยท `search_pockets` ยท `pull_content` No key needed to browse. Pulling content returns a 402 with signup instructions โ register at the URL provided, attach your API key as a Bearer token, and pull licensed content with full attribution and audit trail.
Provides automated access to Brazilian legal precedents from the Supreme Federal Court, Superior Court of Justice, and Superior Labor Court using web scraping to handle dynamic content and authentication for legal research and case law analysis.
Google Search, web scraping, and multi-source research tools.
Provides access to 120+ AI models and live web data โ scraping, crawling, and structured extraction โ through a single API key.
Fetch Browser enables headless web content retrieval and Google searching without API keys, supporting multiple output formats for web scraping and content analysis tasks.
Provides a bridge to the Status Invest platform for accessing Brazilian stock market data, including payment dates, stock information, and financial indicators through an API that scrapes the website.
Provides conversational access to Palo Alto Networks Cortex Cloud platform documentation through web scraping and intelligent indexing with automatic caching, relevance scoring, and separate tools for general documentation versus API-specific content.
Read any page past its anti-bot wall, crawl sites, search without an API key, solve captchas locally
Scrape, clean and match product and price data from any website
One API for 65 platforms and 572 endpoints: social, commerce, retail, jobs, finance, web scraping.
Scraping API for Google Search, Maps, Flights, Jobs, YouTube, LinkedIn, Trustpilot and more.
Extract data from any website with thousands of scrapers, crawlers, and automations on Apify Store โก
Cloud scraping & crawling API for AI agents. Turn any URL into clean, LLM-ready markdown.
Scrape any web page, search Google/Amazon/Walmart/YouTube and extract data via the ScrapingBee API.
One key to 1,000+ paid data APIs: enrichment, SEO/SERP, scraping, places, news. Pay per call.
Customer-voice research: scrape reviews and comments from 30+ sites, find ads, hooks, creators.
Instagram: Instagram public data scraper API for search, users, posts, hashtags, locations and more.
Live web search, image search, topic filters and full-text fetch over our own crawled index.
SEC filings, financials, insider trades; company data; web scraping; OFAC, email; HVAC calcs. x402
23 API tools for Claude Code โ email, DNS, SSL, NLP, scraping. Free tier included.
Web search, scraping, Google Trends and data lookups. Paid per call in USDC on Base via x402.
FAQ
698 Web Scraping & Crawling MCP servers are indexed on Lulu MCPs, ranked by trust score and aggregated from the official MCP Registry, Glama, PulseMCP and Smithery. Every listing links back to its source registry and installs in one click for Claude Code, Cursor and VS Code.
Web Scraping & Crawling is a sub-filter within Lulu MCPs' Data & APIs category, matched against a validated term set specific to web scraping & crawling. Servers can appear under more than one sub-filter when their functionality spans multiple areas.