The web data API to search, scrape, and interact at scale. 🔥
-
Updated
Sep 21, 2026 - TypeScript
The web data API to search, scrape, and interact at scale. 🔥
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ
Python scraper based on AI
The SDK to extract data and interact with any site on the web. Get started with Claude Code, Codex, Eve, Mastra, and more.
Turn any website into a structured API. Extract, automate, search and monitor the web.
Declarative data automation language and Go runtime for structured extraction workflows.
Browser automation CLI built for AI agents. Break through anti-bot walls, hand off to humans across platforms when stuck. Parallel multi-task execution, independent multi-session operation, isolated multi-account browsing.
Extract Keywords from sentence or Replace keywords in sentences.
A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.
The undetected self-hosted browser automation platform. Powered by Camoufox (Firefox) for 0% detection rates. Built for speed, privacy, and scalability.
Python package for scraping recipes data
ContextGem: Effortless LLM extraction from documents
To extract article from given URL
Converts a pdf file into a text file while keeping the layout of the original pdf. Useful to extract the content from a table in a pdf file for instance. This is a subclass of PDFTextStripper class (from the Apache PDFBox library).
🚚 Agile Data Preparation Workflows made easy with Pandas, Dask, cuDF, Dask-cuDF, Vaex and PySpark
A beginner-friendly yet powerful Python toolkit for financial analysis and automation — built to make modern investing accessible to everyone
Lightweight library for scraping web-sites with LLMs
⛏️ The extraction engine behind Maigret: turn any profile URL into a structured OSINT record across 150+ sites
Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
To associate your repository with the data-extraction topic, visit your repo's landing page and select "manage topics."