
Overview
The Web connector crawls publicly accessible websites and ingests their content into PipesHub. Use it to bring product docs, public wikis, competitor sites, or reference material into the same searchable layer as internal knowledge.
How it works
PipesHub fetches a single page or recursively crawls from a root URL, following links found in HTML. It extracts page titles and cleaned text, and processes linked documents including PDFs, Office files, CSV, JSON, XML, and Markdown. URLs are normalized to avoid duplicates and content is stored with parent-child hierarchy. Pages behind logins, cookies, or private APIs are out of scope.
Configure
Create the connector under Workspace Settings and choose personal or team scope
Enter the root URL, this cannot be changed later
Set crawl type (single page or recursive), depth, page limit, and max file size
Apply domain restrictions and path filters, and restrict file extensions if needed
Enable manual approval or content-type indexing toggles, then save to start crawling

Web
Crawl public sites and documents, extend agent knowledge beyond internal sources.
You Might be Interested In
GitHub

GitHub stars in 12 weeks.
Discord

500+
Active developers on Discord.







