Overview


The Web connector crawls publicly accessible websites and ingests their content into PipesHub. Use it to bring product docs, public wikis, competitor sites, or reference material into the same searchable layer as internal knowledge.


How it works

PipesHub fetches a single page or recursively crawls from a root URL, following links found in HTML. It extracts page titles and cleaned text, and processes linked documents including PDFs, Office files, CSV, JSON, XML, and Markdown. URLs are normalized to avoid duplicates and content is stored with parent-child hierarchy. Pages behind logins, cookies, or private APIs are out of scope.


Configure
  • Create the connector under Workspace Settings and choose personal or team scope

  • Enter the root URL, this cannot be changed later

  • Set crawl type (single page or recursive), depth, page limit, and max file size

  • Apply domain restrictions and path filters, and restrict file extensions if needed

  • Enable manual approval or content-type indexing toggles, then save to start crawling

Web

Crawl public sites and documents, extend agent knowledge beyond internal sources.

GitHub

…k

GitHub stars in 12 weeks.

Discord

500+

Active developers on Discord.