Scrape a web URL and split its content into structured chunks for RAG pipelines. Ideal for summarizing or processing page content without field extraction.
Capture evidence of a URL hosting stolen or impersonating content. Creates an independent archive snapshot with retrieval timestamp, SHA-256 hash, and PDQ perceptual fingerprint. Run before content is removed.
Initiates a structured web crawl from a specified URL, following internal links to explore site content with configurable depth, breadth, and filtering options for targeted data extraction.
Starts a structured web crawl from a given URL, following internal links to discover and extract content. Control crawl depth, breadth, and focus on specific site sections.