Website Scraping
Searj can read your website content and use it to train your chatbot. This means the bot will be able to answer customer questions based on the information on your website.
What is Website Scraping?
Website scraping is the process of automatically reading text content from your website pages. Searj visits the pages you specify, extracts the important text, then processes it for bot training — just like it does with PDF files.
💡 This feature is perfect if you have a website with information about your services, products, or FAQs.
Adding a URL
Follow these steps to add a website:
From the bot page, open the "Documents" tab.
Click the "Add URL" button to open the add dialog.
Enter the page URL you want to scrape, e.g., https://example.com
Choose the depth level: 1, 2, or 3. See the next section for details.
Searj will begin scraping the content immediately.
Scraping Depth Explained
Scraping depth determines how many levels of links Searj will follow starting from the original URL. This is one of the most important settings when adding a website:
Depth 1 — Single Page Only
Only scrapes the exact page you provided the URL for. Ideal for a single page like an "About Us" page or a specific service page.
Depth 2 — Page + Its Links
Scrapes the original page plus all pages linked from it. Great for a complete section of your site, like a services page and all individual service details.
Depth 3 — Three Levels Deep
Scrapes three levels: the original page, its links, and links found on those pages. Best for scraping an entire website.
| Depth | What Gets Scraped | Best For |
|---|---|---|
| 1 | The specified page only | A single specific page |
| 2 | The page + all links on it | A section of your site |
| 3 | Page + its links + links on those pages | Entire website |
💡 Tip: Always start with depth 1 to test the results, then increase depth if you need more content.
Plan Limits for Scraping
Each plan has different limits for scraping depth and page count:
| Plan | Max Depth | Max Pages |
|---|---|---|
| Free | 1 | 10 |
| Starter | 2 | 30 |
| Pro | 3 | 50 |
| Business | 3 | 100 |
| Enterprise | 3 | 200 |
Best Practices
- Start with depth 1 — Test results first before increasing depth to ensure the quality of extracted content.
- Use informative pages — Your FAQ page and About page are excellent starting points.
- Arabic websites work perfectly — Searj supports scraping Arabic content without any issues.
- Some websites may block scraping — If no content is extracted, contact our support team.
Troubleshooting
Empty Content
If the extracted content is empty, the website may use JavaScript to load content dynamically. Try adding a different internal page.
Timeout
If scraping takes too long and fails, try reducing the depth or using a simpler page.
Website Blocks Access
Some websites block automated access. In this case, you can manually copy the content and use the "Manual Text" option instead.
⚠️ Please ensure you have the right to scrape the website content. Respect each website's terms of service and do not scrape content you don't have permission to use. Searj is designed for scraping your own website or publicly available content.