Skip to main content

Crawling Manager API

The Crawling Manager API provides endpoints for scheduling, managing, and monitoring data crawling jobs across various connector types including Google Workspace, OneDrive, SharePoint Online, Slack, and Confluence.

Base URL

All endpoints are prefixed with /api/v1/crawlingManager

Authentication

All endpoints require authentication via Authorization header with valid JWT token. In addition:
  • Endpoints for one connector check that the caller may manage it: a team connector’s schedule only by an admin, and a personal connector’s only by the person who created it
  • Admin privileges for DELETE /schedule/all, because it removes the schedules of every connector in the organization, including other members’ personal connectors
Headers:

Path Parameters

Endpoints for one connector take two path parameters:
  • connector — the connector’s type name (for example drive, onedrive or jira). The crawling manager accepts any connector type; it doesn’t keep its own list.
  • connectorId — the ID of the connector instance. The request is checked against this instance: it must exist (otherwise 404), and the caller must be allowed to manage it (otherwise 403).
Internally, each connector’s schedule is keyed as crawl-{connector}-{connectorId}-{orgId}, with the connector name lowercased and spaces replaced by hyphens. That key is the job ID of a one-time (once) schedule and the id reported for a paused schedule. A repeating schedule’s job ID is the one the job queue (BullMQ) assigns.

API Endpoints

Schedule a new crawling job for a specific connector type.Endpoint: POST /api/v1/crawlingManager/:connector/:connectorId/scheduleParameters:
  • connector (string, path) - The connector type
  • connectorId (string, path) - The connector instance ID
Schedule Configuration Types:All schedule configurations inherit base properties:
  • scheduleType (required) - The type of schedule
  • isEnabled (boolean, default: true) - Whether the schedule is enabled
  • timezone (string, default: “UTC”) - Timezone for schedule execution
Hourly Schedule:
Daily Schedule:
Weekly Schedule:
Monthly Schedule:
Custom Schedule (Cron):
Once Schedule:
Interval Schedule: the interval goes inside a nested scheduleConfig object.
Retrieve the status of a scheduled crawling job for a specific connector.Endpoint: GET /api/v1/crawlingManager/:connector/:connectorId/scheduleParameters:
  • connector (string, path) - The connector type
  • connectorId (string, path) - The connector instance ID
Status: 200 OKA scheduled, repeating job that is running now:
  • id is the ID the job queue (BullMQ) gave this run; repeating runs look like repeat:<key>:<time in ms>. A one-time (once) job’s id is crawl-{connector}-{connectorId}-{orgId}, and so is a paused schedule’s.
  • name is crawl-{connector}-{connectorId}.
  • Fields with no value, such as finishedOn or failedReason on a job that hasn’t finished, are left out.
A paused schedule is reported from memory, with "state": "paused", "progress": 0, "attemptsMade": 0, and timestamp set to when it was paused.Job States:
  • waiting or prioritized - Queued, waiting for a worker
  • delayed - Due later (the next repeating run, or a one-time run’s scheduled time)
  • active - Being processed now
  • failed - The last run failed after its retries
  • paused - Paused through this API
Runs that completed successfully (completed) aren’t reported. Failed runs are, for as long as the queue keeps them (see Job history retention below), so a schedule that was removed, or a one-time sync that failed, can still be reported with "state": "failed". The endpoint answers 404 “No scheduled job found for this connector” only when the connector has no paused schedule and no waiting, prioritized, delayed, active or failed job.Progress Tracking: Jobs report progress at key stages: 10% (start), 20% (task service obtained), 100% (completion). Failed jobs may show partial progress.
Retrieve all scheduled crawling jobs for the organization.Endpoint: GET /api/v1/crawlingManager/schedule/all
Status: 200 OKThe list holds, for each connector type (the connector value, such as drive, across all of that type’s instances), up to its 10 most recent queued, delayed, active or failed jobs, followed by every paused schedule in the organization. Runs that completed successfully aren’t listed; failed ones are, including those of a schedule that has since been removed. Each entry has the same fields as Get Job Status; this example shortens scheduleConfig.
Remove a scheduled crawling job for a specific connector.Endpoint: DELETE /api/v1/crawlingManager/:connector/:connectorId/removeParameters:
  • connector (string, path) - The connector type
  • connectorId (string, path) - The connector instance ID
Status: 200 OK
Remove all scheduled crawling jobs for the organization. Admins only: any other caller is refused with 400 “Admin access required”, and nothing is removed.Endpoint: DELETE /api/v1/crawlingManager/schedule/all
Status: 200 OK
Pause a scheduled crawling job for a specific connector.Endpoint: POST /api/v1/crawlingManager/:connector/:connectorId/pauseParameters:
  • connector (string, path) - The connector type
  • connectorId (string, path) - The connector instance ID
Status: 200 OK
Resume a paused crawling job for a specific connector.Endpoint: POST /api/v1/crawlingManager/:connector/:connectorId/resumeParameters:
  • connector (string, path) - The connector type
  • connectorId (string, path) - The connector instance ID
Status: 200 OK
Retrieve statistics about the crawling job queue.Endpoint: GET /api/v1/crawlingManager/stats
Status: 200 OK
Statistics Fields:
  • waiting - Number of jobs waiting to be processed
  • active - Number of jobs currently being processed
  • completed - Number of completed jobs
  • failed - Number of failed jobs
  • delayed - Number of delayed jobs
  • paused - Number of paused jobs
  • repeatable - Number of repeatable/scheduled jobs
  • total - Total number of jobs across all states

Data Types

System Configuration

The API includes built-in rate limiting and concurrency controls:
  • Maximum 5 concurrent jobs per queue
  • Jobs are automatically retried with exponential backoff (5000ms initial delay) on failure
  • Stalled jobs are detected after 30 seconds
  • Maximum 3 retry attempts for failed jobs
  • Job history retention: when a run finishes, the queue keeps at most the 10 most recent completed runs and the 10 most recent failed runs, counted across the whole queue rather than per connector. Removing or replacing a connector’s schedule also trims that connector instance’s finished runs to its 10 most recent. Completed runs are kept in the queue and counted by GET /stats, but Get Job Status and Get All leave them out; failed runs appear in all three
  • Jobs are removed and recreated when updating schedules (no job modification). The new schedule’s time, cron expression and timezone are checked first, so a schedule refused for those reasons never removes the one already in place. A schedule sent with isEnabled: false still clears the current one
  • Removing or pausing a schedule also cancels its runs that are waiting to start, including one-time runs. If the job queue can’t be reached, the request fails with 500 instead of reporting success, and a schedule is marked paused only after its runs are removed

Error Handling

Errors raised by these endpoints use the API’s standard error shape:
code is HTTP_ followed by the status name (HTTP_BAD_REQUEST, HTTP_UNAUTHORIZED, HTTP_FORBIDDEN, HTTP_NOT_FOUND), or VALIDATION_ERROR when the request doesn’t match the schema. A 500 comes in two kinds. When the job queue can’t be reached, the code is INTERNAL_ERROR with the message “Something went wrong on PipesHub’s side. Please try again; if it keeps happening, ask your admin for help.” When the queue answered but removing a schedule or a queued run failed, the code is HTTP_INTERNAL_SERVER_ERROR with a specific message that says what may still be in place; each endpoint’s tab lists them. requestId is included when the request has one. The only exception is the 404 for a connector with no scheduled job, shown under Get Job Status. Common HTTP Status Codes:
  • 200 - Success
  • 201 - Created (for scheduling jobs)
  • 400 - Bad Request (validation errors, invalid configuration, or a non-admin calling Remove All)
  • 401 - Unauthorized (missing or invalid authentication)
  • 403 - Forbidden (not allowed to manage this connector, or an OAuth token missing the scope)
  • 404 - Not Found (connector instance not found, or no scheduled job)
  • 500 - Internal Server Error