Crawling Manager API
The Crawling Manager API provides endpoints for scheduling, managing, and monitoring data crawling jobs across various connector types including Google Workspace, OneDrive, SharePoint Online, Slack, and Confluence.Base URL
All endpoints are prefixed with/api/v1/crawlingManager
Authentication
All endpoints require authentication viaAuthorization header with valid JWT token. In addition:
- Endpoints for one connector check that the caller may manage it: a team connector’s schedule only by an admin, and a personal connector’s only by the person who created it
- Admin privileges for
DELETE /schedule/all, because it removes the schedules of every connector in the organization, including other members’ personal connectors
Path Parameters
Endpoints for one connector take two path parameters:connector— the connector’s type name (for exampledrive,onedriveorjira). The crawling manager accepts any connector type; it doesn’t keep its own list.connectorId— the ID of the connector instance. The request is checked against this instance: it must exist (otherwise404), and the caller must be allowed to manage it (otherwise403).
crawl-{connector}-{connectorId}-{orgId}, with the connector name lowercased and spaces replaced by hyphens. That key is the job ID of a one-time (once) schedule and the id reported for a paused schedule. A repeating schedule’s job ID is the one the job queue (BullMQ) assigns.
API Endpoints
POST /:connector/:connectorId/schedule - Schedule Job
POST /:connector/:connectorId/schedule - Schedule Job
Schedule a new crawling job for a specific connector type.Endpoint: Schedule Configuration Types:All schedule configurations inherit base properties:Daily Schedule:Weekly Schedule:Monthly Schedule:Custom Schedule (Cron):Once Schedule:Interval Schedule: the interval goes inside a nested
POST /api/v1/crawlingManager/:connector/:connectorId/scheduleParameters:connector(string, path) - The connector typeconnectorId(string, path) - The connector instance ID
- Request Body
- Success Response
- Error Responses
scheduleType(required) - The type of scheduleisEnabled(boolean, default: true) - Whether the schedule is enabledtimezone(string, default: “UTC”) - Timezone for schedule execution
scheduleConfig object.GET /:connector/:connectorId/schedule - Get Job Status
GET /:connector/:connectorId/schedule - Get Job Status
Retrieve the status of a scheduled crawling job for a specific connector.Endpoint:
GET /api/v1/crawlingManager/:connector/:connectorId/scheduleParameters:connector(string, path) - The connector typeconnectorId(string, path) - The connector instance ID
- Success Response
- Error Responses
Status:
200 OKA scheduled, repeating job that is running now:idis the ID the job queue (BullMQ) gave this run; repeating runs look likerepeat:<key>:<time in ms>. A one-time (once) job’sidiscrawl-{connector}-{connectorId}-{orgId}, and so is a paused schedule’s.nameiscrawl-{connector}-{connectorId}.- Fields with no value, such as
finishedOnorfailedReasonon a job that hasn’t finished, are left out.
"state": "paused", "progress": 0, "attemptsMade": 0, and timestamp set to when it was paused.Job States:waitingorprioritized- Queued, waiting for a workerdelayed- Due later (the next repeating run, or a one-time run’s scheduled time)active- Being processed nowfailed- The last run failed after its retriespaused- Paused through this API
completed) aren’t reported. Failed runs are, for as long as the queue keeps them (see Job history retention below), so a schedule that was removed, or a one-time sync that failed, can still be reported with "state": "failed". The endpoint answers 404 “No scheduled job found for this connector” only when the connector has no paused schedule and no waiting, prioritized, delayed, active or failed job.Progress Tracking: Jobs report progress at key stages: 10% (start), 20% (task service obtained), 100% (completion). Failed jobs may show partial progress.GET /schedule/all - Get All Job Statuses
GET /schedule/all - Get All Job Statuses
Retrieve all scheduled crawling jobs for the organization.Endpoint:
GET /api/v1/crawlingManager/schedule/all- Success Response
- Error Responses
Status:
200 OKThe list holds, for each connector type (the connector value, such as drive, across all of that type’s instances), up to its 10 most recent queued, delayed, active or failed jobs, followed by every paused schedule in the organization. Runs that completed successfully aren’t listed; failed ones are, including those of a schedule that has since been removed. Each entry has the same fields as Get Job Status; this example shortens scheduleConfig.DELETE /:connector/:connectorId/remove - Remove Job
DELETE /:connector/:connectorId/remove - Remove Job
Remove a scheduled crawling job for a specific connector.Endpoint:
DELETE /api/v1/crawlingManager/:connector/:connectorId/removeParameters:connector(string, path) - The connector typeconnectorId(string, path) - The connector instance ID
- Success Response
- Error Responses
Status:
200 OKDELETE /schedule/all - Remove All Jobs
DELETE /schedule/all - Remove All Jobs
Remove all scheduled crawling jobs for the organization. Admins only: any other caller is refused with
400 “Admin access required”, and nothing is removed.Endpoint: DELETE /api/v1/crawlingManager/schedule/all- Success Response
- Error Responses
Status:
200 OKPOST /:connector/:connectorId/pause - Pause Job
POST /:connector/:connectorId/pause - Pause Job
Pause a scheduled crawling job for a specific connector.Endpoint:
POST /api/v1/crawlingManager/:connector/:connectorId/pauseParameters:connector(string, path) - The connector typeconnectorId(string, path) - The connector instance ID
- Success Response
- Error Responses
Status:
200 OKPOST /:connector/:connectorId/resume - Resume Job
POST /:connector/:connectorId/resume - Resume Job
Resume a paused crawling job for a specific connector.Endpoint:
POST /api/v1/crawlingManager/:connector/:connectorId/resumeParameters:connector(string, path) - The connector typeconnectorId(string, path) - The connector instance ID
- Success Response
- Error Responses
Status:
200 OKGET /stats - Get Queue Statistics
GET /stats - Get Queue Statistics
Retrieve statistics about the crawling job queue.Endpoint:
GET /api/v1/crawlingManager/stats- Success Response
- Error Responses
Status: Statistics Fields:
200 OKwaiting- Number of jobs waiting to be processedactive- Number of jobs currently being processedcompleted- Number of completed jobsfailed- Number of failed jobsdelayed- Number of delayed jobspaused- Number of paused jobsrepeatable- Number of repeatable/scheduled jobstotal- Total number of jobs across all states
Data Types
- Enums
- Interfaces
System Configuration
- Rate Limiting
- Job Lifecycle
- Sync Events
The API includes built-in rate limiting and concurrency controls:
- Maximum 5 concurrent jobs per queue
- Jobs are automatically retried with exponential backoff (5000ms initial delay) on failure
- Stalled jobs are detected after 30 seconds
- Maximum 3 retry attempts for failed jobs
- Job history retention: when a run finishes, the queue keeps at most the 10 most recent completed runs and the 10 most recent failed runs, counted across the whole queue rather than per connector. Removing or replacing a connector’s schedule also trims that connector instance’s finished runs to its 10 most recent. Completed runs are kept in the queue and counted by GET /stats, but Get Job Status and Get All leave them out; failed runs appear in all three
- Jobs are removed and recreated when updating schedules (no job modification). The new schedule’s time, cron expression and timezone are checked first, so a schedule refused for those reasons never removes the one already in place. A schedule sent with
isEnabled: falsestill clears the current one - Removing or pausing a schedule also cancels its runs that are waiting to start, including one-time runs. If the job queue can’t be reached, the request fails with
500instead of reporting success, and a schedule is marked paused only after its runs are removed
Error Handling
Errors raised by these endpoints use the API’s standard error shape:code is HTTP_ followed by the status name (HTTP_BAD_REQUEST, HTTP_UNAUTHORIZED, HTTP_FORBIDDEN, HTTP_NOT_FOUND), or VALIDATION_ERROR when the request doesn’t match the schema. A 500 comes in two kinds. When the job queue can’t be reached, the code is INTERNAL_ERROR with the message “Something went wrong on PipesHub’s side. Please try again; if it keeps happening, ask your admin for help.” When the queue answered but removing a schedule or a queued run failed, the code is HTTP_INTERNAL_SERVER_ERROR with a specific message that says what may still be in place; each endpoint’s tab lists them. requestId is included when the request has one. The only exception is the 404 for a connector with no scheduled job, shown under Get Job Status.
Common HTTP Status Codes:
200- Success201- Created (for scheduling jobs)400- Bad Request (validation errors, invalid configuration, or a non-admin calling Remove All)401- Unauthorized (missing or invalid authentication)403- Forbidden (not allowed to manage this connector, or an OAuth token missing the scope)404- Not Found (connector instance not found, or no scheduled job)500- Internal Server Error