Index GDrive File

Index a single Google Drive file (by file ID) into a collection (headless).

Path parameters

collection_namestringRequired

Headers

authorizationstring or nullOptional

Request

This endpoint expects an object.
service_account_jsonstringRequired

Google Cloud service-account JSON (as a string). Its OAuth client ID must be authorized for the drive.readonly scope in the customer’s Workspace Domain-wide Delegation settings.

subject_emailstringRequired
Email of the Workspace user to impersonate. Must belong to the domain that authorized the service account. Drive content is read as this user.
file_idstringRequired

Google Drive file ID (the ‘/d/<id>/’ segment of a file URL). Native Google Docs/Sheets/Slides are exported to PDF/XLSX automatically.

processing_typeenumRequired

Document processing type. ‘advanced’ uses agentic OCR with AI-enhanced extraction for complex layouts, tables, figures, charts, and documents containing images. ‘basic’ provides reliable OCR optimized for general document indexing and high-volume processing.

Allowed values:
skip_existingbooleanOptionalDefaults to true

When true, files already indexed in the collection are skipped and will not be re-indexed with incoming changes. When false, all incoming files are indexed regardless of whether they already exist.

mask_piibooleanOptionalDefaults to false

When true, detected PII (emails, phone numbers, SSNs, credit cards, names, and locations) is masked in the parsed content before it is embedded and stored — replaced with entity tags like <PERSON> and <EMAIL_ADDRESS>. For images (including images embedded in PDFs), PII text visible in the image is also pixel-redacted. Opt-in; defaults to false, which leaves content unchanged.

pii_engineenumOptionalDefaults to kev

Masking engine for this job. kev (default): the Kev classifier model on Captain’s infrastructure; supports pii_fields and pii_instructions; no fallback, so a file kev cannot mask fails to index and the rest of the job continues. jev: TypeSafe’s hosted Jev model, slightly more accurate, best-effort availability; supports custom fields and instructions and may list fallbacks in pii_fallback. presidio: the legacy pattern engine, built-in categories only (rejects pii_fields or pii_instructions with 422), never a fallback. Requires mask_pii: true. The earlier names captain-jev and captain-presidio keep working as before; new requests should use kev, jev or presidio.

Allowed values:
pii_fallbackobject or booleanOptional

Ordered fallback engines for jev, as {"engines": [...], "retry_budget_seconds": n}. Omitted or an empty engines list means no fallback. Only valid with pii_engine: "jev"; with kev or presidio a non-empty list answers 422. On the multipart file upload endpoint, send it as a JSON string form field. With the earlier engine name captain-jev it is true (default, fall back to captain-presidio) or false.

pii_fieldslist of objects or nullOptional

Additional PII categories to mask for this job, on top of the built-in set. Requires mask_pii to be true. Up to 20 fields. Each masked value is replaced with the field’s tag in angle brackets and appears in the masking report under the field’s name.

pii_instructionsstring or nullOptional<=1000 characters

Plain-language guidance that steers what counts as personal data for this job, for the built-in categories and any pii_fields. For example: ‘Names of hospital staff may stay; patient names, record numbers and bed assignments must be masked.’ Requires mask_pii to be true. Up to 1,000 characters.

overwrite_existingbooleanOptionalDefaults to false

When true, files that already exist in the collection are re-indexed and replaced with zero downtime: the new version is built alongside the live one and atomically swapped in when complete, so the previous version keeps serving search results throughout the rebuild. The document keeps the same document_id across overwrites, and its status reads ‘updating’ in the document listing while the rebuild runs. Requires skip_existing=false. Setting both to true returns a 400 error.

transcription_languagestring or nullOptional

AWS Transcribe language code for the spoken audio (e.g. ‘es-US’, ‘pt-BR’). Omit to auto-detect per file. Video and audio files only. Supported codes: https://docs.aws.amazon.com/transcribe/latest/dg/supported-languages.html

custom_metadatamap from strings to strings or integers or doubles or booleans or lists of strings or nullOptional

Custom metadata to attach to all chunks from this file. Keys must be strings. Values: str, int, float, bool, or List[str].

parsing_scriptstring or nullOptional

Relative path to a JS parsing script for JSON files (e.g. ‘research/paper-parser’).

Response

Successful Response
job_idstring
statusstringOptionalDefaults to pending
custom_metadatamap from strings to any or nullOptional

The custom_metadata Captain accepted for this job, echoed back as validated. Null when none was supplied.

Errors

400
Bad Request Error
© 2026 Captain