Welcome Aboard

Captain File Search API

Captain is a deterministic File Search API for unstructured data.
Connect cloud storage, index files, and retrieve source chunks.

File Search API

  • Deterministic Retrieval: Ask questions in plain English and get retrieved chunks with source metadata.
  • Cloud Storage Integration: Connect S3, GCS, Azure, R2, and other storage providers. Captain processes and indexes files over a single API call.
  • Multi-Tenancy: Organize collections to scope different teams, folders, projects, etc.
  • Chunk-Level Workflow: Use stable chunk IDs, regions, custom metadata, and relations for source-grounded applications.
  • PII Masking: Redact sensitive text and faces before content is indexed. See PII Masking.
  • Stable Document Identity: Every document has a document_id derived from its source identity, which defaults to the object’s URI. Single-file index requests accept a source_identity so the same content indexed from a staging location and later from its canonical location stays one document, and skip_existing matches on identity rather than file name. Each document also records the checksum of the bytes its current version was indexed from.
Complex Docs, Images, and Sheets
Complex Docs, Images, and Sheets

Captain can search across very large documents, text-heavy or visual images, and multi-faceted spreadsheets.


Automatic VLM, OCR, and computer vision pipelines support search over visual and text-heavy content.

Getting Help

Ready to Start?

© 2026 Captain