Document ingestion for AI builders

Prepare your knowledge for retrieval.

Choose from 12 source types, 6 vector stores, and 4 embedding providers. Configure the services you need for each project.

Connect sources or bring your files

Browse supported connected sources, select documents, upload files, or crawl a website. Connection methods and monitoring support vary by source.

Confluence

Select pages from your team wiki.

  • OAuth connection
  • Space and page browsing
Setup guide for Confluence
Google Drive

Sync documents and files from selected folders.

  • OAuth connection
  • Folder browsing

Folder monitoring supported

Setup guide for Google Drive
Jira

Prepare selected issues for search and retrieval.

  • OAuth connection
  • Project browsing and JQL search
Setup guide for Jira
Dropbox

Select and sync files from Dropbox folders.

  • OAuth connection
  • File and folder browsing

Folder monitoring supported

Setup guide for Dropbox
Notion

Load accessible workspace pages and database content.

  • OAuth connection
  • Page selection
Setup guide for Notion
Amazon S3

Load files from S3 or compatible object storage.

  • Access key connection
  • Bucket and prefix browsing

Folder monitoring supported

Setup guide for Amazon S3
Supabase Storage

Load files from storage buckets, rather than database tables.

  • Your Supabase project URL and API key
  • Bucket and folder browsing

Folder monitoring supported

Setup guide for Supabase Storage
GitHub

Sync selected repository files, issues, and discussions.

  • OAuth connection
  • Repository browsing
Setup guide for GitHub
Azure Blob Storage

Load documents from storage containers.

  • Connection string
  • Container and directory browsing

Folder monitoring supported

Setup guide for Azure Blob Storage
Google Cloud Storage

Load documents from cloud storage buckets.

  • Service account connection
  • Bucket and directory browsing

Folder monitoring supported

Setup guide for Google Cloud Storage
Website Crawler

Crawl accessible pages on a selected website.

  • URL-based setup
  • Depth control and SSRF protection
Setup guide for Website Crawler
File Upload

Upload local documents for processing.

  • File selection or drag and drop
  • Uploaded files use billable storage
Setup guide for File Upload

Store vectors in your configured database

Provide a reachable endpoint and the required credentials. Self-hosted services must be reachable from the hosted app; a server on your laptop is not automatically accessible.

Supabase pgvector

Store vectors in a configured Supabase PostgreSQL project.

  • Project URL and service key required
  • External projects may require one-time SQL setup
Setup guide for Supabase pgvector
PostgreSQL / pgvector

Connect a public PostgreSQL database with pgvector over TLS.

  • Encrypted project credential
  • Tenant-scoped table and HNSW initialization
Setup guide for PostgreSQL / pgvector
Pinecone

Use a managed vector index with project namespaces.

  • API key required
  • Index initialization
Setup guide for Pinecone
Cloudflare Vectorize

Use a globally distributed Vectorize index for semantic retrieval.

  • Scoped API token
  • Project namespaces and metadata isolation
Setup guide for Cloudflare Vectorize
Qdrant

Use a Qdrant collection on a server reachable by the app.

  • Configure an API key for the sync workflow
  • Collection initialization
Setup guide for Qdrant
Milvus

Use Milvus or a compatible Zilliz Cloud endpoint.

  • Configure a token for the sync workflow
  • Collection initialization
Setup guide for Milvus

Choose an embedding provider

Provider charges are separate. Changing models or vector dimensions can require reprocessing documents and configuring a compatible destination.

OpenAI

Choose an embedding model for your project.

  • text-embedding-3-small / large
  • Your provider API key
Setup guide for OpenAI
Cohere

Choose English or multilingual embeddings.

  • embed-english-v3.0 / embed-multilingual-v3.0
  • Your provider API key
Setup guide for Cohere
Google Gemini

Configure a supported Gemini embedding model.

  • gemini-embedding-001 and other configured models
  • Your provider API key
Setup guide for Google Gemini
Ollama

Use embeddings from your own reachable Ollama service.

  • Configurable model and dimensions
  • Server must be reachable from the hosted app
Setup guide for Ollama

Prepare, maintain, and test your data

Structure-aware chunks
Preserve headings and source links in chunk metadata so retrieved passages retain context.
Scheduled maintenance
Schedule checks of existing documents. Monitor supported folders for new files; timing depends on polling and processing.
Separate projects
Choose embedding and vector-store settings per project. Organize documents with tags and collaborate using team roles.
Search and chat playground
Test retrieval against synced documents before connecting your application. Chat requires a configured LLM provider.
Connection details
Use your configured vector store with external tools. Apply the required organization and project filters and the appropriate access controls.
Enhanced processing
Add AI-generated context to chunks. With configured OCR, extracted image descriptions can become searchable text; provider charges and image storage apply.