pydo.genai.create_knowledge_base_data_source()

Generated on 7 Jul 2026 from pydo version v0.39.0

Usage

client.genai.create_knowledge_base_data_source(
    knowledge_base_uuid="\"123e4567-e89b-12d3-a456-426614174000\"",
    body={
        "aws_data_source": {...},
        "chunking_algorithm": "CHUNKING_ALGORITHM_UNKNOWN",
        "chunking_options": {...},
        ...,
    },
)
Returns JSONRaises HttpResponseError

Description

To add a data source to a knowledge base, send a POST request to /v2/gen-ai/knowledge_bases/{knowledge_base_uuid}/data_sources.

Parameters

knowledge_base_uuid string required

Knowledge base id

aws_data_source object optional

AWS S3 Data Source

Show child properties
bucket_name string optional

Example: example name

Spaces bucket name

item_path string optional

Example: example string

key_id string optional

Example: 123e4567-e89b-12d3-a456-426614174000

The AWS Key ID

region string optional

Example: example string

Region of bucket

secret_key string optional

Example: example string

The AWS Secret Key

chunking_algorithm string optional

One of: CHUNKING_ALGORITHM_UNKNOWN, CHUNKING_ALGORITHM_SECTION_BASED, CHUNKING_ALGORITHM_HIERARCHICAL, CHUNKING_ALGORITHM_SEMANTIC, CHUNKING_ALGORITHM_FIXED_LENGTH

Default: CHUNKING_ALGORITHM_UNKNOWN

chunking_options object optional
Show child properties
child_chunk_size integer optional

Example: 350

max_chunk_size integer optional

Example: 750

Common options

parent_chunk_size integer optional

Example: 1000

Hierarchical options

semantic_threshold number optional

Example: 0.5

Semantic options

knowledge_base_uuid string optional

Example: "12345678-1234-1234-1234-123456789012"

Knowledge base id

spaces_data_source object optional

Spaces Bucket Data Source

Show child properties
bucket_name string optional

Example: example name

Spaces bucket name

item_path string optional

Example: example string

region string optional

Example: example string

Region of bucket

web_crawler_data_source object optional

WebCrawlerDataSource

Show child properties
base_url string optional

Example: example string

The base url to crawl.

crawling_option string optional

Options for specifying how URLs found on pages should be handled.

- UNKNOWN: Default unknown value
- SCOPED: Only include the base URL.
- PATH: Crawl the base URL and linked pages within the URL path.
- DOMAIN: Crawl the base URL and linked pages within the same domain.
- SUBDOMAINS: Crawl the base URL and linked pages for any subdomain.
- SITEMAP: Crawl URLs discovered in the sitemap.

One of: UNKNOWN, SCOPED, PATH, DOMAIN, SUBDOMAINS, SITEMAP

Default: UNKNOWN

embed_media boolean optional

Example: True

Whether to ingest and index media (images, etc.) on web pages.

exclude_tags array of strings optional

Example: ['example string']

Declaring which tags to exclude in web pages while webcrawling

Request Sample

Show Request Sample
import os
from pydo import Client

client = Client(token=os.environ.get("DIGITALOCEAN_TOKEN"))

req = {
  "aws_data_source": {
    "bucket_name": "example name",
    "item_path": "example string",
    "key_id": "123e4567-e89b-12d3-a456-426614174000",
    "region": "example string",
    "secret_key": "example string"
  },
  "chunking_algorithm": "CHUNKING_ALGORITHM_UNKNOWN",
  "chunking_options": {
    "child_chunk_size": 350,
    "max_chunk_size": 750,
    "parent_chunk_size": 1000,
    "semantic_threshold": 0.5
  },
  "knowledge_base_uuid": "\"12345678-1234-1234-1234-123456789012\"",
  "spaces_data_source": {
    "bucket_name": "example name",
    "item_path": "example string",
    "region": "example string"
  },
  "web_crawler_data_source": {
    "base_url": "example string",
    "crawling_option": "UNKNOWN",
    "embed_media": True,
    "exclude_tags": [
      "example string"
    ]
  }
}

resp = client.genai.create_knowledge_base_data_source(knowledge_base_uuid="\"123e4567-e89b-12d3-a456-426614174000\"", body=req)

Response Example

Show Response Example
{
  "knowledge_base_data_source": {
    "aws_data_source": {
      "bucket_name": "example name",
      "item_path": "example string",
      "region": "example string"
    },
    "bucket_name": "example name",
    "chunking_algorithm": "CHUNKING_ALGORITHM_UNKNOWN",
    "chunking_options": {
      "child_chunk_size": 350,
      "max_chunk_size": 750,
      "parent_chunk_size": 1000,
      "semantic_threshold": 0.5
    },
    "created_at": "2023-01-01T00:00:00Z",
    "dropbox_data_source": {
      "folder": "example string"
    },
    "file_upload_data_source": {
      "original_file_name": "example name",
      "size_in_bytes": "12345",
      "stored_object_key": "example string"
    },
    "google_drive_data_source": {
      "folder_id": "123e4567-e89b-12d3-a456-426614174000",
      "folder_name": "example name"
    },
    "item_path": "example string",
    "last_datasource_indexing_job": {
      "completed_at": "2023-01-01T00:00:00Z",
      "data_source_uuid": "123e4567-e89b-12d3-a456-426614174000",
      "error_details": "example string",
      "error_msg": "example string",
      "failed_item_count": "12345",
      "indexed_file_count": "12345",
      "indexed_item_count": "12345",
      "removed_item_count": "12345",
      "skipped_item_count": "12345",
      "started_at": "2023-01-01T00:00:00Z",
      "status": "DATA_SOURCE_STATUS_UNKNOWN",
      "total_bytes": "12345",
      "total_bytes_indexed": "12345",
      "total_file_count": "12345"
    },
    "region": "example string",
    "spaces_data_source": {
      "bucket_name": "example name",
      "item_path": "example string",
      "region": "example string"
    },
    "updated_at": "2023-01-01T00:00:00Z",
    "uuid": "123e4567-e89b-12d3-a456-426614174000",
    "web_crawler_data_source": {
      "base_url": "example string",
      "crawling_option": "UNKNOWN",
      "embed_media": true,
      "exclude_tags": [
        "example string"
      ]
    }
  }
}

More Information

See /v2/gen-ai/knowledge_bases/{knowledge_base_uuid}/data_sources in the API reference for additional detail on responses, headers, parameters, and more.

We can't find any results for your search.

Try using different keywords or simplifying your search terms.