Skip to main content
A minimal guide to using S3 (or S3-compatible, e.g., MinIO) to ingest data and/or store Cognee’s internal files. Before you start:
  • Complete Quickstart to understand basic operations
  • Ensure you have LLM Providers configured
  • Have S3 credentials and access to an S3 bucket

What S3 Storage Does

  • Ingest from S3: Pass s3://... paths to cognee.add() to load data directly from S3
  • Store Cognee data on S3: Set your data/system roots to S3 URLs to keep all files on S3
  • S3-compatible: Works with MinIO and other S3-compatible services

Prerequisites

Install with AWS extra if needed (boto3/s3fs) and add credentials to .env:

Option A: Ingest from S3

Pass S3 URIs (files or prefixes) directly to remember(). Directories/prefixes expand to files when credentials are set.
This loads data directly from S3 using the s3:// URI. remember() expands prefixes, reads the S3 objects, and builds retrieval-ready memory for each target dataset.
This simple example uses S3 paths for demonstration. In practice, you can mix S3 files with local files, use dataset scoping, and apply custom loaders. The same remember() flow works with S3 paths.

Option B: Store Cognee Data on S3

Keep Cognee’s generated files (text copies, system files) on S3 by pointing roots to S3 URLs. Add this to your .env:
This configures Cognee to store all its internal files (processed data, system files) on S3 instead of locally.
Cognee chooses S3 storage when roots start with s3:// (or when STORAGE_BACKEND=s3 and both roots are S3 URLs). If AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY are not set in .env, Cognee falls back to boto3/s3fs’s default credential chain for native AWS S3 deployments instead of erroring. See Object Storage for provider-specific setup details.

Performance notes

Ingesting a file over the S3 backend costs 4 S3 requests per file (1 PUT, 2 HEAD, 1 GET), down from 13. Cognee now uploads the payload once instead of twice and hashes it from the bytes it already holds, rather than downloading the object back several times to recompute a hash of content it just wrote. Uploads are written in 4 MB chunks rather than read into memory whole, so the peak memory one upload adds is bounded by the chunk size instead of the file size — large files no longer scale RAM with their length. Inside DATA_ROOT_DIRECTORY, an uploaded source file is stored under a content-addressed key, <content_md5>/<original_filename>. Expect one hash-named prefix per distinct payload rather than a flat listing. Derived text keeps its existing flat text_<md5>.txt name. See Hash-based file storage for what this means for deduplication. Cognee raises botocore’s per-client connection pool to 32, above the default per-dataset item concurrency of 20, so concurrent items are limited by the network rather than by the pool. This is a fixed internal value — there is no environment variable for it.

Core Concepts

Understand knowledge graph fundamentals

Setup Configuration

Configure providers and databases

API Reference

Explore API endpoints