Skip to main content

What is a dataset in Cognee?

A dataset is a named container that groups documents and their metadata. It is the main boundary for:
  • Organizing content
  • Running pipelines
  • Applying permissions
Operations that write to a dataset — such as remember, improve, and memify — fall back to a dataset named main_dataset when you don’t name one. Your first write creates it, so you can start without deciding on a dataset layout up front.
Dataset isolation requires specific configuration. See permissions system for details on access control requirements and supported database setups.
  • Remember:
    • Direct new content into a specific dataset (by name or ID)
    • If it doesn’t exist, Cognee creates it and associates your permissions
    • Items ingested are linked to that dataset and deduplicated within it
  • Improve:
    • Runs enrichment against a chosen dataset
    • Loads the dataset’s existing graph, checks rights, and runs the improvement pipeline in dataset scope
    • Lets you deepen or bridge memory without re-ingesting the source data
  • Recall:
    • Queries can be scoped by dataset
    • Results and metrics remain separated by dataset
  • Forget:
    • Removes memory at item, dataset, or full-user scope
    • Uses dataset permissions to decide what the current user can remove

Dataset name vs dataset id

Every dataset carries two identifiers, and they are not interchangeable:
  • dataset_name — the human-readable string you choose ("finance", "main_dataset"). Names cannot contain spaces or dots.
  • dataset_id — the dataset’s UUID, stored as its primary key.
The id is derived from the name, not random: it is a uuid5 of the dataset name combined with the owner’s user id and tenant id. So the same name, used by the same user, always resolves to the same dataset — but the same name used by a different user or tenant resolves to a different UUID. Datasets are shared across users by id only, never by name: to reach a dataset shared with you, pass its dataset_id, because passing the name would resolve to (or create) a separate dataset of your own. On write paths, an unknown name creates a new dataset; an unknown id raises DatasetNotFoundError.

Which identifier each operation accepts

To go from a name to an id, list your datasets:

Access control

  • Permissions (read, write, share, delete) are enforced at the dataset level
  • Share one dataset with a team, keep another private
  • Independently manage who can modify or distribute content

Incremental processing

  • Processing status is tracked per dataset
  • After you remember more data, the underlying cognify step focuses on new or changed items
  • Skips what’s already completed for that dataset

Datasets vs NodeSets

Datasets scope storage, permissions, and pipeline execution; NodeSets are semantic tags within a dataset.
  • During remember(), you can label items with one or more NodeSet names (e.g., “AI”, “FinTech”)
  • The underlying graph-building step propagates those labels into the graph by creating NodeSet nodes and linking derived chunks and entities via belongs_to_set relationships
  • This lets you slice a single dataset’s graph by topic or team without creating new datasets, while dataset-level permissions still control overall access

Remember

Direct content into a dataset

Improve

Enrich memory within a dataset

Recall

Scope queries by dataset

Forget

Remove datasets and data