Skip to main content
Ingest structured relational data — databases, CSV files, and dlt resources — directly into cognee’s knowledge graph. Foreign keys become graph edges, tables become schema nodes, and each row becomes a searchable document, all built deterministically from the schema without LLM extraction.

Why Use This Integration

  • Schema-Aware Graphs: Foreign key relationships are preserved as first-class edges in the knowledge graph
  • Deterministic Graph Construction: Structured data bypasses LLM entity extraction — no hallucination risk
  • Mixed Ingestion: Combine structured (dlt) and unstructured (text, PDF) data in the same dataset
  • Multiple Input Modes: Pass explicit dlt resources, CSV file paths, or database connection strings
  • Write Dispositions: Control how data is synced — merge (upsert), append, or replace

Installation

Or with uv:

Quick Start

1. Ingest a dlt Resource

Define a dlt resource and pass it to cognee.remember(). The dlt-specific structured-ingestion options primary_key, write_disposition, SQL query, and max_rows_per_table are accepted by cognee.remember() and forwarded to the underlying ingestion step. After ingestion, use cognee.recall(...) to query the graph.
dlt automatically detects nested structures (like pets inside each user) and creates separate tables with foreign key relationships.
The lower-level cognee.add(...) + cognee.cognify(...) pair still accepts the same dlt kwargs and remains useful when you need to run ingestion and graph building as separate steps. For the runnable end-to-end version of this walkthrough, see examples/demos/ingestion_and_migration/dlt_ingestion_example.py.

2. Build and Query the Graph

Once remember() finishes ingesting and building the graph, use cognee.recall(...) to query it.

Other Input Modes

CSV Files

CSV files are handled by the dlt_csv_loader, which the loader engine registers above the plain-text csv_loader whenever the dlt extra is installed. Selection happens inside the ingest pipeline rather than before it, so the same route applies to every way a CSV can arrive — a local path, a file:// or s3:// location, or an uploaded file:
Each CSV is staged through dlt and becomes one manifest record per original file. The manifest’s identity is derived from the original file name — not the temporary copy an s3:// download or an upload lands in — so it stays stable across runs and a re-add updates the existing record instead of creating a second one. As with any dlt source, the rows skip chunking and LLM entity extraction. Per-call dlt options for CSVs travel through the loader-config channel rather than as remember() keyword arguments. dlt_csv_loader accepts primary_key, write_disposition, max_rows_per_table, and column_value_columns:
To flatten a CSV into plain text instead — the behavior you get automatically on installs without the dlt extra — request csv_loader explicitly for that call:
A CSV that yields no rows now fails loudly with an IngestionError naming the source, rather than being ingested as an empty record. dlt_csv_loader likewise requires the dataset and user context that the ingest pipeline supplies, so it cannot be invoked outside of remember() / add().

Database Connection String

Ingest tables directly from an existing database:
The connection string is a source cognee reads from, never a destination it writes to — ingested rows land in cognee’s own configured stores, including vector embeddings of each row, so they’re searchable by semantic similarity as well as graph traversal. To change where cognee stores memory, configure the graph, vector, and relational providers.
Supported databases via auto-detection: SQLite, PostgreSQL, MySQL, MSSQL, Oracle. Hosted Postgres providers such as Neon work with their standard postgresql:// connection strings; keep provider-required SSL parameters such as ?sslmode=require. Amazon Redshift is also compatible since it speaks the PostgreSQL wire protocol — use a standard postgresql:// connection string pointing to your Redshift endpoint. For Snowflake and Google BigQuery, construct a dlt source directly and pass it to cognee.remember() (see the Cloud Data Warehouses accordion below). You can optionally filter with a SQL WHERE clause:

Mixed Structured + Unstructured

Combine dlt resources with unstructured text in a single dataset:
Structured data creates deterministic graph nodes from the schema, while unstructured text goes through LLM-based entity extraction. Both are combined in the same knowledge graph.

Write Dispositions

Control how data is synced on repeated runs using the write_disposition parameter:
  • replace (default): Drop and recreate tables on each run. Use for full snapshot refreshes.
  • merge: Upsert by primary key — updates existing rows, inserts new ones. Best for data that changes over time.
  • append: Always insert without deduplication. Use for time-series data and event logs.
write_disposition only controls dlt’s staging snapshot. It does not by itself make a re-run pick up rows a source has gained, because the ingestion skip in add() short-circuits an already-ingested source first — see Re-Ingesting a Source and Re-ingesting a source that keeps growing.

How It Works

  1. Source Detection: cognee identifies dlt resources and connection strings in the input. .csv files are not detected here — they are routed by the loader engine’s dlt_csv_loader inside the ingest pipeline (see CSV Files) and join the flow from step 2 onward
  2. Pipeline Execution: A dlt pipeline loads data into a per-dataset staging database
  3. Schema Extraction: Table schemas, primary keys, and foreign keys are extracted
  4. Graph Construction: Each row becomes a document node; foreign keys become edges between nodes
  5. LLM Bypass: Structured rows skip chunking, entity extraction, and summarization — the graph is built entirely from schema metadata
The primary_key parameter controls upsert behavior when you use write_disposition="merge". If not specified, cognee auto-detects from an id column or falls back to the first column.

Bounding Ingestion

Two opt-in caps guard against large sources; both default to 0, meaning no cap.
  • Rows per table: pass the max_rows_per_table kwarg to remember() / add() to bound the per-table row count for a single call, or set the DLT_MAX_ROWS_PER_TABLE environment variable to change the process-wide default.
  • Column-value length: DLT_MAX_COLUMN_VALUE_LENGTH bounds the length of the cell values that become shared ColumnValue nodes — cell-level graph nodes that link rows sharing the same value in a selected column. (Columns are selected with the column_value_columns kwarg — add() only — or DLT_COLUMN_VALUE_COLUMNS; nothing is selected by default.) A positive value skips selected cells longer than that many characters, dropping rather than truncating them — worthwhile for free-text-heavy columns or wildcard ("*") selection, since long one-off values make poor shared nodes and each unique value costs one embedding. Unlike max_rows_per_table, this cap has no per-call kwarg: set the environment variable or the dlt_max_column_value_length ingestion config field. Cognee previously applied a fixed 256-character cap unconditionally; set DLT_MAX_COLUMN_VALUE_LENGTH=256 to keep that behavior.

Re-Ingesting a Source

Each relational dlt source is stored as a single record whose identity is stable: it is derived from the dataset name and the source name (dlt_source:{dataset_name}:{source_name}) and does not depend on the data. Re-running remember() / add() on the same source therefore never creates a second copy — but what happens to the existing record depends on how you call it:
  • Plain re-add (the default): add() is idempotent. A source that has already been ingested keeps its record, content hash, and cognify status as-is — whether or not the upstream data changed — and nothing is reprocessed.
  • Explicit re-ingest: to pick up upstream changes, call add(..., incremental_loading=False, data_cache=False) — the completed-skip runs whenever either flag is on, so both must be off — or use update() with the record’s UUID. The record then updates in place under the same identity (a content hash over the source’s tables and rows tracks the change), so the source is never absent from the store, and the next cognify() purges the source’s previously derived artifacts from the graph and vector stores before re-emitting the current rows. Rows deleted upstream disappear, and changed rows do not keep their stale values alongside the new ones. Between the re-ingesting add() and cognify(), searches still return the source’s previous rows; they are replaced only once the re-cognify completes.
The purge is a real delete, re-authorized as one: re-ingesting a changed dlt source requires delete permission on the dataset. If that permission is missing, the run fails instead of continuing — silently skipping the purge would leave stale rows in the graph and present them as current.

Renaming a Source

Because the identity is keyed on the dataset and source names, renaming either one is a remove + add, not an in-place rename. The new name ingests from scratch as a fresh source, and the records under the old name stay in the dataset — a re-ingest only reconciles the source names it just ingested, so it will not treat the old name’s records as orphans. Delete them explicitly (for example with cognee.forget(...)) if you do not want both. For the same reason, two dlt sources that share a name within a single add() call would resolve to the same identity. cognee raises an error rather than letting one silently overwrite the other, so give each source a distinct name. CSVs are keyed the same way, but through dlt_csv_loader rather than that check: a CSV’s source name is its file-name stem, with runs of characters other than letters, digits, and underscores replaced by _, leading and trailing separators stripped, and the result lowercased. Two CSVs whose stems normalize to the same string (for example Employees 2024.csv and employees_2024.csv) therefore share one manifest record within a dataset, with the later ingest updating the earlier one. Give CSV files names that stay distinct under that normalization.

Foreign Key Resolution

A foreign key becomes a graph edge only when both the source row and the target row are loaded in the same ingestion run. Two edge cases are worth knowing about — cognee now logs a warning in each so they are diagnosable rather than silent:
  • Target row not loaded: if a foreign key points at a row that wasn’t ingested — most commonly because the target table hit a max_rows_per_table cap you set — the reference is dropped and no edge is created. The warning identifies the dropped references as source_table.column -> ref_table:value. If you see missing edges, raise max_rows_per_table so the referenced rows are included.
  • Duplicate primary keys within a table: if multiple rows in a table share the same primary key, foreign key edges that target that key resolve to the last such row loaded; earlier rows with the same key are shadowed for FK targeting. The warning names the affected table and pk.

Connectors

Connectors are dlt sources for a specific system. The list below keeps the current connector packages visible; the routing details are tucked away for reference.
Cognee supports two DLT connector modes:
  • Relational connectors take the default dlt path described above: each row becomes a schema-context document and foreign keys become edges, all built deterministically without LLM extraction.
  • Document-mode connectors opt each row into normal cognify instead: the row is turned into a text document that goes through LLM entity extraction, just like unstructured text passed to remember().
A dlt source opts into document mode by setting the cognee_document_source attribute (via the document_source_tag() helper in cognee.tasks.ingestion.dlt_utils) to a short source tag. cognee then routes every row from that source through cognify rather than the relational schema-context path:
  • Each row is built from its title and content columns (rendered as # {title}\n\n{content}, or just the content when there is no title), with optional url and id columns preserved in metadata.
  • external_metadata["source"] is set to the connector’s own tag (for example "notion") instead of "dlt", alongside title and, when present, url and external_id (from the row’s id).
Because the tag is connector-provided, the shared ingestion engine stays connector-agnostic: a connector declares its own nature rather than being hard-coded by name.Sync and orphan cleanup. Document sources read back their full current snapshot (max_rows_per_table=0) and honor the write_disposition you pass: use replace for snapshot sources with no delete feed and merge with a hard-delete tombstone column for incremental sources that emit real deletions. primary_key defaults to id. Orphan cleanup is scoped to the source tag, so reconciling a document source only removes that source’s rows, and relational ("dlt") rows and other sources’ rows are never cross-deleted in a mixed dataset. Cleanup is skipped when the fresh read-back is empty, so an empty snapshot is treated as a failed sync rather than a signal to delete everything.
Orphan cleanup now runs in the foreground of add() / remember(): blocking runs execute it synchronously after the fresh rows are committed, so upstream deletions are reflected within the same call. Background runs (run_in_background=True) perform it up front instead.

Gmail Connector

The Gmail connector is a first-class dlt source that turns your inbox into cognee memory. It reuses the same remember() + dlt path described above, so it gets incremental re-sync and forget-on-delete for free. gmail_source() returns a dlt resource that you hand directly to cognee.remember().
This connector reads the content of your email. It is strictly opt-in — nothing is fetched until you construct a source and call remember(). Scope what you ingest with label_ids / query, keep the OAuth token file (token.json) private, and prefer a dedicated dataset so you can wipe it with a single cognee.forget(...).
The Gmail connector ships as the standalone community package cognee-community-connector-gmail, maintained in the cognee-community repository, so core stays free of the Google client SDKs. Install it with pip install cognee-community-connector-gmail, then import gmail_source from cognee_community_connector_gmail as shown below.

Installation

Or with uv:
The standalone package pulls in dlt[sqlalchemy], google-api-python-client, google-auth, and google-auth-oauthlib. The Google client libraries are imported lazily, so the core cognee install stays slim.

One-Time OAuth Setup

The connector authenticates with Gmail via the OAuth2 installed-app (Desktop app) flow using the read-only scope https://www.googleapis.com/auth/gmail.readonly — it never modifies your mailbox.
  1. In the Google Cloud Console, enable the Gmail API, configure an OAuth consent screen (add yourself as a test user), and create an OAuth 2.0 Client ID of type Desktop app.
  2. Download the client-secret JSON and save it as credentials.json (or point credentials_path at it).
  3. The first run opens a browser to consent and caches the resulting user token at token.json (token_path). Later runs reuse and silently refresh that token.

Usage

gmail_source() accepts these keyword-only parameters:The returned resource (gmail_messages) is preconfigured with primary_key="id", write_disposition="merge", and an _deleted hard-delete column, so combined with primary_key="id" on remember() it performs idempotent upserts by Gmail message id.
cognee’s dlt ingestion reads at most max_rows_per_table rows from the dlt destination, and the default is 0 — no cap. For a real inbox, keep it unlimited so forget-on-delete compares against the whole synced corpus rather than a truncated window.

How It Works

  • Incremental sync: The first run does a full (label-scoped) backfill and records the mailbox historyId. This cursor is persisted in dlt’s per-resource state, so re-running remember() on the same dataset resumes where it left off — subsequent runs call users.history.list(startHistoryId=...) and emit only the delta (added / changed / deleted messages).
  • Forget-on-delete: Messages reported as deleted or trashed by the History API are emitted with the _deleted hard-delete marker. dlt removes those rows from its destination on merge, and cognee’s existing orphan_cleanup then purges them from the graph, vector, and relational stores.
  • History expiry: Gmail expires history after roughly a week. If the stored historyId is too old, the History API returns a 404; the connector detects this and falls back to a full backfill so memory re-syncs rather than silently stalling.
For a runnable end-to-end walkthrough that demonstrates the initial backfill followed by an incremental sync, see the cognee-community-connector-gmail package in the cognee-community repository.

Use Cases

Load customer, order, and product tables from a database. Foreign keys between tables (e.g., order.customer_id → customer.id) become graph edges, enabling cross-table queries like “Which customers ordered product X?”
Point cognee at CSV exports from analytics tools. Each row becomes a searchable node in the graph, and you can combine them with unstructured reports in the same dataset.
Use write_disposition="append" to stream event batches into cognee without deduplication. Query across the full event history with natural language.
Use write_disposition="merge" to keep cognee’s graph in sync with a live database. Rows that are removed upstream are cleaned up best-effort; any orphaned rows that fail to delete are logged and retried on the next ingest.
Amazon Redshift speaks the PostgreSQL wire protocol, so the standard connection string auto-detection works:
Snowflake requires constructing a dlt sql_database source manually (install snowflake-sqlalchemy first):
The account_identifier is the part before .snowflakecomputing.com in your Snowflake URL (e.g. myorg-myaccount). Omit table_names to ingest all tables in the schema.Google BigQuery works the same way using dlt’s BigQuery connector — construct the source and pass it directly to cognee.remember(). See the dlt sql_database docs for connector-specific setup.

Remember Operation

Learn more about data ingestion in cognee

dlt Documentation

Official dlt documentation and guides