Migrating a large amount of data? Chat with us and we’ll help you plan it.
What is COGX?
COGX (the Cognee eXchange format) is a common shape that all memory imports are translated into before they enter Cognee. Instead of writing one importer for every memory tool, Cognee defines a single intermediate format and a single loader:cognee.export() writes, so the same
format powers backup, restore, and Cognee-to-Cognee migration. For a full
breakdown of the format and its record kinds, see
COGX Exchange Format.
You never construct COGX records by hand. You hand a source object to
cognee.remember() and it does the rest:
Quickstart: import from Mem0
The import itself is a single call — construct aMem0Source from your export
and hand it to cognee.remember():
Mem0Source, pass it to
cognee.remember(), then query with cognee.recall(). Your Mem0 memories are
now a queryable Cognee knowledge graph.
Simple runnable example
Simple runnable example
This script uses a small inline sample in the exact shape Mem0 returns, so
you can run it without a Mem0 account. Swap in real data using the patterns at
the bottom of this page.
Requires
LLM_API_KEY to be set (in .env or your environment) — both the
import and recall() use the LLM.migrate_from_mem0.py
examples/demos/ingestion_and_migration/migrate_from_mem0/migrate_from_mem0.py.
It imports a bundled sample export of four mem0 memories
(data/mem0_export.json)
in preserve mode — calling cognee.cognify() right after, since a preserve
import alone isn’t recall-queryable (see Import modes below) —
and verifies the result with two recall() queries, then clears everything,
re-imports the same export in re-derive mode, and queries again so you can
compare the two modes side by side. It finishes with forget(everything=True)
so re-runs start clean. Run it with:
Import modes
Every source accepts amode argument that controls how much work Cognee does
on import. Pick it based on whether your source already has an extracted graph
and how much you want to spend on LLM calls.
re-derive mode for Mem0 imports when you want Cognee to build a knowledge
graph from those memories. preserve is also valid for Mem0: each memory is
stored as a raw data item with zero LLM calls — the cheapest way to get the
records in, ready for a later cognify() run.
That makes preserve a two-call pattern whenever you want to query the
imported memories: remember() lands the raw records, and a separate
cognify() builds the graph recall() reads from.
recall() returns nothing for the imported memories. With
re-derive or hybrid the extraction is part of the import, so no extra
cognify() call is needed. The
runnable tutorial
imports the same export both ways so you can compare.
Other sources you can import from
Every source has the same interface — construct it from a file path (or in-memory data) and pass it tocognee.remember().
LangMem (JSON memory exports)
LangMem (JSON memory exports)
LangMemSource reads a LangMem memory export and imports each item as an
atomic memory record. It accepts a file path, an already-parsed list, or a
dict wrapping the list under memories, results, items, or data — any
other shape raises ValueError.Those aliases are scanned in that order, and only dict items count as records.
An alias that is present but empty does not shadow a populated one later in the
order, so {"memories": [], "results": [...]} imports the results records.
A wrapper whose recognized aliases are all empty imports zero records without
raising.Per memory:- content is the first string found among
content,text,memory,data, andmessage; items with none of those are skipped - scope takes
user_id(falling back tonamespace), plusagent_id,session_id, andrun_idwhen present - categories accepts either a single string or a list
- timestamps are read from
created_at/createdAt/timestampandupdated_at/updatedAt - any
metadatais carried over nested underlangmem_metadata
re-derive mode.Letta / MemGPT (.af agent files)
Letta / MemGPT (.af agent files)
LettaSource reads a Letta Agent File (.af, a JSON serialization of one
or more agents) and imports:- core memory blocks → memory blocks in the graph (
COGXMemoryBlock) - message history → one conversation episode per agent (
COGXEpisode) - archival memory → one document per passage (
COGXDocument)
- text is read from
content, falling back to thetextalias whencontentis missing or explicitlynull— Letta serializers that write unset fields asnullrather than omitting them import correctly. An explicitly emptycontent("") counts as a message with no text and does not fall through totext - either key may hold a plain string or a list of typed parts, of which only the text parts are imported
- messages that end up with no text, and messages whose role is
systemortool, are skipped; if that leaves an agent with no messages, no conversation episode is emitted for it
Zep / Graphiti (graph exports)
Zep / Graphiti (graph exports)
ZepSource reads a JSON export of a Zep or Graphiti knowledge graph and
imports:- episodes (verbatim ingested content) →
COGXEpisode - entity nodes →
COGXEntity - relation edges (“facts”) →
COGXFact, including their bi-temporalvalid_at/invalid_atvalidity windows
session_id the same way: from
group_id, falling back to a session_id key when group_id is absent. When
a record carries both, group_id wins.It defaults to hybrid mode because Zep/Graphiti keep both verbatim episodes
and a derived graph.Another Cognee instance (COGX archive)
Another Cognee instance (COGX archive)
COGXArchiveSource re-imports an archive produced by cognee.export(..., format="cogx"). This is the restore half of backup/restore and the receiving
end of Cognee-to-Cognee migration. It defaults to preserve mode (zero-LLM),
because a Cognee archive already carries a fully extracted graph.Runnable tutorial: Letta + Zep
A runnable walkthrough of the two accordions above ships with the repo:examples/demos/ingestion_and_migration/migrate_from_letta_and_zep/migrate_from_letta_and_zep.py.
It runs in three parts, each starting from forget(everything=True):
- Letta — imports a bundled sample agent file
(
sample_letta_dump.json: one agent with two core memory blocks, a three-message history, and two archival passages) inre-derivemode, then verifies it with tworecall()queries. - Zep — imports a bundled sample graph export
(
sample_zep_dump.json: two episodes, four entities, and four facts, each carrying avalid_attimestamp and none aninvalid_atend — every sample fact is still valid) inhybridmode, then runs threerecall()queries against it. - Both together — imports the same two dumps into one graph and queries across them, so you can see a combined Letta + Zep memory answer questions that draw on both sources.
forget(everything=True), so it leaves nothing behind
and re-runs start clean.
Because the dumps are bundled, you can run all three parts without a Letta or
Zep account — only LLM_API_KEY is needed (re-derive, hybrid, and recall()
all call the LLM).
Write your own source
If your memory tool isn’t listed above, you can add it yourself. A source is a small adapter class: subclassMemorySource, name the system it reads from,
and implement one async generator — records() — that yields COGX records.
Everything else (loading, modes, deduplication, graph storage) is handled by
the shared machinery, exactly as for the built-in sources.
Everything you need is importable from cognee.migration: the MemorySource
base class and the record models (COGXDocument, COGXEpisode, COGXTurn,
COGXEntity, COGXFact, COGXMemory, COGXMemoryBlock, COGXScope).
Here is a complete source for a hypothetical notes app that exports a JSON
list of {"id", "text", "user", "tags"} objects:
cognee.remember() like any
built-in source:
Choosing record kinds
Pick the record type that matches what the source holds; the COGX concept page describes each in detail. What happens to a record depends on its kind and the import mode:
Nothing you emit is silently dropped except raw nodes and entities with no
description, both of which carry no standalone text for re-derive to
re-extract.
A COGXFact references its endpoints by subject_ref / object_ref — use the
external_id of an entity record you also emit, or a plain entity name. A
plain-name reference that doesn’t match an emitted entity becomes a new entity
of that name, so a source can emit facts on their own. A reference that looks
like a UUID but matches no emitted record is skipped and logged, never turned
into an entity named by a UUID. A fact’s valid_at / invalid_at validity
window is preserved as edge properties in the graph.
Rules the loader relies on
- Stable
external_ids make re-import idempotent. Each record’s identity in Cognee is derived deterministically from(external_system, external_id), so importing the same export twice doesn’t duplicate data. Use the source system’s own ids; only fall back to synthetic ids (e.g. an index) for records that genuinely have none. records()should be re-callable. The streamingpreserve-mode import passes over the records three times — once to store the raw content, then twice for the graph (nodes first, then facts) — callingrecords()once per pass. File-backed sources get this for free by re-reading the file. If your source is a one-shot cursor (e.g. a live API stream), set the class attributereplayable = Falseto make the loader buffer records instead.- Preserve scope and timestamps. Fill
COGXScope(user_id/agent_id/session_id/run_id) andcreated_at/updated_atwhere the source has them — ownership and time information survive the migration only if the source carries them across. Anything that doesn’t fit a typed field can go in the record’s free-formmetadatadict. - Pick a sensible default mode. Match the import mode to whether your source carries a pre-built graph; callers can always override.
cognee/modules/migration/sources/
— mem0.py is the smallest, and zep.py shows entity/fact emission. If your
source would be useful to others, PRs adding it there (exported in the
package’s __init__.py, with a sample-export test) are welcome.
Export: Cognee → COGX
Migration runs both ways.cognee.export() writes a dataset’s graph to a
portable COGX archive that you can back up, move to another Cognee instance, or
re-import later with COGXArchiveSource:
Using real data
The Mem0 example above uses an inline sample. Here’s how to point any source at real data. From a provider’s client (live API). Fetch with the provider’s own SDK and pass the response straight in — sources accept already-parsed Python lists and dicts, not just file paths:Mem0Source scans a wrapper dict for results, memories, then items, in
that order, counting only dict items as records; an alias that is present but
empty does not shadow a populated one later in the order, so
{"results": [], "memories": [...]} imports the memories records. A wrapper
whose recognized aliases are all empty imports zero records without raising,
and any other shape raises ValueError.
From an exported file. Point the source at the export on disk:
Next steps
remember()
How import and ingestion work under the hood
recall()
Query your migrated memory
COGX Exchange Format
The portable format behind every import and export
Configuration
Configure LLM, embedding, and storage backends